Search NASASearch

SEARCH · Search NASA

Results for “Data Intensive Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

NASA Center for Climate Simulation (NCCS) Presentation

The NASA Center for Climate Simulation (NCCS) offers integrated supercomputing, visualization, and data interaction technologies to enhance NASA's weather and climate prediction capabilities. It serves hundreds of users at NASA Goddard Space Flight Center, as well as other NASA centers, laboratories, and universities across the US. Over the past year, NCCS has continued expanding its data-centric computing environment to meet the increasingly data-intensive challenges of climate science. We doubled our Discover supercomputer's peak performance to more than 800 teraflops by adding 7,680 Intel Xeon Sandy Bridge processor-cores and most recently 240 Intel Xeon Phi Many Integrated Core (MIG) co-processors. A supercomputing-class analysis system named Dali gives users rapid access to their data on Discover and high-performance software including the Ultra-scale Visualization Climate Data Analysis Tools (UV-CDAT), with interfaces from user desktops and a 17- by 6-foot visualization wall. NCCS also is exploring highly efficient climate data services and management with a new MapReduce/Hadoop cluster while augmenting its data distribution to the science community. Using NCCS resources, NASA completed its modeling contributions to the Intergovernmental Panel on Climate Change (IPCG) Fifth Assessment Report this summer as part of the ongoing Coupled Modellntercomparison Project Phase 5 (CMIP5). Ensembles of simulations run on Discover reached back to the year 1000 to test model accuracy and projected climate change through the year 2300 based on four different scenarios of greenhouse gases, aerosols, and land use. The data resulting from several thousand IPCC/CMIP5 simulations, as well as a variety of other simulation, reanalysis, and observationdatasets, are available to scientists and decision makers through an enhanced NCCS Earth System Grid Federation Gateway. Worldwide downloads have totaled over 110 terabytes of data.

Webster, William P.

Provenance Challenges for Earth Science Dataset Publication

Modern science is increasingly dependent on computational analysis of very large data sets. Organizing, referencing, publishing those data has become a complex problem. Published research that depends on such data often fails to cite the data in sufficient detail to allow an independent scientist to reproduce the original experiments and analyses. This paper explores some of the challenges related to data identification, equivalence and reproducibility in the domain of data intensive scientific processing. It will use the example of Earth Science satellite data, but the challenges also apply to other domains.

Tilmes, Curt

Telecommunications and data acquisition systems support for the Viking 1975 mission to Mars

The background for the Viking Lander Monitor Mission (VLMM) is given, and the technical and operational aspects of the tracking and data acquisition support that the Network was called upon to provide are described. An overview of the science results obtained from the imaging, meteorological, and radio science data is also given. The intensive efforts that were made to recover the mission are described.

Mudgway, D. J.

Software architecture for large scale, distributed, data-intensive systems

This paper presents our experience with OODT, a novel software architectual style, and middlware-based implementation for data-intensive systems. To date, OODT has been successfully evaluated in several different science domains including Cancer Research with the National Cancer Institute (NCI), and Planetary Science with NASA's Planetary Data System (PDS).

software architecture

Leveraging Data Intensive Computing to Support Automated Event Services

A large portion of Earth Science investigations is phenomenon- or event-based, such as the studies of Rossby waves, mesoscale convective systems, and tropical cyclones. However, except for a few high-impact phenomena, e.g. tropical cyclones, comprehensive records are absent for the occurrences or events of these phenomena. Phenomenon-based studies therefore often focus on a few prominent cases while the lesser ones are overlooked. Without an automated means to gather the events, comprehensive investigation of a phenomenon is at least time-consuming if not impossible. An Earth Science event (ES event) is defined here as an episode of an Earth Science phenomenon. A cumulus cloud, a thunderstorm shower, a rogue wave, a tornado, an earthquake, a tsunami, a hurricane, or an EI Nino, is each an episode of a named ES phenomenon," and, from the small and insignificant to the large and potent, all are examples of ES events. An ES event has a finite duration and an associated geolocation as a function of time; its therefore an entity in four-dimensional . (4D) spatiotemporal space. The interests of Earth scientists typically rivet on Earth Science phenomena with potential to cause massive economic disruption or loss of life, but broader scientific curiosity also drives the study of phenomena that pose no immediate danger. We generally gain understanding of a given phenomenon by observing and studying individual events - usually beginning by identifying the occurrences of these events. Once representative events are identified or found, we must locate associated observed or simulated data prior to commencing analysis and concerted studies of the phenomenon. Knowledge concerning the phenomenon can accumulate only after analysis has started. However, except for a few high-impact phenomena. such as tropical cyclones and tornadoes, finding events and locating associated data currently may take a prohibitive amount of time and effort on the part of an individual investigator. And even for these high-impact phenomena, the availability of comprehensive records is still only a recent development. A major reason for the lack of comprehensive ,records for the majority of the ES phenomena is the perception that they do not pose immediate and/or severe threat to life and property and are thus not consistently tracked. monitored, and catalogued. Many phenomena even lack commonly accepted criteria for definitions. However. the lack of comprehensive records is also due to the increasingly prohibitive volume of observations and model data that must be examined. NASA Earth Observing System Data Information System (EOSDIS) alone archives several petabytes (PB) of satellite remote sensing data and steadily increases. All of these factors contribute to the difficulty of methodically identifying events corresponding to a given phenomenon and significantly impede systematic investigations. In the following we present a couple motivating scenarios, demonstrating the issues faced by Earth scientists studying ES phenomena.

Clune, Thomas L.

Earth Radiation Budget Experiment - Preliminary seasonal results

Over the previous four years the Earth Radiation Budget Experiment (ERBE) instruments have been gathering data on two satellites, the Earth Radiation Budget Satellite and the the operational NOAA-9 satellite. The ERBE science team recently completed the validation of an initial sampling of these data involving intensive examination of data in four months during 1985 and 1986. The data being placed in the National Space Science Data Center to acquaint the scientific community with their availability are discussed. The ERBE archival data products are also presented.

Barkstrom, Bruce R.

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele

Representation-Independent Iteration of Sparse Data Arrays

An approach is defined that describes a method of iterating over massively large arrays containing sparse data using an approach that is implementation independent of how the contents of the sparse arrays are laid out in memory. What is unique and important here is the decoupling of the iteration over the sparse set of array elements from how they are internally represented in memory. This enables this approach to be backward compatible with existing schemes for representing sparse arrays as well as new approaches. What is novel here is a new approach for efficiently iterating over sparse arrays that is independent of the underlying memory layout representation of the array. A functional interface is defined for implementing sparse arrays in any modern programming language with a particular focus for the Chapel programming language. Examples are provided that show the translation of a loop that computes a matrix vector product into this representation for both the distributed and not-distributed cases. This work is directly applicable to NASA and its High Productivity Computing Systems (HPCS) program that JPL and our current program are engaged in. The goal of this program is to create powerful, scalable, and economically viable high-powered computer systems suitable for use in national security and industry by 2010. This is important to NASA for its computationally intensive requirements for analyzing and understanding the volumes of science data from our returned missions.

James, Mark

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES

Atmospheric Data-Driven Visualization of Air Quality Variation in Megacities

The growth and spread of human settlements and increasing urban density are important processes in global change. Urbanization has accelerated with population growth, and more densely populated urban areas have major effects on the local and regional environments. Air quality in megacities has been a concern for decades. Data for anthropogenic emissions, pollutants, and particulate matter can inform research on the interactions between urban landscapes and the atmosphere, assist policy makers in developing sustainable, healthy environments, and inform the general public. To better assist researchers, students, the public, and policymakers in understanding the annual and seasonal variation of aerosol and gas intensity in megacities, the Science Outreach Team at NASA Langley Research Center’s Atmospheric Science Data Center (ASDC) Distributed Active Archive Center (DAAC) demonstrate data products at the ASDC that can be used to visualize these parameters. The presentation uses data from the ASDC-supported NASA missions Measurements Of Pollution In The Troposphere (MOPITT), Cloud-Aerosol and Infrared Pathfinder Satellite Observation (CALIPSO), Tropospheric Emission Spectrometer (TES), and Multi-angle Imaging Spectro Radiometer (MISR)

Air Quality

Population Viability Analysis (PVA) as a Platform for Predicting Outcomes of Management Options for the Florida Scrub-Jay in Brevard County

The Florida scrub-jay (FSJ) is a species in decline because of extinction debt caused by habitat fragmentation and degradation. Understanding and managing the species and its habitats is challenging due to the complex interactions among the social system, a dynamic mosaic of scrub habitat, and active management. To provide insights into the possible fates of the FSJ populations of mainland, cape, and island sections of Brevard County, we modeled the population dynamics using the Vortex population viability analysis (PVA) software. Vortex is an individual-based simulation that allowed us to include such factors as demographic rates dependent on habitat state, impact of helpers on breeding success, and helper to breeder transition probabilities responding to availability of nearby vacant optimal habitat and vacancies due to the death of breeders. Detailed modeling of the FSJ was possible only because a lot of data about the species and its habitats have been gathered over decades of intensive research. We followed a phased approach to constructing population models that incorporated the best available science and data to address a variety of conservation actions. The first phase focused on constructing a model that incorporates sociobiology and source-sink habitat dynamics. This model allowed us to address some of the most important questions about population size and habitat quality. Once the model framework was in place, we considered the real landscapes and actual local populations, rather than just generic representations of typical FSJ dynamics. After examining the viability of the metapopulations under current conditions, we explored the likely consequences of various management actions that might slow or reverse population declines.

Population Viability Analysis

Expanding Repository Data Available For Sharing and Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

Biology

Harnessing Satellite Data Alone for Mapping Global Thermal Anisotropy

Mapping thermal anisotropy across global lands is critical for advancing a wide range of Earth science studies. However, a comprehensive understanding of global thermal anisotropy intensity (TAI) and its governing factors remains missing. We introduce a novel data-driven methodology to quantify global TAI exclusively using multi-angle MODIS land surface temperature time series observations. Our analysis reveals distinct seasonal and diurnal TAI patterns, with global mean summertime TAI exceeding 2.9°C. Furthermore, we identify strong associations between TAI and key surface and atmospheric parameters, such as leaf area index and downward shortwave radiation. Our findings advocate for a paradigm shift from model-based to data-driven approaches in correcting thermal anisotropy, thereby addressing a critical bottleneck in Earth observation.

54 ENVIRONMENTAL SCIENCES

Simple Parametric Model for Intensity Calibration of Cassini Composite Infrared Spectrometer Data

Accurate intensity calibration of a linear Fourier-transform spectrometer typically requires the unknown science target and the two calibration targets to be acquired under identical conditions. We present a simple model suitable for vector calibration that enables accurate calibration via adjustments of measured spectral amplitudes and phases when these three targets are recorded at different detector or optics temperatures. Our model makes calibration more accurate both by minimizing biases due to changing instrument temperatures that are always present at some level and by decreasing estimate variance through incorporating larger averages of science and calibration interferogram scans.

Brasunas, J.

Comparison of TRMM Ground Validation and Satellite Rain Intensity Estimates

The Tropical Rainfall Measuring Mission (TRMM) Ground Validation (GV) Program began in the late 1980's and has provided a wealth of data and resources for validating TRMM satellite estimates. The TRMM GV program's main operational task is to provide rainfall products for four sites: Darwin, Australia (DARW); Houston, Texas (HSTN); Kwajalein, Republic of the Marshall Islands (KWAJ); and, Melbourne, Florida (MELB). A comparison between TRMM Ground Validation (Version 5) and Satellite (Version 6) rain intensity estimates is presented. The full suite of Version 6 satellite data is currently being generated by the TRMM Science Data and Information System (TSDIS) and should be completed some time near the end of 2005. The gridded satellite product (3G68) will be compared to GV Level II rain-intensity and -type maps (2A53 and 2A54, respectively). The 3G68 product represents a 0.5 deg x 0.5 deg data grid providing estimates of rain intensities from the TRMM Precipitation Radar (PR), Microwave Imager (TMI) and Combined (COM) algorithms. The comparisons will be sub-setted according to geographical type (land, coast and ocean). A bias statistic will be presented that provides quantification of the relative differences between the various estimators. Previous comparisons of an interim satellite product (Version 6a) showed that all of the estimates (GV and satellite) are converging, with some expected discrepancies. The convergence of the GV and satellite estimates bodes well for expectations for the proposed Global Precipitation Measurement (GPM) program and this study and others are being leveraged towards planning GV goals for GPM.

Wolff, David B.

A Pragmatic Path to Investigating Europa's Habitability

Assessment of Europa's habitability, as an overarching science goal, will progress via a comprehensive investigation of Europa's subsurface ocean, chemical composition, and internal dynamical processes, The National Research Council's Planetary Decadal Survey placed an extremely high priority on Europa science but noted that the budget profile for the Jupiter Europa Orbiter (1EO) mission concept is incompatible with NASA's projected planetary science budget Thus, NASA enlisted a small Europa Science Definition Team (ESDT) to consider more pragmatic Europa mission options, In its preliminary findings (May, 2011), the ESDT embraces a science scope and instrument complement comparable to the science "floor" for JEO, but with a radically different mission implementation. The ESDT is studying a two-element mission architecture, in which two relatively low-cost spacecraft would fulfill the Europa science objectives, An envisioned Europa orbital element would carry only a very small geophysics payload, addressing those investigations that are best carried out from Europa orbit An envisioned separate multiple Europa flyby element (in orbit about Jupiter) would emphasize remote sensing, This mission architecture would provide for a subset of radiation-shielded instruments (all relatively low mass, power, and data rate) to be delivered into Europa orbit by a modest spacecraft, saving on propellant and other spacecraft resources, More resource-intensive remote sensing instruments would achieve their science objectives through a conservative multiple-flyby approach, that is better situated to handle larger masses and higher data volumes, and which aims to limit radiation exposure, Separation of the payload into two spacecraft elements, phased in time, would permit costs to be spread more uniformly over mUltiple years, avoiding an excessively high peak in the funding profile, Implementation of each spacecraft would be greatly simplified compared to previous Europa mission concepts, minimizing new development while achieving the key Europa science objectives. We will report on the status of this evolving concept, and will solicit community feedback, as we pursue an innovative and low-cost ways to explore Europa and investigate its habitability.

Pappalardo

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science