Data-Intensive Science Meets Inquiry-Driven Pedagogy: Interactive Big Data Exploration, Threshold Concepts, and Liminality
No abstract available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
No abstract available
Threshold concepts in any discipline are the core concepts an individual must understand in order to master a discipline. By their very nature, these concepts are troublesome, irreversible, integrative, bounded, discursive, and reconstitutive. Although grasping threshold concepts can be extremely challenging for each learner as s/he moves through stages of cognitive development relative to a given discipline, the learner's grasp of these concepts determines the extent to which s/he is prepared to work competently and creatively within the field itself. The movement of individuals from a state of ignorance of these core concepts to one of mastery occurs not along a linear path but in iterative cycles of knowledge creation and adjustment in liminal spaces - conceptual spaces through which learners move from the vaguest awareness of concepts to mastery, accompanied by understanding of their relevance, connectivity, and usefulness relative to questions and constructs in a given discipline. For example, challenges in the teaching and learning of atmospheric science can be traced to threshold concepts in fluid dynamics. In particular, Dynamic Meteorology is one of the most challenging courses for graduate students and undergraduates majoring in Atmospheric Science. Dynamic Meteorology introduces threshold concepts - those that prove troublesome for the majority of students but that are essential, associated with fundamental relationships between forces and motion in the atmosphere and requiring the application of basic classical statics, dynamics, and thermodynamic principles to the three dimensionally varying atmospheric structure. With the explosive growth of data available in atmospheric science, driven largely by satellite Earth observations and high-resolution numerical simulations, paradigms such as that of dataintensive science have emerged. These paradigm shifts are based on the growing realization that current infrastructure, tools and processes will not allow us to analyze and fully utilize the complex and voluminous data that is being gathered. In this emerging paradigm, the scientific discovery process is driven by knowledge extracted from large volumes of data. In this presentation, we contend that this paradigm naturally lends to inquiry-driven pedagogy where knowledge is discovered through inductive engagement with large volumes of data rather than reached through traditional, deductive, hypothesis-driven analyses. In particular, data-intensive techniques married with an inductive methodology allow for exploration on a scale that is not possible in the traditional classroom with its typical problem sets and static, limited data samples. In addition, we identify existing gaps and possible solutions for addressing the infrastructure and tools as well as a pedagogical framework through which to implement this inductive approach.
This call to action by Drs. Johnson and Wilkinson is part of a mosaic of voices sharing tangible progress within the climate movement 1,2. This call speaks to us as environmental and Earth scientists motivated by the urgency of climate change and social inequity and who contribute to finding science-driven climate solutions as part of our daily jobs. Unfortunately, we are often unable to efficiently move this critical and urgent work forward because we are impeded by cumbersome daily workflows and restrictive workplace cultures. Our workplaces have not kept pace with the modern realities of data-intensive science: increasing data volumes and storage needs, rapidly evolving technology, new skill requirements, and a growing need for extensive and diverse collaboration. Struggling with old approaches and learning new ones in isolation can fuel burnout and turnover, preventing us from working on science-driven climate solutions effectively.
We describe the problem of regular small telemetry losses incurred during coherency mode transitions in Cassini's telecommunication. The project did not originally plan any corrective steps for avoiding these data losses, because of 1) the disparity between the small durations of the transitions (1-2 min) and large playback capability losses (15 min) needed for bracketing the transition time spans and 2) the unpredictable content of data downlinking during the transitions. However, as the intense science data return from the tour began, it became apparent that the impact of these small losses can sometimes be significant. We provide two examples of the impact on Radar-dedicated Titan flybys.
We describe the problem of regular small telemetry losses incurred during coherency mode transitions in Cassini’s telecommunication. The project did not originally plan any corrective steps for avoiding these data losses, because of 1) the disparity between the small durations of the transitions (1-2 min) and large playback capability losses (15 min) needed for bracketing the transition time spans and 2) the unpredictable content of data downlinking during the transitions. However, as the intense science data return from the tour began, it became apparent that the impact of these small losses can sometimes be significant. We provide two examples of the impact on Radar-dedicated Titan flybys. In general, the impacts are larger for high-rate data and for data acquired during a targeted flyby of Titan and other icy satellites. Although the content of data during a transition for every downlink pass is unpredictable, we are certain that some important data will be lost on downlink passes dedicated to transmit the flyby data and it does not matter what part of the data will be hit by the transitions. We collected more than 200 days of data from Cassini tour operations between June 2004 and February 2005 to analyze the distributions of the start time and duration of the transitions. We found that the occurrence of a transition can be predicted within a 5-min window, with 95 percent confidence. Given that, it is possible to eliminate the data losses by pausing playback at the beginning of a transition for 5 minutes and resuming playback after transition completion. We briefly describe three operational fixes as to how to implement the playback pause, with the pros and cons for each method. Finally, we report the results of the method chosen by the project and implemented on the spacecraft for several Titan and icy satellites flybys between September and October, 2005.
NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs
High capacity communications from Martian distances, required for the envisioned human exploration and desirable for data-intensive science missions, is challenging. NASA s Deep Space Network currently requires large antennas to close RF telemetry links operating at kilobit-per-second data rates. To accommodate higher rate communications, NASA is considering means to achieve greater effective aperture at its ground stations. This report, focusing on the return link from Mars to Earth, demonstrates that without excessive research and development expenditure, operational Mars-to-Earth RF communications systems can achieve data rates up to 1 Gbps by 2020 using technology that today is at technology readiness level (TRL) 4-5. Advanced technology to achieve the needed increase in spacecraft power and transmit aperture is feasible at an only moderate increase in spacecraft mass and technology risk. In addition, both power-efficient, near-capacity coding and modulation and greater aperture from the DSN array will be required. In accord with these results and conclusions, investment in the following technologies is recommended:(1) lightweight (1 kg/sq m density) spacecraft antenna systems; (2) a Ka-band receive ground array consisting of relatively small (10-15 m) antennas; (3) coding and modulation technology that reduces spacecraft power by at least 3 dB; and (4) efficient generation of kilowatt-level spacecraft RF power.
The Global Learning and Observations to Benefit the Environment (GLOBE) citizen science program has recently conducted a series of month-long intensive observation periods (IOPs), asking the public to submit daily reports on cloud and sky conditions from all regions of Earth. This provides a wealth of crowdsourced observations from the ground, which complements other conventional scientific cloud data. In addition, the GLOBE reports are matched in space and time with geostationary and low Earth orbit satellites, which allows for a straightforward comparison of cloud properties, and minimizes the biases associated with mismatched sampling between participants and satellites. The matched GLOBE dataset is used to calculate the mean observed cloud cover by atmospheric level both worldwide and by region. The overall magnitudes of cloud cover between the GLOBE participants and the matched satellites agree within 10%, which is notable given the distinctly different natures of the data sources. The mean vertical cloud profiles show GLOBE reporting more low-level clouds and fewer high-level clouds than satellites. The low cloud disagreement is likely related to satellites missing low clouds when high clouds block their view. Conversely, the high cloud disagreement is related primarily to cloud opacity, as satellites may miss some optically thin clouds. Monte Carlo testing shows the results to be robust, and the tripled amount of IOP data reduces uncertainty by half. These findings also highlight ways in which citizen science IOP data may be used to support scientific research while accounting for their unique properties. Plain Language Summary: Citizen science is becoming an increasingly prominent aspect of scientific research, and so it important to study how citizen science data can be used effectively. For example, The GLOBE Program has recently conducted a series of special data-collecting events, or “challenges”, which gathered large numbers of reports on cloud and sky conditions. Because NASA GLOBE Clouds matches the participant reports with cloud observations from satellites, we can use these data to get a combined view of clouds from above and below. When looking at the average cloud cover for different atmospheric levels across Earth, we find that the GLOBE participants and the satellites agree quite closely. This is a surprising and fascinating find, given how different in nature volunteer ground reports are to satellite measurements. However, there are some small but notable disagreements between GLOBE participants and satellites about the distribution of cloud cover at different levels. In addition, by testing the data for uncertainty, we show that the results from the GLOBE data are reliable, and that more public participation improves the reliability. So, by carefully designing the analysis methodology, and by testing for the uncertainty of the data, citizen science can make a meaningful contribution to scientific research.
Explore the source record for details and available documents.
The NASA Center for Climate Simulation (NCCS) offers integrated supercomputing, visualization, and data interaction technologies to enhance NASA's weather and climate prediction capabilities. It serves hundreds of users at NASA Goddard Space Flight Center, as well as other NASA centers, laboratories, and universities across the US. Over the past year, NCCS has continued expanding its data-centric computing environment to meet the increasingly data-intensive challenges of climate science. We doubled our Discover supercomputer's peak performance to more than 800 teraflops by adding 7,680 Intel Xeon Sandy Bridge processor-cores and most recently 240 Intel Xeon Phi Many Integrated Core (MIG) co-processors. A supercomputing-class analysis system named Dali gives users rapid access to their data on Discover and high-performance software including the Ultra-scale Visualization Climate Data Analysis Tools (UV-CDAT), with interfaces from user desktops and a 17- by 6-foot visualization wall. NCCS also is exploring highly efficient climate data services and management with a new MapReduce/Hadoop cluster while augmenting its data distribution to the science community. Using NCCS resources, NASA completed its modeling contributions to the Intergovernmental Panel on Climate Change (IPCG) Fifth Assessment Report this summer as part of the ongoing Coupled Modellntercomparison Project Phase 5 (CMIP5). Ensembles of simulations run on Discover reached back to the year 1000 to test model accuracy and projected climate change through the year 2300 based on four different scenarios of greenhouse gases, aerosols, and land use. The data resulting from several thousand IPCC/CMIP5 simulations, as well as a variety of other simulation, reanalysis, and observationdatasets, are available to scientists and decision makers through an enhanced NCCS Earth System Grid Federation Gateway. Worldwide downloads have totaled over 110 terabytes of data.
Modern science is increasingly dependent on computational analysis of very large data sets. Organizing, referencing, publishing those data has become a complex problem. Published research that depends on such data often fails to cite the data in sufficient detail to allow an independent scientist to reproduce the original experiments and analyses. This paper explores some of the challenges related to data identification, equivalence and reproducibility in the domain of data intensive scientific processing. It will use the example of Earth Science satellite data, but the challenges also apply to other domains.
The background for the Viking Lander Monitor Mission (VLMM) is given, and the technical and operational aspects of the tracking and data acquisition support that the Network was called upon to provide are described. An overview of the science results obtained from the imaging, meteorological, and radio science data is also given. The intensive efforts that were made to recover the mission are described.
This paper presents our experience with OODT, a novel software architectual style, and middlware-based implementation for data-intensive systems. To date, OODT has been successfully evaluated in several different science domains including Cancer Research with the National Cancer Institute (NCI), and Planetary Science with NASA's Planetary Data System (PDS).
A large portion of Earth Science investigations is phenomenon- or event-based, such as the studies of Rossby waves, mesoscale convective systems, and tropical cyclones. However, except for a few high-impact phenomena, e.g. tropical cyclones, comprehensive records are absent for the occurrences or events of these phenomena. Phenomenon-based studies therefore often focus on a few prominent cases while the lesser ones are overlooked. Without an automated means to gather the events, comprehensive investigation of a phenomenon is at least time-consuming if not impossible. An Earth Science event (ES event) is defined here as an episode of an Earth Science phenomenon. A cumulus cloud, a thunderstorm shower, a rogue wave, a tornado, an earthquake, a tsunami, a hurricane, or an EI Nino, is each an episode of a named ES phenomenon," and, from the small and insignificant to the large and potent, all are examples of ES events. An ES event has a finite duration and an associated geolocation as a function of time; its therefore an entity in four-dimensional . (4D) spatiotemporal space. The interests of Earth scientists typically rivet on Earth Science phenomena with potential to cause massive economic disruption or loss of life, but broader scientific curiosity also drives the study of phenomena that pose no immediate danger. We generally gain understanding of a given phenomenon by observing and studying individual events - usually beginning by identifying the occurrences of these events. Once representative events are identified or found, we must locate associated observed or simulated data prior to commencing analysis and concerted studies of the phenomenon. Knowledge concerning the phenomenon can accumulate only after analysis has started. However, except for a few high-impact phenomena. such as tropical cyclones and tornadoes, finding events and locating associated data currently may take a prohibitive amount of time and effort on the part of an individual investigator. And even for these high-impact phenomena, the availability of comprehensive records is still only a recent development. A major reason for the lack of comprehensive ,records for the majority of the ES phenomena is the perception that they do not pose immediate and/or severe threat to life and property and are thus not consistently tracked. monitored, and catalogued. Many phenomena even lack commonly accepted criteria for definitions. However. the lack of comprehensive records is also due to the increasingly prohibitive volume of observations and model data that must be examined. NASA Earth Observing System Data Information System (EOSDIS) alone archives several petabytes (PB) of satellite remote sensing data and steadily increases. All of these factors contribute to the difficulty of methodically identifying events corresponding to a given phenomenon and significantly impede systematic investigations. In the following we present a couple motivating scenarios, demonstrating the issues faced by Earth scientists studying ES phenomena.
Over the previous four years the Earth Radiation Budget Experiment (ERBE) instruments have been gathering data on two satellites, the Earth Radiation Budget Satellite and the the operational NOAA-9 satellite. The ERBE science team recently completed the validation of an initial sampling of these data involving intensive examination of data in four months during 1985 and 1986. The data being placed in the National Space Science Data Center to acquaint the scientific community with their availability are discussed. The ERBE archival data products are also presented.
An approach is defined that describes a method of iterating over massively large arrays containing sparse data using an approach that is implementation independent of how the contents of the sparse arrays are laid out in memory. What is unique and important here is the decoupling of the iteration over the sparse set of array elements from how they are internally represented in memory. This enables this approach to be backward compatible with existing schemes for representing sparse arrays as well as new approaches. What is novel here is a new approach for efficiently iterating over sparse arrays that is independent of the underlying memory layout representation of the array. A functional interface is defined for implementing sparse arrays in any modern programming language with a particular focus for the Chapel programming language. Examples are provided that show the translation of a loop that computes a matrix vector product into this representation for both the distributed and not-distributed cases. This work is directly applicable to NASA and its High Productivity Computing Systems (HPCS) program that JPL and our current program are engaged in. The goal of this program is to create powerful, scalable, and economically viable high-powered computer systems suitable for use in national security and industry by 2010. This is important to NASA for its computationally intensive requirements for analyzing and understanding the volumes of science data from our returned missions.
The growth and spread of human settlements and increasing urban density are important processes in global change. Urbanization has accelerated with population growth, and more densely populated urban areas have major effects on the local and regional environments. Air quality in megacities has been a concern for decades. Data for anthropogenic emissions, pollutants, and particulate matter can inform research on the interactions between urban landscapes and the atmosphere, assist policy makers in developing sustainable, healthy environments, and inform the general public. To better assist researchers, students, the public, and policymakers in understanding the annual and seasonal variation of aerosol and gas intensity in megacities, the Science Outreach Team at NASA Langley Research Center’s Atmospheric Science Data Center (ASDC) Distributed Active Archive Center (DAAC) demonstrate data products at the ASDC that can be used to visualize these parameters. The presentation uses data from the ASDC-supported NASA missions Measurements Of Pollution In The Troposphere (MOPITT), Cloud-Aerosol and Infrared Pathfinder Satellite Observation (CALIPSO), Tropospheric Emission Spectrometer (TES), and Multi-angle Imaging Spectro Radiometer (MISR)
The Florida scrub-jay (FSJ) is a species in decline because of extinction debt caused by habitat fragmentation and degradation. Understanding and managing the species and its habitats is challenging due to the complex interactions among the social system, a dynamic mosaic of scrub habitat, and active management. To provide insights into the possible fates of the FSJ populations of mainland, cape, and island sections of Brevard County, we modeled the population dynamics using the Vortex population viability analysis (PVA) software. Vortex is an individual-based simulation that allowed us to include such factors as demographic rates dependent on habitat state, impact of helpers on breeding success, and helper to breeder transition probabilities responding to availability of nearby vacant optimal habitat and vacancies due to the death of breeders. Detailed modeling of the FSJ was possible only because a lot of data about the species and its habitats have been gathered over decades of intensive research. We followed a phased approach to constructing population models that incorporated the best available science and data to address a variety of conservation actions. The first phase focused on constructing a model that incorporates sociobiology and source-sink habitat dynamics. This model allowed us to address some of the most important questions about population size and habitat quality. Once the model framework was in place, we considered the real landscapes and actual local populations, rather than just generic representations of typical FSJ dynamics. After examining the viability of the metapopulations under current conditions, we explored the likely consequences of various management actions that might slow or reverse population declines.