Search NASASearch

Engineering topics

Christopher Lynnes

Publications and source records attributed to Christopher Lynnes.

Introduction to Analysis Methods for Big Earth Data

Big Earth Data are too big to be tractable to simple data inspection. Thus, they typically require models to make sense of all the data. Useful models for Big Earth Data may be physical, statistical, or machine learning based. While physical models are ideal for understanding the data, they are not always feasible, particularly when our ability to observe at finer scales exceeds our ability to incorporate the physics. Statistical models are more generalized, but computationally intensive for many Earth Observation datasets. Machine Learning models generally scale well but are sometimes limited in the physical understanding they can offer. Hybrid models combine attributes—and advantages—of two or more of these types.

Christopher Lynnes

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. At the same time in the Earth Observation area, technology advances are enabling new sensors and satellites that will increase data volume, velocity and application variety. Scaling up can also be seen when operational applications expand from small, local studies to larger spatial scales with more analysis targets.

Christopher Lynnes

Reusability in NASA's Earth Observation System Data and Information System (EOSDIS)

NASA's Earth Observation System Data and Information System (EOSDIS) has been operational since 1994. The term FAIR is relatively recent compared to the operational life of EOSDIS. However, given the evolutionary nature of EOSDIS in response to both the state of the art and community expectations, it is useful to assess where EOSDIS stands relative to the FAIR data management. In this presentation we evaluate EOSDIS against the Reusability principle, which calls for a clear and accessible data usage license, rich description of data, detailed provenance, and compliance with domain-relevant community standards.Achieving consistent compliance for datasets is particularly challenging with the diverse community of missions and scientists providing data to EOSDIS. To address this, EOSDIS's constituent Distributed Active Archive Centers work with data producers to utilize the applicable standards and conventions. In addition, a guidebook is under development to provide comprehensive and understandable guidance on how to construct compliant, and thus more reusable, data products.

Hampapuram (Rama) Ramapriyan

Cloud Optimized Data Formats

Cloud computing offers the promise of being able to analyze Big Data earth Observations at scale, by allowing scientists to deploy many nodes at once to analyze the data. However, in order to take full advantage of cloud scalability, it is often necessary to reorganize and reformat the data to enable fine-grained, parallel access to the data in Web Object Storage. NASA recently conducted a study of several formats that are optimized for analysis in the cloud: Parquet, zarr, HDF (Hierarchical Data Format) in the Cloud, and Cloud-Optimized GeoTIFF (Tagged Image File Format). They were compared against non-cloud-optimized formats, netCDF (network Common Data Form) and GeoTIFF, with criteria based both on stewardship and analysis performance.

Christopher Lynnes

Usage-Based Discovery of Earth Observations

Most providers of Earth Observation data enable search via dataset characteristics, (e.g., quantity being measured, instrument, location, and time). However, the Earth Science Information Partners Federation (ESIP) is attempting to improve dataset discovery by capturing information on how datasets are used. At a usage-based discovery hackfest at the July 2020 ESIP Summer Meeting, participants pooled their collective skills to implement a prototype based on the connections between datasets and data usage, which can reveal unanticipated patterns in dataset connectivity that are not apparent through traditional approaches. A survey of natural hazards websites, focusing on hydrology-related hazards such as floods and hydrology datasets such as rainfall, illustrated the difficulty of gleaning which dataset was actually used and how it was accessed. This difficulty highlights the need for greater transparency on dataset usage in website design. An application wireframe for a usage-based discovery database was conceptualized as a tool to help applications data specialists rapidly assemble a fit-for-purpose website as part of a disaster response, where customer feedback was integral to the manner in which the website would be developed and delivered. Capturing connections among datasets and instances of usage, either for research or applications, could help users identify datasets with a proven track record of utility or those that have been supplanted by more accurate and(or) precise representations. It may also help Earth Observation providers maximize returns on investments in data services. Social discovery, i.e, following discovery paths laid down by previous users of the data, could be a critical tool for improving the timeliness and applicability of dataset delivery at times and in circumstances where practitioners don’t have the capacity to evaluate the deluge of datasets that are returned by today’s search tools.

search

NASA Cloud Data Access

Explore the source record for details and available documents.

Cloud Computing

NASA Briefing to Unidata

The NASA report to Unidata for October, 2021 covers several topics of mutual interest to Unidata. These include plans for hosting NASA data in the cloud, data analysis-in-place, and NASA plans for Open Source Science. Major collaborations with other organizations are also discussed.

NetCDF

Fostering Open Science Inclusiveness for Interdisciplinary Users of Earth Observations

The term Open Science is subject to a variety of interpretations because of a key (and useful) ambiguity in the meaning of “Open”. Open in the sense of Transparency enables more trust in science research by making the details of the scientific process visible and accessible to anyone. “Open” in the sense of Inclusiveness enables more scientists from other disciplines to participate in research in a given discipline, thus producing more interdisciplinary research. Data Systems can play a major role in enabling Open (Inclusive) Science by making it easier for users from other disciplines to work with data within a given discipline. This is challenging for Earth Observation datasets, most of which are the product of advanced instrumentation and sophisticated, specialized variable retrieval algorithms and code. Serving the “extra-disciplinary”communities begins with simple things, like accessible, readable data documentation with adequate scaffolding. But just as important is provisioning Analysis-Ready data that does not require expert pre-processing. Disciplines also often have dominant toolsets, such as R in the biomass community or GIS in many applications communities. Ensuring that EO data are easy to use in the tools favored in other communities will enable more interdisciplinary research. Ideally, interdisciplinary research also benefits from scientists with different domain expertise. Platforms and frameworks that facilitate frictionless collaboration with discipline experts, together with capacity building efforts in those external disciplines also improve the inclusiveness aspect of Open Science. In short, Open Science is at root a way of thinking about how users from diverse discipline can best access and use data and services from a particular discipline.

Christopher Lynnes

Populating a Graph Database to Run a Usage-Based Discovery Tool

Most dataset discovery tools for Earth Observation data rely on descriptions and other metadata of the datasets, using keyword searches or attribute filtering to determine relevance. However, these descriptions often do not include the potential uses of the data. Thus, a user working on floods will rarely see few if any rainfall datasets show up in such a search. The Usage Based Discovery tool, on the other hand, offers usage instances to the user, either research articles or applications, along with the datasets that those usage instances used. This allows a user, particularly one new to the world of Earth Observation data, to investigate which datasets are used in similar cases. The information that powers Usage-Based Discovery is a graph database of relationships of usage to dataset and usage to topic, allowing the user to narrow their search for similar cases. In order to scale out to a graph database rich enough to provide a satisfactory user experience, we combine manual and automated processes to populate the graph. The initial content of the graph has been seeded primarily via human-aided data curation methods, using sites like Google Scholar. To scale up this effort, we’ve employed crowdsourcing. It is easy for anyone to contribute to our graph using their Open Researcher and Contributor Identifier for authorization. We’re now experimenting with Machine Learning and Natural Language Processing to help automate population of the graph, starting with the classification of research articles by topic. Finding adequate training data in the absence of a comprehensive and open research article API continues to be a significant challenge.

Vincent Inverso

A Framework for Assessing Earth Observation Metadata Quality: Implications for Data Discovery and Open Science

The Common Metadata Repository (CMR) contains metadata records describing NASA’s collection of over 8,000 Earth observation data products. The Analysis and Review of CMR (ARC) Team at Marshall Space Flight Center assesses the quality of these metadata records. Metadata, rather than the data itself, is indexed for search in both discipline-specific datacenters and global or aggregated catalogs (such as Earth data Search), making it essential for determining whether a data product is appropriate for a given research question or application need. Since metadata connects users to data, it should be as accurate and complete as possible in addition to meeting minimum database requirements. The ARC team has developed a metadata quality framework by which to assess quality. The framework consists of a set of quality criteria that converge around the dimensions of correctness, completeness, and consistency, with the goal of improving the discoverability, accessibility, and usability of NASA’s Earth Observation data. The application of the framework has resulted in a measurable improvement in NASA’s metadata quality. Key aspects of the framework’s success are the ability to systematically evaluate metadata and provide actionable quality improvement recommendations. Lessons learned from the project will be shared along with implementation details which may be relevant to other science disciplines. By aiming to make data more discoverable and accessible to a broad user community, the ARC metadata quality framework helps contribute to NASA’s commitment to open science.

Jeanne Le Roux

Exploring the Landscape of Earth and Space Science Informatics using Latent Topic Modeling

AGU Earth and Space Science Informatics (ESSI) is at the forefront of data management, analysis, large scale experimentation, and infrastructure development pertaining to Earth and Space Science interests. The key topics of interest within ESSI are also evolving and diversifying over time. We aim to observe and quantify the various topics covered in ESSI, analyze their trends over time, and identify the contributors’ affiliations to gain an understanding of the landscape of ESSI and the direction of the research and management. The data for this work are abstracts submitted to AGU’s ESSI Fall meeting; They serve as a proxy for key research and development areas within ESSI. We use an unsupervised topic modeling technique called Latent Dirichlet Allocation to observe the underlying topics covered in ESSI and their trends over time. With this presentation, we showcase our results from the analysis and insights gained.

Muthukumaran Ramasubramanian