Search NASA⌕ Search

SEARCH · Search NASA

Results for “data access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Integrating Multi-agency Data Products in a Cloud-based Platform for Streamlined Discovery, Visualization, and Use

Earth science data users almost always have an interest in utilizing geospatial data from multiple agencies. As computing capability and cloud-based infrastructures accelerate the pace at which scientific research can be done, there is a growing need to enable search, discovery, and use of multi-agency geospatial observations relevant for a common use case - without undergoing the search and discovery process in a less efficient, disparate path with each agency. NASA’s Earth Observing System Data and Information System (EOSDIS) and NOAA’s National Environmental Satellite, Data and Information Service (NESDIS) both support a wide range of Earth science disciplines’ research, operations, and applications activities. Presently, however, there are few examples of data discovery frameworks supporting an inquiry of both NASA’s and NOAA’s extensive archives of Earth observations that are equally suitable for a particular science scenario, regardless of the agency that “owns” the data. NASA and NOAA are collaborating on a data expedition platform for exploring fire weather using data products from both agencies. Users will be able to search, discover, and visualize NASA and NOAA products in one interface. Each agency will curate metadata for its respective datasets, providing for a rich search experience. The collaboration will pilot a shared search interface into these metadata datastores. Data products will be stored in the cloud in cloud-optimized format(s). These formats will allow for optimized data access and visualization to support the “data expedition”. Avenues for further development and application of this cloud-based, multi-agency data provisioning platform will also be discussed.

cloud-based technology↗

ASDC Distribution and Services of TEMPO Data

The Tropospheric Emissions: Monitoring of POllution (TEMPO) instrument represents a groundbreaking advancement in remote sensing technology, providing real-time and high-resolution measurements of atmospheric pollutants and air quality monitoring. This presentation will highlight the significance of TEMPO and its data distribution by NASA’s Atmospheric Science Data Center (ASDC) to facilitate the analyses by the research and end user communities. TEMPO's geostationary orbit allows for continuous and high-resolution measurements of key atmospheric pollutants, including nitrogen dioxide (NO 2 ), ozone (O 3 ), and formaldehyde (HCHO). As a result, researchers can investigate the distribution patterns of these pollutants on an hourly basis, providing valuable information for understanding regional and temporal variations in air quality. Furthermore, this presentation will highlight the usability of TEMPO data in complementing ground-based air quality monitoring networks. ASDC distribution services and tools facilitate efficient data handling and analysis to support research and application uses of TEMPO data. As part of NASA’s Earthdata ecosystem, TEMPO will be available through Earthdata Search and Worldview, as well as have variable and spatial subsetting capabilities. This presentation will provide an overview of data access and services available for TEMPO data.

Hazem Mahmoud↗

An Update on Global Satellite-Based Precipitation Products and Services at NASA GES DISC

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is home to major NASA satellite precipitation measurement missions including the Tropical Rainfall Measuring Mission (TRMM) and the Global Precipitation Measurement (GPM) as well as other NASA projects. Accurate and timely available global and regional precipitation products play an important role in research and applications around the world. The GES DISC provides near-real-time, near-global precipitation products (e.g.IMERG) to support a wide variety of interdisciplinary research and operational activities including flood modeling, landslides, crop monitoring and assessment, vector-borne diseases, etc. Climate data record products (e.g.GPCP3) are essential for climate research, assessment, model evaluation and applications. To facilitate data access and exploration, the GES DISC has developed data services such asGiovanni, an online visualization and analysis tool for easy access to over 2000 satellite- and model-based variables. Established in the mid 1980s, the GES DISC also distributes data in other disciplines including hydrology, atmospheric chemistry, atmospheric dynamics, etc. In this presentation, we will present an update on global precipitation products and data services including the new IMERG V06B suite, the latest version of GPCP (Version 3) and value-added data subsetting services (L34RS, L2S).

Liu, Zhong↗

Geophysical data analysis and visualization using the Grid Analysis and Display System

Several problems posed by the rapidly growing volume of geophysical data are described, and a selected set of existing solutions to these problems is outlined. A recently developed desktop software tool called the Grid Analysis and Display System (GrADS) is presented. The GrADS' user interface is a natural extension of the standard procedures scientists apply to their geophysical data analysis problems. The basic GrADS operations have defaults that naturally map to data analysis actions, and there is a programmable interface for customizing data access and manipulation. The fundamental concept of the GrADS' dimension environment, which defines both the space in which the geophysical data reside and the 'slice' of data which is being analyzed at a given time, is expressed The GrADS' data storage and access model is described. An argument is made in favor of describable data formats rather than standard data formats. The manner in which GrADS users may perform operations on their data and display the results is also described. It is argued that two-dimensional graphics provides a powerful quantitative data analysis tool whose value is underestimated in the current development environment which emphasizes three dimensional structure modeling.

Doty, Brian E.↗

Zero-Copy Objects System

Zero-Copy Objects System software enables application data to be encapsulated in layers of communication protocol without being copied. Indirect referencing enables application source data, either in memory or in a file, to be encapsulated in place within an unlimited number of protocol headers and/or trailers. Zero-copy objects (ZCOs) are abstract data access representations designed to minimize I/O (input/output) in the encapsulation of application source data within one or more layers of communication protocol structure. They are constructed within the heap space of a Simple Data Recorder (SDR) data store to which all participating layers of the stack must have access. Each ZCO contains general information enabling access to the core source data object (an item of application data), together with (a) a linked list of zero or more specific extents that reference portions of this source data object, and (b) linked lists of protocol header and trailer capsules. The concatenation of the headers (in ascending stack sequence), the source data object extents, and the trailers (in descending stack sequence) constitute the transmitted data object constructed from the ZCO. This scheme enables a source data object to be encapsulated in a succession of protocol layers without ever having to be copied from a buffer at one layer of the protocol stack to an encapsulating buffer at a lower layer of the stack. For large source data objects, the savings in copy time and reduction in memory consumption may be considerable.

Burleigh, Scott C.↗

The Challenges of Searching, Finding, Reading, Understanding and Using Mars Mission Datasets for Science Analysis

This viewgraph presentation reviews the problems that non-mission researchers have in accessing data to use in their analysis of Mars. The increasing complexity of Mars datasets results in custom software development by instrument teams that is often the only means to visualize and analyze the data. The solutions to the problem are to continue efforts toward synergizing data from multiple missions and making the data, s/w, derived products available in standardized, easily-accessible formats, encourage release of "lite" versions of mission-related software prior to end-of-mission, and planetary image data should be systematically processed in a coordinated way and made available in an easily accessed form. The recommendations of Mars Environmental GIS Workshop are reviewed.

data sets↗

Binding profiles for 961 Drosophila and C. elegans transcription factors reveal tissue-specific regulatory relationships

A catalog of transcription factor (TF) binding sites in the genome is critical for deciphering regulatory relationships. Here, we present the culmination of the efforts of the modENCODE (model organism Encyclopedia of DNA Elements) and modERN (model organism Encyclopedia of Regulatory Networks) consortia to systematically assay TF binding events in vivo in two major model organisms,Drosophila melanogaster(fly) andCaenorhabditis elegans(worm). These data sets comprise 605 TFs identifying 3.6 M sites in the fly and 356 TFs identifying 0.9 M sites in the worm, and represent the majority of the regulatory space in each genome. We demonstrate that TFs associate with chromatin in clusters termed “metapeaks,” that larger metapeaks have characteristics of high-occupancy target (HOT) regions, and that the importance of consensus sequence motifs bound by TFs depends on metapeak size and complexity. Combining ChIP-seq data with single-cell RNA-seq data in a machine-learning model identifies TFs with a prominent role in promoting target gene expression in specific cell types, even differentiating between parent–daughter cells during embryogenesis. These data are a rich resource for the community that should fuel and guide future investigations into TF function. To facilitate data accessibility and utility, all strains expressing green fluorescent protein (GFP)-tagged TFs are available at the stock centers for each organism. The chromatin immunoprecipitation sequencing data are available through the ENCODE Data Coordinating Center, GEO, and through a direct interface that provides rapid access to processed data sets and summary analyses, as well as widgets to probe the cell-type-specific TF–target relationships.

Biochemistry & Molecular Biology↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗

Rdesign: A data dictionary with relational database design capabilities in Ada

Data Dictionary is defined to be the set of all data attributes, which describe data objects in terms of their intrinsic attributes, such as name, type, size, format and definition. It is recognized as the data base for the Information Resource Management, to facilitate understanding and communication about the relationship between systems applications and systems data usage and to help assist in achieving data independence by permitting systems applications to access data knowledge of the location or storage characteristics of the data in the system. A research and development effort to use Ada has produced a data dictionary with data base design capabilities. This project supports data specification and analysis and offers a choice of the relational, network, and hierarchical model for logical data based design. It provides a highly integrated set of analysis and design transformation tools which range from templates for data element definition, spreadsheet for defining functional dependencies, normalization, to logical design generator.

Lekkos, Anthony A.↗

Analysis of Water and Energy Budgets and Trends Using the NLDAS Monthly Data Sets

The North American Land Data Assimilation System (NLDAS) is a collaborative project between NASA GSFC, NOAA, Princeton University, and the University of Washington. NLDAS has created surface meteorological forcing data sets using the best-available observations and reanalyses. The forcing data sets are used to drive four separate land-surface models (LSMs), Mosaic, Noah, VIC, and SAC, to produce data sets of soil moisture, snow, runoff, and surface fluxes. NLDAS hourly data, accessible from the NASA GES DISC Hydrology Data Holdings Portal, http://disc.sci.gsfc.nasa.gov/hydrology/data-holdings, are widely used by various user communities in modeling, research, and applications, such as drought and flood monitoring, watershed and water quality management, and case studies of extreme events. More information is available at http://ldas.gsfc.nasa.gov/. To further facilitate analysis of water and energy budgets and trends, NLDAS monthly data sets have been recently released by NASA GES DISC.

Vollmer, Bruce E.↗

Technologies and Methods Used at the Laboratory for Atmospheric and Space Physics (LASP) to Serve Solar Irradiance Data

The Laboratory for Atmospheric and Space Physics (LASP) at the University of Colorado in Boulder, USA operates the Solar Radiation and Climate Experiment (SORCE) NASA mission, as well as several other NASA spacecraft and instruments. Dozens of Solar Irradiance data sets are produced, managed, and disseminated to the science community. Data are made freely available to the scientific immediately after they are produced using a variety of data access interfaces, including the LASP Interactive Solar Irradiance Datacenter (LISIRD), which provides centralized access to a variety of solar irradiance data sets using both interactive and scriptable/programmatic methods. This poster highlights the key technological elements used for the NASA SORCE mission ground system to produce, manage, and disseminate data to the scientific community and facilitate long-term data stewardship. The poster presentation will convey designs, technological elements, practices and procedures, and software management processes used for SORCE and their relationship to data quality and data management standards, interoperability, NASA data policy, and community expectations.

Pankratz, Chris↗

Memory Optimizations for Sparse Linear Algebra on GPU Hardware

An effort to maximize memory bandwidth utilization for a sparse linear algebra kernel executing on NVIDIA® Tesla V100 and A100 Graphics Processing Units (GPUs) is described. The kernel consists of a block-sparse matrix-vector product and a series of forward/backward triangular solves. The computation is memory-bound and exhibits low arithmetic intensity. Along with a relatively small block size, the data layout poses a challenge to effectively utilize the available memory bandwidth on common GPU architectures. An earlier implementation using a warp to process a single row of the matrix was found to yield good memory performance on the V100 architecture. However, anew approach, which assigns a warp to six rows of the matrix, is proposed for the A100. In addition, two new features offered by the A100 architecture are explored.L2residency control enables a portion of theL2cache to be used for persistent data access, and the asynchronous copy instruction allows data to be loaded directly from main memory into shared memory. Demonstrations show that the new implementation improves memory bandwidth utilization from 71.5% to 81.2% of the peak available on theA100 architecture.

GPU↗

Air Quality (AQ) Monitoring From Space By NASA Using Tempo

In an era marked by escalating environmental concerns, understanding the Earth's atmosphere and its complex interactions is paramount. The Tropospheric Emission Monitoring of POllution (TEMPO) instrument is a cutting-edge venture by NASA that stands at the forefront of Earth observation technology and exemplifies NASA's commitment to unraveling the intricacies of the air we breathe. TEMPO was launched with the primary objective of monitoring air quality. This story map explores the innovative technology behind the TEMPO instrument, its mission objectives, tools and services provided by the Atmospheric Science Data Center (ASDC) for data access and the potential impact of TEMPO data findings on our health and Earth's environmental future.

Hazem Mahmoud↗

OPeNDAP and HDF5 in the Cloud: Techniques and Best Practices for Serving HDF5 Data

In our talk we will discuss considerations and best practices for organizing data and metadata in HDF5 files when serving ESDIS products in S3 with OPeNDAP. The first part of the talk will be devoted to the best practices for organizing data and user-defined metadata including CF conventions in HDF5. We will also address interoperability with netCDF-4 and its role in accessing data in HDF5. In the second part of the talk we will go over OPeNDAP considerations for S3 access to HDF5 data.

HDF5↗

VAS operational procedures and results at the Kansas City Satellite Field Services Station

An operational assessment of VAS data by using a Man-computer Interactive Data Access System (McIDAS) terminal linked by a 9600 band telephone line is discussed. Seven hours of VAS data were processed and edited daily. Data was scheduled 16 hours a day, 7 days a week; however, during this time period there were very few days with 16 hours of data to evalute. The McIDAS terminal, which has 10 display frames and 5 graphics, provide access to the sounding data processed. These data are processed using two procedures. The dwell sounding data are generated by using all 12 spectral channels with a spin budget of 39. To provide coverage for most of the United States, soundings are made starting at 18 minutes after the hour from approximately 49 deg N to 36 deg N and at 48 minutes after the hour from 36 deg N to 26 deg N. The dwell imaging mode uses 11 channels but the spin budge is 17. With the reduced spin budget, retrievals can be made at 18 or 48 minutes after the hour for approximately 44 deg N to 27 deg N. With these constraints a schedule, of data sets was proposed to use the schedule and how the data set could be used are shown.

Heckman, B.↗

The Pilot Land Data System: Report of the Program Planning Workshops

An advisory report to be used by NASA in developing a program plan for a Pilot Land Data System (PLDS) was developed. The purpose of the PLDS is to improve the ability of NASA and NASA sponsored researchers to conduct land-related research. The goal of the planning workshops was to provide and coordinate planning and concept development between the land related science and computer science disciplines, to discuss the architecture of the PLDs, requirements for information science technology, and system evaluation. The findings and recommendations of the Working Group are presented. The pilot program establishes a limited scale distributed information system to explore scientific, technical, and management approaches to satisfying the needs of the land science community. The PLDS paves the way for a land data system to improve data access, processing, transfer, and analysis, which land sciences information synthesis occurs on a scale not previously permitted because of limits to data assembly and access.

Source record↗