Search NASA⌕ Search

SEARCH · Search NASA

Results for “Web Archive Holdings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Developing Bloom Filters for Web Archives’ Holdings (Final Project Report)

The main goal of the project was to develop a framework for web archives to create Bloom filters based on their holdings of archived web resources. A Bloom filter (BF) is a data structure that, in our scenario, contains hash values of all (or a subset of) URLs, of which an archive has one or more archival copies. Two main use cases fall in scope for this project and are supported by a BF implementation: 1) Sharing of an archive’s holdings (URLs) and 2) Querying the holdings of one or more archives. Since URL strings are hashed before ingested into the BF, the index of an archive is not shared in plain text when the BF is shared with trusted parties. On the other hand, a BF implementation does allow queries for a URL to confirm if an archive indeed has one or more archival copies of that URL. This collaborative project between Los Alamos National Laboratory (LANL) and the Croatian Web Archive (HAW), from the National and University Library in Zagreb (NSK), developed by NSK and University of Zagreb University Computing Center SRCE), aimed at developing software to create BFs, evaluate the scalability of the approach, pilot a search service based on BFs, and design a framework for archives to share their filters with trusted parties. In the remainder of this document, we will report on the work completed with respect to the individual deliverables, outline aspects of future work, and conclude with our recommendations for the use of BFs for the web archiving community.

97 MATHEMATICS AND COMPUTING↗

Bloom Filter framework for Web Archives

The software provides a framework to build a fast look up layer that reflects the holdings of an archive or database and operates between search/discovery system and disk storage. The software utilizes Bloom filter (BF) data structure. The discovery service powered by the Bloom filter layer comes with a very high level of accuracy yet space-saving, since BF compresses data to fit into RAM.

Balakireva, Lyudmila↗

Modern Era Retrospective Restrospective-Analysis for Research and Applications (MERRA) Data and Services at the GES DISC

The Modern Era Retrospective-analysis for Research and Applications (MERRA) dataset is a NASA satellite era, 30 year (1979 - present), reanalysis using the Goddard Earth Observing System Data Assimilation System, Version 5 (GEOS-5). The project, run out of NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center, provides the science and application communities with a state-of-the-art global analysis with emphasis on improved estimates of the hydrological cycle over a broad range of weather and climate time scales. MERRA products are generated as a long-term synthesis that places the NASA EOS suite of observations in a climate context. The MERRA analysis is performed at a horizontal resolution of 2/3 longitude x 1/2 latitude (540x361 global gridpoints) with observational analyses every 6 hours. The MERRA output data will include 3 dimensional state fields for every 6 hourly analysis cycle on 42 pressure levels (or 72 terrain following model coordinate levels) from the surface through the stratosphere. Several data products are specifically designed to support chemistry and stratosphere transport modeling. The 2 dimensional surface and atmospheric diagnostics (numbering 259) are being stored on the native grid at 1 hourly intervals. These include radiation and vertical integrals of the atmosphere for water and energy budget studies and also surface diagnostics where the diurnal cycle is important. The one hourly surface and near surface data product will also facilitate research on the integrated analysis of Earth system observations in the land, ocean and cryosphere. The MERRA products are archived and distributed by the Goddard Earth Sciences Data and Information Services Center (GES DISC) through its Modeling DISC Web (MDISC) portal. Multiple data access methods and services are available for MERRA data through MDISC: (1) Mirador offers a quick, comprehensive search of MERRA and all GES DISC archived data holdings, allowing searches on keywords, location names or latitude/longitude box, and date/time, with responses within a few seconds. (2) Giovanni is a GES DISC developed Web application that provides data visualization and analysis online. Giovanni features popular visualizations such as latitude-longitude maps, animations, cross sections, profiles, time series, etc. and some basic statistical analysis functions such as scatter plots and correlation coefficient maps. Users are able to download results in several different formats, including Google Earth. (3) On-the-fly parameter subsetting of data within a spatial/temporal window is provided through a simple select and click Web page. (4) MERRA data are also available via OPeNDAP, GrADS Data Server (GDS) and can be converted to netCDF on the fly.

Berrick, Stephen W.↗

Availability of Previously Unprocessed ALSEP Raw Instrument Data, Derivative Data, and Metadata Products

In year 2010, 440 original data archival tapes for the Apollo Lunar Science Experiment Package (ALSEP) experiments were found at the Washington National Records Center. These tapes hold raw instrument data received from the Moon for all the ALSEP instruments for the period of April through June 1975. We have recently completed extraction of binary files from these tapes, and we have delivered them to the NASA Space Science Data Cordinated Archive (NSSDCA). We are currently processing the raw data into higher order data products in file formats more readily usable by contemporary researchers. These data products will fill a number of gaps in the current ALSEP data collection at NSSDCA. In addition, we have estabilished a digital, searcheable archive of ALSEP document and metadata as part of the web portal of the Lunar and Planetary Institute. It currently holds approx. 700 documents totaling approx. 40,000 pages

ALSEP↗

NASA's Earth Observing Data and Information System

NASA's Earth Observing System Data and Information System (EOSDIS) has been a central component of NASA Earth observation program for over 10 years. It is one of the largest civilian science information system in the US, performing ingest, archive and distribution of over 3 terabytes of data per day much of which is from NASA s flagship missions Terra, Aqua and Aura. The system supports a variety of science disciplines including polar processes, land cover change, radiation budget, and most especially global climate change. The EOSDIS data centers, collocated with centers of science discipline expertise, archive and distribute standard data products produced by science investigator-led processing systems. Key to the success of EOSDIS is the concept of core versus community requirements. EOSDIS supports a core set of services to meet specific NASA needs and relies on community-developed services to meet specific user needs. EOSDIS offers a metadata registry, ECHO (Earth Observing System Clearinghouse), through which the scientific community can easily discover and exchange NASA s Earth science data and services. Users can search, manage, and access the contents of ECHO s registries (data and services) through user-developed and community-tailored interfaces or clients. The ECHO framework has become the primary access point for cross-Data Center search-and-order of EOSDIS and other Earth Science data holdings archived at the EOSDIS data centers. ECHO s Warehouse Inventory Search Tool (WIST) is the primary web-based client for discovering and ordering cross-discipline data from the EOSDIS data centers. The architecture of the EOSDIS provides a platform for the publication, discovery, understanding and access to NASA s Earth Observation resources and allows for easy integration of new datasets. The EOSDIS also has developed several methods for incorporating socioeconomic data into its data collection. Over the years, we have developed several methods for determining needs of the user community including use of the American Customer Satisfaction Index and a broad metrics program.

Mitchell, Andrew E.↗

Protein Data Bank: A Comprehensive Review of 3D Structure Holdings and Worldwide Utilization by Researchers, Educators, and Students

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB), funded by the United States National Science Foundation, National Institutes of Health, and Department of Energy, supports structural biologists and Protein Data Bank (PDB) data users around the world. The RCSB PDB, a founding member of the Worldwide Protein Data Bank (wwPDB) partnership, serves as the US data center for the global PDB archive housing experimentally-determined three-dimensional (3D) structure data for biological macromolecules. As the wwPDB-designated Archive Keeper, RCSB PDB is also responsible for the security of PDB data and weekly update of the archive. RCSB PDB serves tens of thousands of data depositors (using macromolecular crystallography, nuclear magnetic resonance spectroscopy, electron microscopy, and micro-electron diffraction) annually working on all permanently inhabited continents. RCSB PDB makes PDB data available from its research-focused web portal at no charge and without usage restrictions to many millions of PDB data consumers around the globe. It also provides educators, students, and the general public with an introduction to the PDB and related training materials through its outreach and education-focused web portal. This review article describes growth of the PDB, examines evolution of experimental methods for structure determination viewed through the lens of the PDB archive, and provides a detailed accounting of PDB archival holdings and their utilization by researchers, educators, and students worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Improvements in Space Geodesy Data Discovery at the CDDIS

The Crustal Dynamics Data Information System (CDDIS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to store space geodesy and geodynamics related data products in a central data bank. to maintain information about the archival of these data, and to disseminate these data and information in a timely manner to a global scientific research community. The archive consists of GNSS, laser ranging, VLBI, and DORIS data sets and products derived from these data. The CDDIS is one of NASA's Earth Observing System Data and Information System (EOSDIS) distributed data centers; EOSDIS data centers serve a diverse user community and arc tasked to provide facilities to search and access science data and products. Several activities are currently under development at the CDDIS to aid users in data discovery, both within the current community and beyond. The CDDIS is cooperating in the development of Geodetic Seamless Archive Centers (GSAC) with colleagues at UNAVCO and SIO. TIle activity will provide web services to facilitate data discovery within and across participating archives. In addition, the CDDIS is currently implementing modifications to the metadata extracted from incoming data and product files pushed to its archive. These enhancements will permit information about COOlS archive holdings to be made available through other data portals such as Earth Observing System (EOS) Clearinghouse (ECHO) and integration into the Global Geodetic Observing System (GGOS) portal.

Noll, C.↗

Enabling API Access to the Space Weather Services at the Community Coordinated Modeling Center

Over the span of 20 years, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) has been leading a number of community-driven services and applications that provide a free and open access to the cutting-edge space weather and Heliophysics models through a simple web-based interface. CCMC also oversees an open archive of user model simulations and related metadata, maintains space weather-related data streams, curates validation and event datasets, and more. To maximize utility of its complementary services and data holdings, CCMC has been gradually building up an ad-hoc set of interfaces and specifications that facilitate coupling and interconnection within the organization, while also simplifying management and monitoring of the data. As the models continue to grow in maturity and complexity, CCMC is looking to reduce the complexity for its end users by making public some of the internal APIs as well as implementing dedicated interfaces as required by the community. In this presentation, we will overview the current run services at CCMC and will describe our current and near-future efforts in providing interfaces to these services.

space weather↗

Intelligent Systems Technologies and Utilization of Earth Observation Data

The addition of raw data and derived geophysical parameters from several Earth observing satellites over the last decade to the data held by NASA data centers has created a data rich environment for the Earth science research and applications communities. The data products are being distributed to a large and diverse community of users. Due to advances in computational hardware, networks and communications, information management and software technologies, significant progress has been made in the last decade in archiving and providing data to users. However, to realize the full potential of the growing data archives, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system (KBS). Potential Intelligent Archive concepts include: 1) Mining archived data holdings to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services; 3) Recognizing the value of results, indexing and formatting them for easy access; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building; and 5) Being aware of other nodes in the KBS, participating in open systems interfaces and protocols for virtualization, and achieving collaborative interoperability.

Ramapriyan, H. K.↗

Intelligent Systems Technologies to Assist in Utilization of Earth Observation Data

With the launch of several Earth observing satellites over the last decade, we are now in a data rich environment. From NASA's Earth Observing System (EOS) satellites alone, we are accumulating more than 3 TB per day of raw data and derived geophysical parameters. The data products are being distributed to a large user community comprising scientific researchers, educators and operational government agencies. Notable progress has been made in the last decade in facilitating access to data. However, to realize the full potential of the growing archives of valuable scientific data, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system. Potential Intelligent Archive concepts include: 1) Mining archived data holdings using Intelligent Data Understanding algorithms to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services involved in a scientific enterprise; 3) Recognizing the value of results, indexing and formatting them for easy access, and delivering them to concerned individuals; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building (i.e., the transformations from data to information to knowledge) instead of just data pipelining; and 5) Being aware of other nodes in the knowledge building system, participating in open systems interfaces and protocols for virtualization, and collaborative interoperability. This paper presents some of these concepts and identifies issues to be addressed by research in future intelligent systems technology.

Ramapriyan, Hampapuram K.↗

BIOME: A browser-aware search and order system

The Oak Ridge National Laboratory (ORNL) Distributed Active Archive Center (DAAC), which is associated with NASA's Earth Observing System Data and Information System (EOSDIS), provides access to a large number of tabular and imagery datasets used in ecological and environmental research. Because of its large and diverse data holdings, the challenge for the ORNL DAAC is to help users find data of interest from the hundreds of thousands of files available at the DAAC without overwhelming them. Therefore, the ORNL DAAC developed the Biogeochemical Information Ordering Management Environment (BIOME), a search and order system for the World Wide Web (WWW). The WWW provides a new vehicle that allows a wide range of users access to the data. This paper describes the specialized attributes incorporated into BIOME that allow researchers easy access to an otherwise bewildering array of data products.

Grubb, Jon W.↗

HRP Data Management Plan

The purpose of Human Research Program Data Management Plan (DMP) is to define the processes and activities required for the overall management of the research data collected and managed by HRP throughout their life cycle. New updates to the Data Management Plan in 2023 include 1. CAPABILITIES AND SERVICES Data Repositories. Principal Investigators (PIs) funded by HRP may be asked to submit data to one of several NASA data repositories. HRP archives data in the NASA Life Sciences Portal (NLSP) that it considers to be unique and high value. This includes data from human subjects in space flight (ISS and commercial flights) and ground analogs to spaceflight; spaceflight tech demos involving humans; human omics data including the microbiome; parabolic flight studies; and the NASA Space Radiation Laboratory (NSRL). The Open Science Data Repository (OSDR) includes The Ames Life Sciences Data Archive (ALSDA), used to archive non-human biological data (e.g., animal) generated by the Human Research program, and GeneLab, available to HRP PIs to archive non-human omics data. Catalog for search and retrieval. A catalog of non-human HRP life science experiments, with all associated descriptions (mission, payload, hardware, and personnel related information), and biospecimens is provided on the NLSP public web site for search and retrieval. 2. IRB ROLE IN RETURN OF INDIVIDUAL RESEARCH RESULTS The NASA IRB manages the process for incidental findings and for returning results to subjects for studies for which NASA IRB is the IRB of record. Omics data, especially genomics data, may generate information significant to the health of or risk to a research subject. These data potentially hold the keys to understand lifetime risks of chronic diseases, such as cancer, as well as risks associated with exposures common in space flight. 3. UPDATE OF TERMS – IDENTIFIABLE AND ATTRIBUTABLE DATA HRP now follows Federal and NASA policy by using “identifiable” instead of “attributable” for Personally Identifiable Information (PII). 4. POLICY ABOUT INTERNAL NON-RESEARCH USE OF DATA The HRP Chief Scientist grants access to data from HRP-funded research for non-research internal use that includes program management, customer facilitation, strategic planning, and risk research planning. Typical HRP personnel granted access to HRP research data for internal use include the Element Scientist, Subject Matter Experts (SME), and Data/bioinformatics Scientists. If data accessed for Internal Use is provided to an intramural or extramural scientist for hypothesis driven research, all Federal and NASA regulations (e.g., IRB review) regarding human subject research apply.

Data Management Plan↗

NASA Scientific Data Purchase Project: From Collection to User

NASA's Scientific Data Purchase (SDP) project is currently a $70 million operation managed by the Earth Science Applications Directorate at Stennis Space Center. The SDP project was developed in 1997 to purchase scientific data from commercial sources for distribution to NASA Earth science researchers. Our current data holdings include 8TB of remote sensing imagery consisting of 18 products from 4 companies. Our anticipated data volume is 60 TB by 2004, and we will be receiving new data products from several additional companies. Our current system capacity is 24 TB, expandable to 89 TB. Operations include tasking of new data collections, archive ordering, shipment verification, data validation, distribution, metrics, finances, customer feedback, and technical support. The program has been included in the Stennis Space Center Commercial Remote Sensing ISO 9001 registration since its inception. Our operational system includes automatic quality control checks on data received (with MatLab analysis); internally developed, custom Web-based interfaces that tie into commercial-off-the-shelf software; and an integrated relational database that links and tracks all data through operations. We've distributed nearly 1500 datasets, and almost 18,000 data files have been downloaded from our public web site; on a 10-point scale, our customer satisfaction index is 8.32 at a 23% response level. More information about the SDP is available on our Web site.

Nicholson, Lamar↗

Value-added Data Services at the Goddard Earth Sciences Data and Information Services Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), in addition to serving the Earth Science community as one of the major Distributed Active Archives Centers (DAACs), provides much more than just data. Among the value-added services available to general users are subsetting data spatially and/or by parameter, online analysis (to avoid downloading unnecessarily all the data), and assistance in obtaining data from other centers. Services available to data producers and high-volume users include consulting on building new products with standard formats and metadata and construction of data management systems. A particularly useful service is data processing at the DISC (i.e., close to the input data) with the users algorithm. This can take a number of different forms: as a configuration-managed algorithm within the main processing stream; as a stand-alone program next to the on-line data storage; as build-it-yourself code within the Near-Archive Data Mining (NADM) system; or as an on-the-fly analysis with simple algorithms embedded into the web-based tools. Partnerships between the GES DISC and scientists, both producers and users, allow the scientists to concentrate on science, while the GES DISC handles the data management, e.g., formats, integration, and data processing. The existing data management infrastructure at the GES DISC supports a wide spectrum of options: from simple data support to sophisticated on-line analysis tools, producing economies of scale and rapid time-to-deploy. At the same time, such partnerships allow the GES DISC to serve the user community more efficiently and to better prioritize on-line holdings. Several examples of successful partnerships are described in the presentation.

Leptoukh, Gregory G.↗

BIOME: A scientific data archive search-and-order system using browser-aware, dynamic pages

The Oak Ridge National Laboratory's (ORNL) Distributed Active Archive Center (DAAC) is a data archive and distribution center for the National Air and Space Administration's (NASA) Earth Observing System Data and Information System (EOSDIS). Both the Earth Observing System (EOS) and EOSDIS are components of NASA's contribution to the US Global Change Research Program through its Mission to Planet Earth Program. The ORNL DAAC provides access to data used in ecological and environmental research such as global change, global warming, and terrestrial ecology. Because of its large and diverse data holdings, the challenge for the ORNL DAAC is to help users find data of interest from the hundreds of thousands of files available at the DAAC without overwhelming them. Therefore, the ORNL DAAC has developed the Biogeochemical Information Ordering Management Environment (BIOME), a customized search and order system for the World Wide Web (WWW). BIOME is a public system located at http://www-eosdis. ornl.gov/BIOME/biome.html.

CLIMATIC CHANGE↗

Progress and Plans in Support of the Polar Community

Feedback provided by the Antarctic community has proven instrumental in positively influencing the direction of the GCMD's development. For example, in response to requests for a stand alone metadata authoring tool, a new shareable software package called docBUILDER solo will be released to the public in March 2006. This tool permits researchers to document their data during experiments and observational periods in the field. The international polar community has also played a key role in encouraging support for the foreign language character set in the metadata display and tools (10% of the records in the AMD hold foreign characters). In the upcoming release, the full ISO character set, which also includes mathematical symbols, will be supported. Additional upgrades include the ability for users to search for data sets based on pre-selected temporal and spatial resolution ranges. Data providers are strongly encouraged to populate the resolution fields for their data sets, although these fields are not currently required. In prior versions, browser incompatibilities often resulted in unreliable performance for users attempting to initiate a spatial search using a map based on Java applet technology. The GCMD will offer an integrated Google map and date search, replacing the applet technology and enhancing the geospatial and temporal searches. It is estimated that 30% of the records in the AMD have direct access to data. A growing number of these records can be accessed through data service links. Related data services are therefore becoming valuable assets in facilitating the use and visualization of data. Users will gain the ability to refine services using the same options as those available for data set searches. Data providers are encouraged to describe available data-related services through the directory. Future plans include offering web services through a SOAP interface and extending semantic queries for the polar regions through the use of ontologies. The Open Archives Initiative's (OAI) Protocol for Metadata Harvesting (PMH) has been successfully tested with several organizations and appears to be a prime candidate for sharing metadata within the community. The GCMD anticipates contributing to the design of the data management system for the International Polar Year and to the ongoing efforts in the years to come. Further enhancements will be discussed at the meeting.

Olsen, Lola M.↗