Search NASA⌕ Search

SEARCH · Search NASA

Results for “Web Archiving”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Hashes are not suitable to verify fixity of the public archived web

Web archives, such as the Internet Archive, preserve the web and allow access to prior states of web pages. We implicitly trust their versions of archived pages, but as their role moves from preserving curios of the past to facilitating present day adjudication, we are concerned with verifying the fixity of archived web pages, or mementos, to ensure they have always remained unaltered. A widely used technique in digital preservation to verify the fixity of an archived resource is to periodically compute a cryptographic hash value on a resource and then compare it with a previous hash value. If the hash values generated on the same resource are identical, then the fixity of the resource is verified. We tested this process by conducting a study on 16,627 mementos from 17 public web archives. We replayed and downloaded the mementos 39 times using a headless browser over a period of 442 days and generated a hash for each memento after each download, resulting in 39 hashes per memento. The hash is calculated by including not only the content of the base HTML of a memento but also all embedded resources, such as images and style sheets. We expected to always observe the same hash for a memento regardless of the number of downloads. However, our results indicate that 88.45% of mementos produce more than one unique hash value, and about 16% (or one in six) of those mementos always produce different hash values. We identify and quantify the types of changes that cause the same memento to produce different hashes. These results point to the need for defining an archive-aware hashing function, as conventional hashing functions are not suitable for replayed archived web pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Developing Bloom Filters for Web Archives’ Holdings (Final Project Report)

The main goal of the project was to develop a framework for web archives to create Bloom filters based on their holdings of archived web resources. A Bloom filter (BF) is a data structure that, in our scenario, contains hash values of all (or a subset of) URLs, of which an archive has one or more archival copies. Two main use cases fall in scope for this project and are supported by a BF implementation: 1) Sharing of an archive’s holdings (URLs) and 2) Querying the holdings of one or more archives. Since URL strings are hashed before ingested into the BF, the index of an archive is not shared in plain text when the BF is shared with trusted parties. On the other hand, a BF implementation does allow queries for a URL to confirm if an archive indeed has one or more archival copies of that URL. This collaborative project between Los Alamos National Laboratory (LANL) and the Croatian Web Archive (HAW), from the National and University Library in Zagreb (NSK), developed by NSK and University of Zagreb University Computing Center SRCE), aimed at developing software to create BFs, evaluate the scalability of the approach, pilot a search service based on BFs, and design a framework for archives to share their filters with trusted parties. In the remainder of this document, we will report on the work completed with respect to the individual deliverables, outline aspects of future work, and conclude with our recommendations for the use of BFs for the web archiving community.

97 MATHEMATICS AND COMPUTING↗

The DSA Toolkit Shines Light Into Dark and Stormy Archives

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web archive collections are often configured to revisit the same original resource multiple times. This is incredibly useful for understanding an unfolding news story or the evolution of an organization. Unfortunately, over time, some of these original resources can go off-topic and no longer suit the purpose for which the collection was originally created. They can go off-topic due to web site redesigns, changes in domain ownership, financial issues, hacking, technical problems, or because their content has moved on from the original topic. Even though they are off-topic, the archiving system will still capture them, thus it becomes imperative to anyone performing research on these collections to identify these off-topic mementos. Hence, we present the Off-Topic Memento Toolkit, which allows users to detect off-topic mementos within web archive collections. The mementos identified by this toolkit can then be separately removed from a collection or merely excluded from downstream analysis. The following similarity measures are available: byte count, word count, cosine similarity, Jaccard distance, Sørensen-Dice distance, Simhash using raw text content, Simhash using term frequency, and Latent Semantic Indexing via the gensim library. We document the implementation of each of these similarity measures. We possess a gold standard dataset generated by manual analysis, which contains both off-topic and on-topic mementos. Using this gold standard dataset, we establish a default threshold corresponding to the best F1 score for each measure. We also provide an overview of potential future directions that the toolkit may take.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The SHADOZ Data Base: History, Archive Web Guide, and Sample Climatologies

SHADOZ (Southern Hemisphere Additional Ozonesonde) is a project to augment and archive ozonesonde data from ten tropical and subtropical ozone stations. Started in 1998 by NASA's Goddard Space Flight Center and other US and international co-investigators, SHADOZ is an important tool for tropospheric ozone research in the equatorial region. The rationale for SHADOZ is to: (1) validate and improve remote sensing techniques (e.g., the Total Ozone Mapping Spectrometer (TOMS) satellite) for estimating tropical ozone, (2) contribute to climatology and trend analyses of tropical ozone and (3) provide research topics to scientists and educate students, especially in participating countries. SHADOZ is envisioned as a data service to the global scientific community by providing a central public archive location via the internet: http://code9l6.gsfc.nasa.gov/Data_services/shadoz. While the SHADOZ website maintains a standard data format for the archive, it also informs the data users on the differing stations' preparation techniques and data treatment. The presentation navigates through the SHADOZ website to access each station's sounding data and summarize each station's characteristics. Since the start of the project in 1998, the SHADOZ archive has accumulated over 600 ozonesonde profiles and received over 30,000 outside data requests. Data also includes launches from various SHADOZ supported field campaigns, such as, the Indian Ocean Experiment (INDOEX), Sounding of Ozone and Water in the Equatorial Region (SOWER) and Aerosols99 Atlantic Cruise. Using data from the archive, sample climatologies and profiles from selected stations and campaigns will be shown.

White, J. C.↗

Bloom Filter framework for Web Archives

The software provides a framework to build a fast look up layer that reflects the holdings of an archive or database and operates between search/discovery system and disk storage. The software utilizes Bloom filter (BF) data structure. The discovery service powered by the Bloom filter layer comes with a very high level of accuracy yet space-saving, since BF compresses data to fit into RAM.

Balakireva, Lyudmila↗

Summer 2021 Internship Report

The Library Research and Prototyping team (Proto Team) under the research library explores aspects of scholarly communication, covering aspects of infrastructure, interoperability, and persistence. One of the key contributions of the team is the contribution in standardizing and forming the infrastructure for Mementos in web archiving. The Memento framework in web archiving enables access to digital resources in prior versions serving both persistence and the interoperability of the information. During the internship program, I worked on improving the Memento infrastructure by updating the memento validator. The Memento validator provides the functionality of testing the compliance of the web resources to the Memento specification. The application is intended to be used by the community at large such as web archives, researchers, librarians, and general users, with different levels of technical expertise.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Trace Crawler SOFTWARE

The trace crawler is a tool for selective web crawling to archive web resources with well-defined boundaries. The specific web navigation steps (or trace) are formulated for the families of webpages, where layout or HTML structure can be similar but the content is different, for example, GitHub, Slideshare, blogs, etc. The trace is recorded in a json file format.

Balakireva, Lyudmila↗

Summer 2021 Internship Report

During this internship program, I was assigned with developing a Web application to robustify web resources referenced by links (URI-Rs) in PDFs. It was intended for use in LANL Research 2 Library systems such as RASSTI (Review and Approval System for Scientific and Technical Information), where users submit scholarly PDFs that may contain URI-Rs. For such PDFs, the application should, using Web Archives, create robust snapshots (i.e., Mementos) of the web resources referenced by URI-Rs, and notify this robustification to LANL Research Library systems. The Mementos were to be created using the Robust Links Service, and the notifications were to be sent using the Linked Data Notifications (LDN) Protocol.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

JPL, NASA and the Historical Record: Key Events/Documents in Lunar and Mars Exploration

This document represents a presentation about the Jet Propulsion Laboratory (JPL) historical archives in the area of Lunar and Martian Exploration. The JPL archives documents the history of JPL's flight projects, research and development activities and administrative operations. The archives are in a variety of format. The presentation reviews the information available through the JPL archives web site, information available through the Regional Planetary Image Facility web site, and the information on past missions available through the web sites. The presentation also reviews the NASA historical resources at the NASA History Office and the National Archives and Records Administration.

Hooks, Michael Q.↗

ScholarGuard

The ScholarGuard framework aims to address the gap in archiving and preservation efforts for scholarly artifacts beyond traditional research papers, such as software source code, datasets, presentation slides, workflows, protocols, videos, and more. It introduces a prototype system designed to automatically track researchers' outputs across various scholarly productivity portals on the open web, including platforms like GitHub, Slideshare, Figshare, and Wikipedia. The system detects the availability of new scholarly artifacts and applies modern web archiving technology to create a durable archival record, including high-level metadata for each artifact. This metadata is displayed within the system, linking both to the live version and the archived version of the resource, ensuring long-term accessibility and preservation of diverse research outputs. The software serves as a critical tool for preserving the broader spectrum of scholarly contributions, facilitating visibility, searchability, and long-term access to research artifacts beyond the traditional scope of journal publications.

Balakireva, Lyudmila↗

A Virtual Collaborative Environment for Mars Surveyor Landing Site Studies

Over the past year and a half, the Center for Mars Exploration (CMEX) at NASA Ames Research Center (ARC) has been working with the Mars Surveyor Project Office at JPL to promote interactions among the planetary community and to coordinate landing site activities for the Mars Surveyor Project Office. To date, CMEX has been responsible for organizing the first two Mars Surveyor Landing Site workshops, web-archiving resulting information from these workshops, aiding in science evaluations of candidate landing sites, and serving as a liaison between the community and the Project. Most recently, CMEX has also been working with information technologists at Ames to develop a state-of-the-art collaborative web site environment to foster interaction of interested members of the planetary community with the Mars Surveyor Program and the Project Office. The web site will continue to evolve over the next several years as new tools and features are added to support the ongoing Mars Surveyor missions.

Gulick, V.C.↗

A Virtual Collaborative Environment for Mars Surveyor Landing Site Studies

Over the past year and a half, the Center for Mars Exploration (CMEX) at NASA Ames Research Center (ARC) has been working with the Mars Surveyor Project Office at JPL to promote interactions among the planetary community and to coordinate landing site activities for the Mars Surveyor Project Office. To date, CMEX has been responsible for organizing the first two Mars Surveyor Landing Site workshops, web-archiving resulting information from these workshops, aiding in science evaluations of candidate landing sites, and serving as a liaison between the community and the Project. Most recently, CMEX has also been working with information technologists at Ames to develop a state-of-the-art collaborative web site environment to foster interaction of interested members of the planetary community with the Mars Surveyor Program and the Project Office. The web site will continue to evolve over the next several years as new tools and features are added to support the ongoing Mars Surveyor missions.

Gulick, V. C.↗

Memento Protocol Validator Software

The Memento Protocol Validator software provides a suite of tools to validate compliance with the Memento protocol. The Memento protocol is an interoperability solution for accessing web archiving information with the DateTime dimension.

Balakireva, Lyudmila↗

TimeStitch-Memento-Aggregator

Memento Aggregator is a Java service that federates web archives worldwide: given an Original-URL and optional datetime, it discovers mementos and exposes standards-compliant Memento TimeGate and TimeMap APIs.

Balakireva, Lyudmila↗

Integrating Wind Profiling Radars and Radiosonde Observations with Model Point Data to Develop a Decision Support Tool to Assess Upper-Level Winds for Space Launch

On the day-of-launch, the 45th Weather Squadron (45 WS) Launch Weather Officers (LWOs) monitor the upper-level winds for their launch customers to include NASA's Launch Services Program and NASA's Ground Systems Development and Operations Program. They currently do not have the capability to display and overlay profiles of upper-level observations and numerical weather prediction model forecasts. The LWOs requested the Applied Meteorology Unit (AMU) develop a tool in the form of a graphical user interface (GUI) that will allow them to plot upper-level wind speed and direction observations from the Kennedy Space Center (KSC) 50 MHz tropospheric wind profiling radar, KSC Shuttle Landing Facility 915 MHz boundary layer wind profiling radar and Cape Canaveral Air Force Station (CCAFS) Automated Meteorological Processing System (AMPS) radiosondes, and then overlay forecast wind profiles from the model point data including the North American Mesoscale (NAM) model, Rapid Refresh (RAP) model and Global Forecast System (GFS) model to assess the performance of these models. The AMU developed an Excel-based tool that provides an objective method for the LWOs to compare the model-forecast upper-level winds to the KSC wind profiling radars and CCAFS AMPS observations to assess the model potential to accurately forecast changes in the upperlevel profile through the launch count. The AMU wrote Excel Visual Basic for Applications (VBA) scripts to automatically retrieve model point data for CCAFS (XMR) from the Iowa State University Archive Data Server (http://mtarchive.qeol.iastate.edu) and the 50 MHz, 915 MHz and AMPS observations from the NASA/KSC Spaceport Weather Data Archive web site (http://trmm.ksc.nasa.gov). The AMU then developed code in Excel VBA to automatically ingest and format the observations and model point data in Excel to ready the data for generating Excel charts for the LWO's. The resulting charts allow the LWOs to independently initialize the three models 0-hour forecasts against the observations to determine which is the best performing model and then overlay the model forecasts on time-matched observations during the launch countdown to further assess the model performance and forecasts. This paper will demonstrate integration of observed and predicted atmospheric conditions into a decision support tool and demonstrate how the GUI is implemented in operations.

Bauman, William H., III↗