Search NASA⌕ Search

SEARCH · Search NASA

Results for “Web Archiving”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Hashes are not suitable to verify fixity of the public archived web

Web archives, such as the Internet Archive, preserve the web and allow access to prior states of web pages. We implicitly trust their versions of archived pages, but as their role moves from preserving curios of the past to facilitating present day adjudication, we are concerned with verifying the fixity of archived web pages, or mementos, to ensure they have always remained unaltered. A widely used technique in digital preservation to verify the fixity of an archived resource is to periodically compute a cryptographic hash value on a resource and then compare it with a previous hash value. If the hash values generated on the same resource are identical, then the fixity of the resource is verified. We tested this process by conducting a study on 16,627 mementos from 17 public web archives. We replayed and downloaded the mementos 39 times using a headless browser over a period of 442 days and generated a hash for each memento after each download, resulting in 39 hashes per memento. The hash is calculated by including not only the content of the base HTML of a memento but also all embedded resources, such as images and style sheets. We expected to always observe the same hash for a memento regardless of the number of downloads. However, our results indicate that 88.45% of mementos produce more than one unique hash value, and about 16% (or one in six) of those mementos always produce different hash values. We identify and quantify the types of changes that cause the same memento to produce different hashes. These results point to the need for defining an archive-aware hashing function, as conventional hashing functions are not suitable for replayed archived web pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Developing Bloom Filters for Web Archives’ Holdings (Final Project Report)

The main goal of the project was to develop a framework for web archives to create Bloom filters based on their holdings of archived web resources. A Bloom filter (BF) is a data structure that, in our scenario, contains hash values of all (or a subset of) URLs, of which an archive has one or more archival copies. Two main use cases fall in scope for this project and are supported by a BF implementation: 1) Sharing of an archive’s holdings (URLs) and 2) Querying the holdings of one or more archives. Since URL strings are hashed before ingested into the BF, the index of an archive is not shared in plain text when the BF is shared with trusted parties. On the other hand, a BF implementation does allow queries for a URL to confirm if an archive indeed has one or more archival copies of that URL. This collaborative project between Los Alamos National Laboratory (LANL) and the Croatian Web Archive (HAW), from the National and University Library in Zagreb (NSK), developed by NSK and University of Zagreb University Computing Center SRCE), aimed at developing software to create BFs, evaluate the scalability of the approach, pilot a search service based on BFs, and design a framework for archives to share their filters with trusted parties. In the remainder of this document, we will report on the work completed with respect to the individual deliverables, outline aspects of future work, and conclude with our recommendations for the use of BFs for the web archiving community.

97 MATHEMATICS AND COMPUTING↗

The DSA Toolkit Shines Light Into Dark and Stormy Archives

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web archive collections are often configured to revisit the same original resource multiple times. This is incredibly useful for understanding an unfolding news story or the evolution of an organization. Unfortunately, over time, some of these original resources can go off-topic and no longer suit the purpose for which the collection was originally created. They can go off-topic due to web site redesigns, changes in domain ownership, financial issues, hacking, technical problems, or because their content has moved on from the original topic. Even though they are off-topic, the archiving system will still capture them, thus it becomes imperative to anyone performing research on these collections to identify these off-topic mementos. Hence, we present the Off-Topic Memento Toolkit, which allows users to detect off-topic mementos within web archive collections. The mementos identified by this toolkit can then be separately removed from a collection or merely excluded from downstream analysis. The following similarity measures are available: byte count, word count, cosine similarity, Jaccard distance, Sørensen-Dice distance, Simhash using raw text content, Simhash using term frequency, and Latent Semantic Indexing via the gensim library. We document the implementation of each of these similarity measures. We possess a gold standard dataset generated by manual analysis, which contains both off-topic and on-topic mementos. Using this gold standard dataset, we establish a default threshold corresponding to the best F1 score for each measure. We also provide an overview of potential future directions that the toolkit may take.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Bloom Filter framework for Web Archives

The software provides a framework to build a fast look up layer that reflects the holdings of an archive or database and operates between search/discovery system and disk storage. The software utilizes Bloom filter (BF) data structure. The discovery service powered by the Bloom filter layer comes with a very high level of accuracy yet space-saving, since BF compresses data to fit into RAM.

Balakireva, Lyudmila↗

Summer 2021 Internship Report

The Library Research and Prototyping team (Proto Team) under the research library explores aspects of scholarly communication, covering aspects of infrastructure, interoperability, and persistence. One of the key contributions of the team is the contribution in standardizing and forming the infrastructure for Mementos in web archiving. The Memento framework in web archiving enables access to digital resources in prior versions serving both persistence and the interoperability of the information. During the internship program, I worked on improving the Memento infrastructure by updating the memento validator. The Memento validator provides the functionality of testing the compliance of the web resources to the Memento specification. The application is intended to be used by the community at large such as web archives, researchers, librarians, and general users, with different levels of technical expertise.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Trace Crawler SOFTWARE

The trace crawler is a tool for selective web crawling to archive web resources with well-defined boundaries. The specific web navigation steps (or trace) are formulated for the families of webpages, where layout or HTML structure can be similar but the content is different, for example, GitHub, Slideshare, blogs, etc. The trace is recorded in a json file format.

Balakireva, Lyudmila↗

Summer 2021 Internship Report

During this internship program, I was assigned with developing a Web application to robustify web resources referenced by links (URI-Rs) in PDFs. It was intended for use in LANL Research 2 Library systems such as RASSTI (Review and Approval System for Scientific and Technical Information), where users submit scholarly PDFs that may contain URI-Rs. For such PDFs, the application should, using Web Archives, create robust snapshots (i.e., Mementos) of the web resources referenced by URI-Rs, and notify this robustification to LANL Research Library systems. The Mementos were to be created using the Robust Links Service, and the notifications were to be sent using the Linked Data Notifications (LDN) Protocol.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ScholarGuard

The ScholarGuard framework aims to address the gap in archiving and preservation efforts for scholarly artifacts beyond traditional research papers, such as software source code, datasets, presentation slides, workflows, protocols, videos, and more. It introduces a prototype system designed to automatically track researchers' outputs across various scholarly productivity portals on the open web, including platforms like GitHub, Slideshare, Figshare, and Wikipedia. The system detects the availability of new scholarly artifacts and applies modern web archiving technology to create a durable archival record, including high-level metadata for each artifact. This metadata is displayed within the system, linking both to the live version and the archived version of the resource, ensuring long-term accessibility and preservation of diverse research outputs. The software serves as a critical tool for preserving the broader spectrum of scholarly contributions, facilitating visibility, searchability, and long-term access to research artifacts beyond the traditional scope of journal publications.

Balakireva, Lyudmila↗

Memento Protocol Validator Software

The Memento Protocol Validator software provides a suite of tools to validate compliance with the Memento protocol. The Memento protocol is an interoperability solution for accessing web archiving information with the DateTime dimension.

Balakireva, Lyudmila↗

TimeStitch-Memento-Aggregator

Memento Aggregator is a Java service that federates web archives worldwide: given an Original-URL and optional datetime, it discovers mementos and exposes standards-compliant Memento TimeGate and TimeMap APIs.

Balakireva, Lyudmila↗

Robustifying Links To Combat Reference Rot

Links to web resources frequently break, and linked content can change at unpredictable rates. These dynamics of the Web are detrimental when references to web resources provide evidence or supporting information. In this paper, we highlight the significance of reference rot, provide an overview of existing techniques and their characteristics to address it, and introduce our Robust Links approach, including its web service and underlying API. Robustifying links offers a proactive, uniform, and machine-actionable way to combat reference rot. In addition, we discuss our reasoning and approach aimed at keeping the approach functional for the long term. To showcase our approach, we have robustified all links in this article.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A new database website for nuclear level densities

We introduce a new open-access, web-based database (http://nld.ascsn.net), Current Archive of Nuclear Density of Levels (CANDL), that hosts experimental nuclear level density (NLD) datasets from a variety of techniques and energy ranges. Built using the Dash framework in Python, the database is designed to be interactive and user-friendly, allowing researchers to search, visualize, fit, and export NLD data with minimal effort. This resource includes data extracted from evaporation spectra, Oslo method variants, and other experimental techniques that cover excitation energies beyond the neutron resonance region. The database supports on-the-fly fitting with two widely-used phenomenological models—the Constant Temperature (CT) model and the Back-Shifted Fermi Gas (BSFG) model—selected for their simplicity and computational efficiency. Future versions aim to include additional datasets and model types, as well as easy-to-use interfaces to data science techniques. Here, this platform offers a vital tool for the nuclear physics, astrophysics, medicine, and reactor design communities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Interactive Web Application for Traffic Simulation Data Management and Visualization

As traffic simulation software becomes more effective for realistically simulating and analyzing traffic dynamics and vehicle interactions on the mesoscopic and microscopic level, the management, dissemination, and collaborative visualization of traffic simulation results produced by individual transportation planners presents a significant challenge. Existing online content management systems have a very limited capability in allowing users to query specific traffic simulation scenarios and geospatially visualize simulation results through shareable and interactive web interfaces. This paper presents a web-based application for promoting the archiving, sharing, and visualization of large-scale traffic simulation outputs. The application is developed to enhance cyber-physical controls, communications, and public education for collaborative transportation planning. Unique features of the web application include: (a) allowing users to upload their new traffic simulation scenarios (parameters and outputs), as well as search existing scenarios using easily accessible interfaces; (b) optimizing simulation output files with heterogeneous data formats and projected coordinate systems for web-based storage and management using a scalable and searchable data/metadata standard; (c) standardizing user-uploaded simulation outputs using web interfaces and data processing libraries with parallel computing capacity; and (d) providing shareable web visual interfaces for visualizing the traffic flow and signal information stored in simulation outputs (e.g., regional traffic patterns and individual vehicle interactions) and visually comparing multiple simulation outputs both spatially and temporally. Furthermore, the paper presents the conceptual design and implementation of this application, and demonstrates the application’s performance for sharing, comparing, and visualizing simulation outputs from VISSIM and SUMO, two commonly used traffic simulation software programs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Slow control and data acquisition development in the Mu2e experiment

The muon campus program at Fermilab includes the Mu2e experiment that will search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus, improving by four orders of magnitude the search sensitivity reached so far.Mu2e’s Trigger and Data Acquisition System (TDAQ) uses {\it otsdaq} solution. Developed at Fermilab, {\it otsdaq} uses the {\it artdaq} DAQ framework and {\it art} analysis framework, for event transfer, filtering, and processing.{\it otsdaq} is an online DAQ software suite with a focus on flexibility and scalability, and provides a multi-user interface accessible through a web browser.A Detector Control System (DCS) for monitoring, controlling, alarming, and archiving has been developed using the Experimental Physics and Industrial Control System (EPICS) open source Platform. The DCS System has also been integrated into {\it otsdaq}, providing a GUI multi-user, web-based control, and monitoring dashboard.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗