Skip to content

Repository files navigation

Wayback-Archive banner

Rescue whole websites from the Wayback Machine.

PyPI version Build License PyPI downloads per month GitHub stars


Wayback-Archive is a command-line tool that downloads an archived website from the Wayback Machine and rebuilds it as a folder you can open offline. It fetches every page and asset it can find in the snapshot, rewrites the links to local paths and removes the Wayback Machine's own toolbar, scripts and URL prefixes, so the copy looks like the site did on that day. wget --mirror and httrack do not understand Wayback Machine URLs; this does.

The python.org homepage of 31 December 2005, rescued and opened from the local copy: original layout, stylesheet and images, no Wayback Machine toolbar

Features

  • Saves HTML, CSS, JavaScript, images and fonts, following links in HTML, CSS and JavaScript until nothing is left to fetch.
  • Every link in the copy points at the local file, so pages open and lead to each other without the Wayback toolbar or its URL prefixes.
  • A page or file missing from the snapshot is looked up at up to three other Wayback timestamps (a day either side and a week earlier) before it is given up on.
  • A capture that is an archived error or a Cloudflare challenge page is replaced with the nearest good capture.
  • Google Fonts are saved locally, and a font the archive returns as an HTML error page is dropped instead of breaking the page.
  • Pages still work when the archive lost jQuery: a copy is fetched from code.jquery.com instead.
  • Social icon groups, button links and cookie banners keep working in the copy.
  • Analytics and ads are stripped by default; external iframes too, with REMOVE_EXTERNAL_IFRAMES=true.
  • Optional HTML, CSS, JavaScript and image minification for a smaller folder.

Quick start

pipx install wayback-archive
WAYBACK_URL="https://web.archive.org/web/20051231235226/http://www.python.org/" MAX_FILES=50 wayback-archive
open output/index.html                     # Linux: xdg-open output/index.html

Use 1.5.0 or later: in 1.4.6 and earlier an archived page can make the tool fetch any host directly and save the response (#55), and pipx upgrade wayback-archive moves an existing install (one made with Python 3.9 needs pipx reinstall --python python3.10 wayback-archive instead). Needs Python 3.10 or newer (pip on 3.9 finds no release it can install); pip install works too, and the other install paths are in Getting started. That run takes about a minute, ends with a Download Complete! block counting the files downloaded, failed and skipped, and opens from disk as the page above: the homepage, its stylesheet and its images are all within the first 50 files. Drop MAX_FILES to fetch the whole site. On another site a short run can stop before a stylesheet that is only reached through @import, because the crawl fetches files in the order it finds them. There are no flags: every setting is an environment variable, listed in Configuration.

Documentation

The full documentation is at https://geiserx.github.io/Wayback-Archive/.

Related projects

Wayback-Diff, Way-CMS, web-mirror, media-download, n8n-nodes-way-cms (archived).

License

GPL-3.0-or-later