Tutorial by Webrecorder.
Browsertrix is a full professional web archiving service. It is a paid service with support included in subscriptions.
With advanced technologies used in large institutional archives it offers possibilities for archiving content that is otherwise difficult to capture, e.g. social media pages. There are still content types that cannot be archived, but Browsertrix can successfully preserve much content that text-based web crawlers cannot handle.
The primary purpose of Browsertrix is preservation and documentation. Content is stored in professional web archive files, in the WACZ file format.
For research purposes the WACZ file format can present difficulties in sorting, extracting and analyzing data. This is easier to do with content stored in web page formats and file folders such as regular text-based crawlers like HTTrack Website Copier or A1 Website Download will provide.
However, the option of preserving very good copies and content that cannot be captured by text-based crawlers may be relevant in some cases.
For handling the WACZ files provided by Browsertrix, they may be replayed with the dedicated application ReplayWeb.page application from the Browsertrix service provider.
The best option for sorting and analyzing archived content is the web archive interface, SolrWayback, developed and maintained by Thomas Egense from the national Danish web archive, Netarkivet, and other contributors.
SolrWayback is available as open source software, but it must be installed to a server. A server and the skills to set up the installation are thus prerequisites for this solution. A description of features in SolrWayback may be found in Web Archives and Web Archiving - Introduction for Scholars and Students by Asger Harlung, 2025, chapter 8.2.5
Service: https://webrecorder.net/browsertrix/
Developer YouTube Channel: https://www.youtube.com/@webrecorder
Works on: