Skip to content

About

A microservice for loading a web page at a specified URL and converting it into MHTML format.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

28 Commits

Folders and files

Repository files navigation

Web Page Archive Service

Русская версия

Web Page Archive Service is a utility designed for downloading individual web pages via web scraping. Functional access is provided through the gRPC API. The service enables capturing the full page content, including embedded assets, and storing it in the MHTML format. The resulting data package is delivered as a ZIP archive.

Project Purpose

The service automates the downloading and archiving of a single web page into a ZIP archive using Playwright for scraping. It exposes a gRPC interface for integration into other systems, simplifying the retrieval of web content.

The main goal is to create a template microservice that can be easily extended for parsing and storing web pages in offline MHTML format.

Technologies

Language: C#
Platform: .NET 8
API: gRPC for high-performance RPC
Web scraping: Playwright (uses Chromium)
Archiving: ZipArchive — built-in System.IO.Compression library for creating ZIP archives

Solution Architecture

The recommended approach for a long-running service that frequently creates browser contexts and pages:

  • Create a single IPlaywright and IBrowser instance for the entire application lifetime.
  • For each page download task, create a new IBrowserContext, IPage, and ICDPSession, and dispose of them after use.

Development

cd ~/YourWorkDir
git clone git@github.com:abaula/web-page-archive-svc.git
cd web-page-archive-svc
# build solution
dotnet build src/WebPageArchive.sln
# Install PowerShell (Ubuntu)
sudo apt-get update && sudo apt-get install -y powershell
# Install Chromium browser (headless only)
pwsh src/WebPageArchive/bin/Debug/net8.0/playwright.ps1 install --with-deps --only-shell chromium

Starting the server

cd ~/YourWorkDir/web-page-archive-svc/src/WebPageArchive
dotnet run

Starting the client

cd ~/YourWorkDir/web-page-archive-svc/src/WebPageArchiveClient
dotnet run

Container Deployment

This project uses Podman.

Build the image

cd ~/YourWorkDir
git clone git@github.com:abaula/web-page-archive-svc.git
cd web-page-archive-svc
podman build -t web-page-archive-svc:1.0.0 .

Create the container

podman run -d \
  --name web-page-archive-svc \
  -p 50001:8000 \
  web-page-archive-svc:1.0.0

Notes

  • -p 50001:8000 means port 50001 on the host maps to port 8000 inside the container.
  • The container port 8000 can be changed in src/WebPageArchive/appsettings.json.

License

GNU GPL 3.0 © 2026
Full text: LICENSE

About

A microservice for loading a web page at a specified URL and converting it into MHTML format.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages