A simple, automated ISO mirroring system for homelabs and personal use. This project utilizes an ephemeral runner architecture via GitHub Actions, fetching operating system images through multi-threaded downloaders (aria2c) and synchronizing them directly to Google Drive (rclone) to generate a static web directory.
Browse the Web Directory: https://tiw302.github.io/iso-deployment-bot
Library Tab — command generator & direct downloads |
Discovery Tab — metadata tags & search filter |
About Tab — homelab guides & statistics |
| Introduction | Setup & Build | Components | Resources |
|---|---|---|---|
| Overview | Requirements | Core Scripts | Web Directory |
| Motivation | Installation | Data Sources | Contributing |
| Design Choices | Cloud Config | Maintenance Tools | License |
os-deployment-library (OSDL) is a modest collection of Python scripts designed to automate the downloading and storing of operating system ISOs.
It was built as a practical workaround for homelab operators or developers who frequently need to provision virtual machines but find it inconvenient to repeatedly download images from slow or distant official mirrors. By utilizing GitHub Actions as a free compute runner and Rclone to upload files to Google Drive, it maintains a personal repository of OS images with minimal manual intervention.
When setting up test environments or home servers, grabbing an ISO directly from upstream sources can sometimes be frustrating:
- Download Speeds: Official mirrors can be slow depending on your geographic location or their current load.
- Archival: Older point-releases are often removed from upstream sites once a newer version drops, making reproducible setups difficult.
- Manual Effort: Keeping track of which distributions have released new versions and manually downloading them is tedious.
This project attempts to mitigate these minor annoyances by running an automated script overnight. It fetches the required files using multi-threaded tools and places them into cloud storage, generating a basic web page so the files are easy to find the next day.
The project is not an enterprise-grade infrastructure tool, but rather a pragmatic script built around a few specific choices to keep it free and easy to maintain:
Stateless Execution: The system does not require a persistent server. It runs entirely on ephemeral GitHub Actions runners. It determines what needs to be downloaded by comparing a local JSON database (src/os_deployment_library/distros.json) against what currently exists in the target Google Drive.
Parallel Downloading: Instead of standard single-thread downloads, it uses aria2c with 16 connections. This helps finish the downloads quickly before the GitHub Action job times out.
Basic Web Index: Browsing files in Google Drive can be slow. To make downloading the ISOs easier, the script automatically generates a static HTML page (web/index.html) using a simple Python template after every successful sync.
Collision Prevention: Since multiple distributions (especially SourceForge downloads) may resolve to the generic name download.iso or latest.iso, the system uses an intelligent filename resolver. It extracts actual filenames from redirected paths or automatically prepends unique parent folders (e.g. dr460nized-latest.iso) to prevent clobbering in the cloud library.
The workflow runs twice a week (Mondays and Thursdays) via a GitHub Actions Cron schedule (.github/workflows/daily_sync.yml). The sequence of events is straightforward:
- Check State:
sync.pyreads the desired list of ISOs fromsrc/os_deployment_library/distros.jsonand uses Rclone to check which ones are missing from your Google Drive. - Download: Any missing files are downloaded to the GitHub runner's temporary storage using
aria2c. - Upload: The successfully downloaded files are moved to Google Drive via Rclone.
- Cleanup: Because GitHub runners have limited disk space (~14GB), the script deletes local files immediately after uploading to make room for the next one.
- Generate Index:
generate_index.pycreates a fresh static HTML page based on what is now available in the database.
| Component | Requirement |
|---|---|
| Runtime | Python 3.8+ |
| Downloader | aria2 |
| Cloud Sync | rclone |
| Storage Target | Google Drive (or any Rclone-compatible cloud storage) with adequate free space. |
If you want to run the scripts manually on your own machine instead of GitHub Actions:
# 1. Tidy up the database (sort entries and remove duplicates)
python3 tools/refactor.py
# 2. (Optional) Fetch the latest ISO checksums
python3 tools/fetch_checksums.py
# 3. Run the main download/upload script
python3 src/scripts/sync.py
# 4. Generate the static web index
python3 src/scripts/generate_index.py
# 5. Run the unit test suite to verify code correctness
python3 -m unittest discover -s tests -p "test_*.py" -v
# 6. Check formatting
ruff check && ruff format --checkTo let GitHub Actions access your Google Drive without committing passwords to the repository, you need to provide your Rclone configuration as a base64-encoded secret.
-
Setup your remote locally using
rclone config. Make sure the remote name in your config matches the one expected insync.py. -
Find where your config file is stored by running:
rclone config file. -
Encode the file's contents to base64:
base64 -w 0 <path_to_rclone.conf>
-
In your GitHub repository, go to Settings > Secrets and variables > Actions.
-
Create a new repository secret named
RCLONE_CONF_DATAand paste the base64 string.
The primary list of all operating systems is kept in a simple JSON file inside src/os_deployment_library/distros.json. To track a new OS, just add an entry:
{
"name": "Ubuntu 24.04 LTS",
"url": "https://releases.ubuntu.com/24.04/ubuntu-24.04-desktop-amd64.iso",
"category": "Linux",
"description": "The latest Long Term Support release of Ubuntu."
}The next time the script runs, it will notice the new entry, download the ISO, upload it to the cloud, and update the web page.
The repository is divided into core operational scripts and a few helper tools used to maintain the database.
src/os_deployment_library/distros.json: The single source of truth containing the list of all ISO URLs.src/scripts/sync.py: The main script that handles the logic of downloading and uploading.src/scripts/generate_index.py: The script responsible for creating the HTML front-end.
These are optional scripts located in the tools/ folder, written to make managing a large list of ISOs less manual:
tools/fetch_top_distros.py: A basic scraper to find popular Linux distributions and add them to the list.tools/fetch_descriptions_wikipedia.py: Reaches out to the Wikipedia API to automatically pull short text descriptions for the OSs in the database.tools/fetch_checksums.py: A best-effort discovery tool that guesses and downloads SHA256/SHA512 checksum files to verify ISO integrity.tools/check_links.py: Pings every URL indistros.jsonto check for 404/dead links, so they can be updated or removed.tools/cleanup_db.py: Compares database entries against what actually exists in Google Drive, checks missing URLs, and removes broken/un-mirrored entries to keep the list clean.tools/refactor.py: A formatting script to sort the list alphabetically and remove any accidental duplicates.tools/sync.sh: A helper bash script to automatically commit and push any changes detected in the repository.
This repository includes automated workflows built on GitHub Actions:
Runs on every push and pull request to master/main branches.
- Syntax Check: Verifies that all Python scripts compile without syntax errors.
- Unit Tests: Executes the comprehensive test suite (
tests/) containing 50+ test cases to verify filename resolution, collision prevention, DB structure, auto-tagging, and checksum discovery heuristics.
Runs automatically every Sunday at midnight UTC.
- Link Verification: Scans the entire
distros.jsondatabase (120+ entries) for 404/dead links. - Automated Reporting: If broken links are detected, the workflow automatically opens or updates a GitHub Issue with the detailed error log, and sends a Discord alert.
This is a personal utility, but suggestions and improvements are welcome. If you find a bug in the synchronization logic, want to add support for a different cloud provider, or just want to fix a broken link in distros.json, feel free to open an issue or a pull request.
This project is licensed under the MIT License - see the LICENSE file for details.