Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to Side Projects

Side Projects

Comic Cruncher: Shrinking a Digital Comic Library and Binding Issues Into Trades

A spread of 1990s comic books fanned out across a dark wooden floor.

Part 5 of the thread Archive conversion workshop

PROJECT AT A GLANCEOpen source
What it is
Desktop app
My role
Creator
Year
2024
Built with
  • Python
  • PyQt6
  • Pillow
  • pdf2image
  • rarfile
THE SHORT VERSION4 points
  • Cruncher mode turns PDF, CBR, CBZ and CB7 comics into CBZ files with pages resized and converted to WebP.
  • Combiner mode groups single issues by series and binds them into trade paperback volumes, 12 issues each by default.
  • It uses every CPU core, skips files it's already done, and writes .backup copies before touching originals.
  • PDFs need Poppler, and CBRs need UnRAR or 7-Zip.

Digital comic collections get big, and they get messy. You end up with a mix of PDFs, CBRs and CBZs, some of them scanned at sizes no screen will ever show, and a folder with forty single issues of one series that you'd really rather read as a handful of volumes.

comic-cruncher handles both halves of that. It's a PyQt6 desktop app with two modes: one that shrinks comics down, and one that binds issues together into trade paperbacks.

Cruncher mode

Point it at an input folder, pick an output folder, and it converts whatever it finds (PDF, CBZ, CBR or CB7) into an optimized CBZ. Along the way it resizes every page to fit inside 2500 by 2500 pixels, keeping the aspect ratio, and re-encodes it as WebP.

It runs files in parallel across all your CPU cores, skips anything that's already been processed, and reports the space saved as it goes. You can drag and drop files or whole folders onto it.

The settings are all in the GUI:

Setting Default What it controls
Max Workers Your CPU count How many files run in parallel
Max Dimension 2500px Longest side of any page
WebP Quality 85 Compression quality, 1 to 100

The README's rough expectations are a 60-75% size reduction on PDFs and 50-65% on CBRs. Pages look normal at reading size, though you may see some quality loss if you zoom way in.

Tip

It writes a .backup file before modifying an original, so a bad batch isn't the end of the world. Still, run a small folder first to check you like the quality setting before you turn it loose on the whole library.

If you'd rather do the CBR-to-CBZ part from a script, I wrote up a PowerShell version that converts CBR to CBZ with optional WebP pages. Same idea, no GUI.

Combiner mode

This is the TPB Creator, added in 2.0. Give it a folder of comics from the same series and it works out the series name and issue numbers from the filenames, using a handful of patterns for common naming schemes. Then it groups them into volumes, 12 issues per volume unless you change it, and names each one properly.

There's an option to delete the original issues once a volume has been built successfully. Leave that off until you've checked a few volumes.

Using it

It needs Python 3.9 or newer, plus two outside tools depending on what you're feeding it:

  • Poppler for PDFs (on Linux, poppler-utils; on macOS, brew install poppler; on Windows, download it and add it to your PATH)
  • UnRAR or 7-Zip for CBRs

Then install the Python packages and launch the GUI:

pip install -r requirements.txt
python comic_cruncher.py

Heads up

If PDFs won't process, Poppler isn't on your PATH. If CBRs fail, it can't find UnRAR or 7-Zip. Those two account for most "it doesn't work" moments. If you run out of memory, process fewer files at once or lower the worker count.

Since CBR support leans on 7-Zip, it's worth knowing your archives are healthy before you start. My script for testing and converting a folder of mixed archives with 7-Zip is a good first pass on an old collection.

What's next

The roadmap in the changelog has a few things queued up: series-specific preferences, better pattern recognition for odd naming conventions, and an undo for recent operations. After that, a config file for persistent settings, quality profiles so you're not fiddling with the slider every time, and a command-line mode for batch runs without the GUI.

It's MIT licensed and open to contributions. If your collection names its files in a way the pattern matcher doesn't catch, that's a good place to start.