[](https://pypi.org/project/mail-deduplicate)
[](https://pypi.org/project/mail-deduplicate)
[](https://pepy.tech/projects/mail-deduplicate)
[](https://github.com/kdeldycke/mail-deduplicate/actions/workflows/tests.yaml?query=branch%3Amain)
[](https://app.codecov.io/gh/kdeldycke/mail-deduplicate)
[](https://github.com/kdeldycke/mail-deduplicate/actions/workflows/docs.yaml?query=branch%3Amain)
[](https://doi.org/10.5281/zenodo.7364256)
**What is Mail Deduplicate?**
Provides the `mdedup` CLI, an utility to deduplicate mails from a set of boxes.
## Features
- Duplicate detection based on cherry-picked and normalized mail headers.
- Fetch mails from multiple sources.
- Reads and writes to `mbox`, `maildir`, `babyl`, `mh`, `mmdf` and `eml` formats.
- Deduplication strategies based on size, timestamp, file path or random choice, chainable as fallbacks.
- Copy, move, delete or hardlink the resulting set of duplicates.
- Dry-run mode.
- Protection against false-positives with safety checks on size and content differences.
- Optional cache reusing the hashes computed by previous runs.
- Parallel hashing and selection over several processes.
- Supports macOS, Linux and Windows.
- [Standalone executables](#executables) for Linux, macOS and Windows.
- Shell auto-completion for Bash, Zsh and Fish.
## Installation
All [installation methods](https://kdeldycke.github.io/mail-deduplicate/install.html) are available in the documentation. Below are the most popular ones:
### Try it now
[`uv`](https://docs.astral.sh/uv/getting-started/installation/) is the fastest way to run `mdedup` on any platform, thanks to its [`uvx` command](https://docs.astral.sh/uv/guides/tools/#running-tools):
```shell-session
$ uvx --from mail-deduplicate -- mdedup
```
### macOS
`mdedup` is part of the official [Homebrew](https://brew.sh) default tap, so you can install it with:
```shell-session
$ brew install mail-deduplicate
```
### Executables
Standalone binaries of `mdedup`'s latest version are available as direct downloads for several platforms and architectures:
| Platform | `arm64` | `x86_64` |
| ----------- | -------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| **Linux** | [Download `mdedup-linux-arm64.bin`](https://github.com/kdeldycke/mail-deduplicate/releases/latest/download/mdedup-linux-arm64.bin) | [Download `mdedup-linux-x64.bin`](https://github.com/kdeldycke/mail-deduplicate/releases/latest/download/mdedup-linux-x64.bin) |
| **macOS** | [Download `mdedup-macos-arm64.bin`](https://github.com/kdeldycke/mail-deduplicate/releases/latest/download/mdedup-macos-arm64.bin) | [Download `mdedup-macos-x64.bin`](https://github.com/kdeldycke/mail-deduplicate/releases/latest/download/mdedup-macos-x64.bin) |
| **Windows** | [Download `mdedup-windows-arm64.exe`](https://github.com/kdeldycke/mail-deduplicate/releases/latest/download/mdedup-windows-arm64.exe) | [Download `mdedup-windows-x64.exe`](https://github.com/kdeldycke/mail-deduplicate/releases/latest/download/mdedup-windows-x64.exe) |
## Quickstart
Duplicate mails pile up whenever boxes get copied around: backups of the same account taken at different times, per-folder exports where a mail tagged with several labels lands in several boxes, archives consolidated from multiple clients, or IMAP synchronizations gone wrong.
`mdedup` cleans this up. Point it at your boxes: it groups copies of the same mail by hashing a curated set of headers, applies a `--strategy` to pick which copy to keep in each group, then performs an `--action` on the result. The default action is the safest one: sources are only read, and the deduplicated selection is written to a brand new box.
So to merge two overlapping archives into a single clean box:
```shell-session
$ mdedup --strategy select-one --export merged.mbox archive-2024.mbox archive-2025.mbox
```
Mails found in both archives are copied once into `merged.mbox`, along with all the mails that were unique to each source.
To remove duplicates in place instead, switch to a destructive action, and rehearse with `--dry-run`:
```shell-session
$ mdedup --strategy select-one --action delete-discarded --dry-run ~/Maildir
```
The [hands-on tutorial](https://kdeldycke.github.io/mail-deduplicate/tutorial.html) builds a small playground of duplicated mails, then walks through strategies, actions and safeguards on it. The [design page](https://kdeldycke.github.io/mail-deduplicate/design.html) explains how duplicates are detected.
> [!NOTE]
> Memory does not grow with the size of your boxes: each mail is reduced to a lightweight stub as soon as it is hashed, so a 215 MB maildir of 1,500 mails peaks at 46 MB. It does grow with their number, at roughly 1 KB of resident memory each, so a large enough collection can still exhaust your machine.
>
> Hashing is where a run spends most of its time, and `--cache` reuses it between runs. The [performance page](https://kdeldycke.github.io/mail-deduplicate/performance.html) covers both, and how to manage that cache.
You can support this project with pull requests, [business support 🤝](https://github.com/sponsors/kdeldycke) and [sponsorship 🫶](https://github.com/sponsors/kdeldycke).