π Extractor
Pull every email, URL, IP, date, UUID, secret β or anything a regex can match β
out of the current file and into a fresh, de-duplicated one.
---
Logs, configs, dumps and docs are full of buried structure. **Extractor** scrapes
it out in one command: choose *what* to pull, and it collects every match from the
file (or your selection), removes duplicates, and opens the results in a new tab β
ready to save, sort or paste elsewhere.
## β¨ Features
- π― **22 built-in extractors** β from emails to Luhn-checked credit-card numbers
- π§© **Custom regex** β type any pattern; capture groups are honored
- π§Ή **De-duplicated output**, optionally case-insensitive and/or sorted
- βοΈ **Whole file or selection** β your choice, configurable
- π±οΈ **Everywhere** β Command Palette, Tools menu, and right-click
- π¦ **Zero dependencies** β pure Python standard library
- π **ST3 & ST4** β runs on both plugin hosts
## π Usage
1. Open the **Command Palette** (`Ctrl/Cmd + Shift + P`).
2. Run **`Extractor: Extract to new fileβ¦`**.
3. Pick what to extract. Done β the matches open in a new tab.
Prefer fewer clicks? Every extractor also has:
- a **direct** Command Palette entry, e.g. `Extractor: Emails`
- an item under **Tools β Extractor**
> π‘ By default Extractor reads the whole file, or just your selection when you
> have one. Change this under *Preferences β Package Settings β Extractor β Settings*.
### Bind a key
Any extractor can be bound directly by passing its `kind`:
```json
{ "keys": ["ctrl+alt+e"], "command": "extractor", "args": { "kind": "emails" } }
```
Omit `args` to open the picker.
## π§° What it can extract
| Extractor | Matches |
| -------------------- | --------------------------------------------------------------- |
| **Emails** | `user@host.tld` |
| **URLs** | `http`, `https`, `ftp` and `www.` links |
| **Domains** | Registrable root domains (`example.co.uk`) |
| **Hostnames** | Fully-qualified host names and `localhost` |
| **IPv4 addresses** | Dotted-quad addresses (`192.168.0.1`) |
| **IPv6 addresses** | Colon-hex addresses (`::1`, `fe80::β¦`) |
| **Ports** | Port numbers from `host:port` references |
| **Phone numbers** | International / local numbers (heuristic) |
| **Dates** | ISO, numeric and month-name dates |
| **Times** | Clock times, optional seconds and AM/PM |
| **Timestamps** | ISO 8601, syslog and Unix-epoch timestamps |
| **File paths** | Unix, Windows and UNC paths |
| **UUIDs** | RFC 4122 UUIDs / GUIDs |
| **MAC addresses** | Ethernet hardware addresses |
| **Hex colors** | `#rgb` and `#rrggbb` |
| **Numbers** | Integers and decimals (thousands-aware) |
| **Hashtags** | `#hashtag` |
| **Mentions** | `@handle` |
| **Credit cards** | 13β19 digit numbers that pass the Luhn check |
| **Secrets & API keys** | AWS / Google / GitHub / Slack keys, JWTs, PEM private keys |
| **Markdown links** | URLs inside `[text](url)` |
| **Custom regexβ¦** | Everything matching a pattern you type |
## π¦ Installation
### Package Control (recommended)
1. Open the Command Palette and run **Package Control: Install Package**.
2. Search for **Extractor** and press Enter.
### Manual install
1. In Sublime Text, open **Preferences β Browse Packagesβ¦**.
2. Create a folder named `Extractor`.
3. Copy the contents of this repository into it.
4. Restart Sublime Text, or run **Tools β Developer β Reload Plugins**.
## βοΈ Settings
`Preferences β Package Settings β Extractor β Settings`
```json
{
"unique": true, // drop duplicate matches
"case_insensitive_dedupe": true, // "Example.com" == "example.com"
"sort": false, // keep first-seen order (true = sort AβZ)
"source": "auto" // "auto" | "selection" | "file"
}
```
## β οΈ Notes on accuracy
Extraction from free-form text is inherently heuristic. A few deliberate trade-offs:
- **Domains / hostnames** are gated to a curated list of common TLDs, so
`main.py` and `README.md` are *not* mistaken for hosts. Exotic TLDs are best
captured with the **Custom regex** extractor.
- **Phone numbers** require a `+` prefix or grouping separators; bare digit runs
(which are indistinguishable from IDs or timestamps) are skipped.
- **File paths** containing spaces are captured up to the first space.
- **Numbers** is intentionally broad and will match digits inside IPs, dates, etc.
## π€ Contributing
Issues and pull requests are welcome β new extractors, better patterns, and test
cases especially.
## π License
Released under the [MIT License](LICENSE).