The pain
You push your dotfiles to a public GitHub repo. Feels safe — "just configs, scripts, nothing secret". Half an hour later you realise:
- Your work email
[email protected]is sitting ingit logbecause some old merge commit dragged the employer domain in. - A colleague's address
[email protected]is in there too — they once sent a PR to your fork, and the fork kept their commits with the work domain. - The
gitleaks.tomlyou wrote yourself to block leaks contains a regex literal with your work domain — meaning you've documented that domain in a public repo on purpose. - Your CLAUDE.md mentions internal codenames for your home-lab server and its RFC1918 address.
Most of this isn't "secret" in the classic sense. No API key, no private key. It's an identity leak: your face, your employer, your home infrastructure, your workflow. Most off-the-shelf scanners miss it — the built-in detectors target AWS/Stripe/JWT, not "a colleague's name in attributable form".
The worst part is you thought you were being careful.
What I found in one day
I started with one task ("drop a script into dotfiles") and found seven different classes of leak:
- Work email in commit metadata for two different people (mine and a colleague's).
- Work domain in literal regex patterns inside the
gitleaksconfig — the file I wrote myself to block leaks. - Internal codenames for my home lab, mentioned as comments in docs.
- RFC1918 LAN IP of my home server in example config snippets.
- Path leak — references to local secret directories in docs as an example. Not a secret per se, but it tells anyone reading that I keep secrets at exactly that path, plaintext, no encryption layer.
- Stale local refs from an old
git fetch pull/N/head:pr-N— after a force-push to main, those branches still drag pre-rewrite commits with the work email back into the local pack. - Lockfile sha512 hashes containing a codename substring as part of a base64 blob → false positive of the new codename rule.
Each one is its own class. No single scanner catches all of them.
The fix — defense in depth
Five layers. Each catches its own class. There's no single source of truth, and there shouldn't be.
Layer 1. Git identity by directory
Default user.email in ~/.gitconfig is a public noreply:
[user]
name = your-handle
email = [email protected]Work identity activates only inside ~/projects/work/:
[includeIf "gitdir/i:~/projects/work/"]
path = ~/.gitconfig-work~/.gitconfig-work (a separate file, NOT in public dotfiles):
[user]
name = Real Name
email = [email protected]Effect: commit inside a work project → work author. Commit anywhere else → public author. No more "oh I forgot to switch identity".
Layer 2. Pre-commit gitleaks scan
Wired through core.hooksPath = ~/.config/git/hooks — globally, once per machine. Every existing and future repo runs the same hook, no per-repo install-hooks.sh.
The hook runs gitleaks protect --staged against two configs:
~/.config/gitleaks.toml— public-safe rules. Extends built-in (~150 detectors: AWS, Stripe, GCP, GitHub, Slack, OpenAI, Anthropic, etc.) plus generic custom rules: RFC1918 ranges,sk-ant-/sk-proj-/Bearer, PEM private keys,~/.claude/secrets/paths. The file lives inside dotfiles, visible to anyone who clones the repo. It contains only generic patterns — no tie to a specific work domain or codename.~/.config/gitleaks.local.toml— personal rules. Lives outside dotfiles, gitignored. Contains the work domain, the home-lab codename, the personal domain. This file is itself a leak if it ends up in a public repo, hence a dedicated gitignore entry plus an installer reminder.
Hook pseudo-code:
gitleaks protect --staged --config ~/.config/gitleaks.toml # public rules
[ -f ~/.config/gitleaks.local.toml ] && \
gitleaks protect --staged --config ~/.config/gitleaks.local.toml # personal rulesPer-repo opt-out: git config hooks.skipLeaks true (e.g. inside a work repo where the work email in every commit is the norm).
Layer 3. Pre-push semantic review via Claude (opt-in)
Regex doesn't catch contextual leaks: "client X said …", "Project X milestone", "internal API call shape". That's semantic — keyword-based regex breaks on false positives.
The pre-push hook pipes the outgoing diff through claude --print --model haiku with a strict prompt — looking for NDA names, client codenames, real-name leaks, business logic. Cheap (haiku ~$0.0001 per push), fast (~1-3s).
Opt-in per repo: git config hooks.claudeReview true. Otherwise it doesn't run.
Layer 4. Clean up historical leaks (one-time, painful)
git filter-repo — the modern replacement for filter-branch. Run it with --mailmap or --email-callback to rewrite email metadata, and with --message-callback to strip unwanted lines from commit messages.
Mailmap example:
your-handle <[email protected]> <[email protected]>Then git push --force --all + --tags. After that, you must:
- Re-add origin —
git filter-repostrips theoriginremote by default ("safety"); add it back withgit remote add origin <url>. - Delete stale remote-tracking refs —
git update-ref -d refs/remotes/origin/<stale-branch>for every branch that's gone from the remote.git fetch --prunedoesn't remove them if the branch had already been deleted earlier. - GC —
git reflog expire --expire=now --all && git gc --prune=now --aggressive. Without this, orphan commits linger locally and re-enter your tree on the next merge / cherry-pick. - Force-push every branch and tag, not just default.
And don't forget: after a force-push, GitHub keeps resolving the old SHA via direct URL https://github.com/<user>/<repo>/commit/<old-sha> for ~30-90 days until cache GC. Full wipe = delete and recreate the repo.
Layer 5. Personal config files outside dotfiles
~/.config/gitleaks.local.toml, ~/.gitconfig-work, ~/.claude/secrets/*.env — never in public dotfiles. Private storage only. The installer (install.sh) on first run scaffolds a template from a public .example file into the right location with an explicit warning:
⚠ ACTION REQUIRED: edit your personal gitleaks rules
~/.config/gitleaks.local.toml
...
To open now:
${EDITOR:-vi} ~/.config/gitleaks.local.tomlWithout that explicit warning, future-you forgets the step and pushes the next commit without the personal guard. Learned by stepping on the rake.
Gotchas (we stepped on them)
filter-reporemoves theoriginremote. Looks like a bug, but it's a safety feature (so you can't accidentally push pre-rewrite state upstream). Re-add manually.- Stale remote-tracking refs aren't cleared by
git fetch --prune. If you've fetchedpull/N/head:pr-N, that ref stays around and drags pre-rewrite metadata. The only way out isgit update-ref -d refs/remotes/origin/pr-N. gitleaks[allowlist]syntax depends on scope. A top-level[allowlist]applies to every rule. A per-rule allowlist is written as[rules.allowlist]immediately after the[[rules]]block — not[[rules.allowlist]](the array syntax dies with a decoder error).- Lockfile sha512 hashes contain base64 blobs of arbitrary letters. A 3-character codename will almost certainly collide with some hash of an npm transitive. Fix: add lockfile paths to the allowlist of each short-codename rule.
- Short-codename collision with a real word. If your codename overlaps with a real word (or its substring), an allowlist can cover compound forms but not the standalone word. Accept the false positive in prose vs. a missed leak in code. Right trade-off, but worth documenting on the rule itself for the next commit author.
- GitHub contributor cache recomputes a few minutes after a push, but full removal from the contributor page after a force-push can take up to an hour. Give it time to settle.
- Force-push doesn't GC remote SHAs. Old commits stay reachable by direct URL for ~30-90 days. If you need an immediate wipe — delete and recreate the repo.
What's measurable
After a full pass (one day, 19 of my repos):
- 0 mentions of the work email in commit metadata of any repo.
- 0 mentions of the work domain in history blobs (verified via
gitleaks detect). - 0 internal codenames across all 22 of my repos. The only hits are inside npm sha512 hashes in lockfiles, allowlisted by path.
- 2 gitleaks configs loaded on every commit — public-safe in dotfiles + personal local outside dotfiles.
- 2 git identities auto-switching by
gitdir— work desk vs. everything else.
Takeaway
Public commits live forever. git push --force hides reachable refs from the default view, but GitHub's SHA cache keeps the old objects for ~90 more days. Once something is public, it's already an archive for anyone who cloned, forked, or got it crawled into a third-party archive. Sanitize before the push — it's not paranoia, it's the only effective moment.
Identity leaks are the cheapest class of leak to miss. Nobody planted an API key for you in plaintext. But:
- your employer in
git log(one old merge commit, once), - a codename in a comment (one doc you wrote yourself),
- a LAN IP in a README (one example that seemed harmless).
Each one is small. Together they're a precise picture of where you work, what infra you run at home, and what your workflow looks like.
The defense is layers, not point rules:
- Identity by directory — two contexts, not one global.
- Pre-commit scan with two configs (public + personal local outside dotfiles).
- Pre-push semantic review for contextual leaks.
- Historical cleanup via
filter-repoplus careful ref management. - Personal config outside dotfiles — never the actual rules in a public repo.
Any single layer slips, another catches. And don't forget that ${EDITOR:-vi} reminder in your installer — otherwise future-you forgets the first step and ships the next commit without the guard.
P.S. — the most reliable layer
Honestly, everything above is compensation for mixing personal and work on the same machine in the first place. The most reliable defense is separate devices: a work laptop for work repos, work email, corporate VPN, NDA-everything; a personal one for your own repos, dotfiles, experiments, AI tooling.
Physical separation:
- No shared
~/.gitconfigto juggle withincludeIf— each device has its own. - No shared
~/.ssh/and~/.aws/where you might accidentally sign a work commit with a personal key, or the reverse. - No shared clipboard / pasteboard / cloud sync funnelling a work-domain snippet into your personal fish config.
- No Claude Code / Codex / Copilot with access to the work codebase and your personal repos in the same
recent filesat the same time. - Browser, password manager, messengers — separate profiles or separate apps.
Middle ground: separate OS users on one machine
If two devices is too expensive (new laptop + second monitor / mic / keyboard × NaN), there's a middle path: one MacBook, two system users. A personal user and a work user. On macOS — Fast User Switching.
What you get:
- Separate
$HOMEper user → separate~/.gitconfig,~/.ssh/,~/.aws/,~/.claude/,~/.codex/,~/Library/Application Support/. No bleed. - Separate Keychain — passwords, tokens, SSH keys isolated.
- Separate browser sessions (not profiles inside one Chrome — actually different user-level instances).
- Separate clipboard / Raycast clipboard history.
- FileVault + automatic lock on switch → if someone gets access to one user, the other stays behind the encryption wall.
What you give up:
- Switching is a pain. Fast User Switch isn't instant, and every time you have to remember which user you are right now. After a couple of weeks you get tired and start "alright, I'll do this tiny task under the personal user, switching is too much friction" — and that's exactly where the whole guard collapses.
- Double resources — two Docker Desktop instances, two IDE setups, two sets of extensions. Painful on a 16 GB RAM machine.
- Software licenses are often per-user (Spotify, JetBrains, etc.).
- Resume/sleep state sometimes drops at switch.
This is one step up from a single user with discipline, and one step down from two physical devices. Works well when:
- you do work part-time or temporarily (a full work-laptop is overkill),
- you want strict data isolation without buying hardware,
- you have a beefy machine (32 GB+ RAM, M-series Apple Silicon) that can carry two user contexts at once.
The five layers above exist because I didn't do this in either sense — neither two-device nor two-user-account. Just one user, one machine, everything together. The layers work, but they're expensive: time, attention, periodic audits, edge cases.
Tldr — the spectrum
- One user + discipline (the five layers above) — cheap, demands discipline and periodic audits.
- Two OS users on one machine — strong data isolation, you pay in UX and switching friction.
- Two physical devices — gold standard, but you double your capex and your peripherals.
Higher on the scale → fewer edge cases and less audit time, but more money / friction. Lower → cheap and convenient, but one forgotten git config user.email and your employer ends up in git log.
If you're standing at that fork — climb up the scale. Future-you will thank you. xD
