● LIVE· № 001 · SANITIZE EVERYTHING THAT HITS GIT: CLEANING PUBLIC REPOS FROM IDENTITY LEAKS · 2026.05.11· № 002 · IMAGEGEN-MCP: A HOMEGROWN MCP SERVER FOR BLOG COVERS · 2026.05.11· № 003 · SMART PASTE: STRIPPING TERMINAL NOISE BEFORE PASTING, WITH ONE HOTKEY · 2026.05.10· № 004 · CLAUDE CODE TEAM TELEMETRY: CENTRALIZED USAGE STATS · 2026.05.07· № 005 · CLAUDE CODE ONBOARDING GUIDE FOR NEWCOMERS · 2026.05.06· 11 POSTS · 0 DRAFTS
EN / RU
·10 MIN

Sanitize everything that hits git: cleaning public repos from identity leaks

Identity leaks are the cheapest class of leak to miss. Your employer in git log, a codename in a comment, a LAN IP in a README. One day of hunting my own traces and five defense layers so it doesn't happen again.

The pain

You push your dotfiles to a public GitHub repo. Feels safe — "just configs, scripts, nothing secret". Half an hour later you realise:

  • Your work email [email protected] is sitting in git log because some old merge commit dragged the employer domain in.
  • A colleague's address [email protected] is in there too — they once sent a PR to your fork, and the fork kept their commits with the work domain.
  • The gitleaks.toml you wrote yourself to block leaks contains a regex literal with your work domain — meaning you've documented that domain in a public repo on purpose.
  • Your CLAUDE.md mentions internal codenames for your home-lab server and its RFC1918 address.

Most of this isn't "secret" in the classic sense. No API key, no private key. It's an identity leak: your face, your employer, your home infrastructure, your workflow. Most off-the-shelf scanners miss it — the built-in detectors target AWS/Stripe/JWT, not "a colleague's name in attributable form".

The worst part is you thought you were being careful.

What I found in one day

I started with one task ("drop a script into dotfiles") and found seven different classes of leak:

  1. Work email in commit metadata for two different people (mine and a colleague's).
  2. Work domain in literal regex patterns inside the gitleaks config — the file I wrote myself to block leaks.
  3. Internal codenames for my home lab, mentioned as comments in docs.
  4. RFC1918 LAN IP of my home server in example config snippets.
  5. Path leak — references to local secret directories in docs as an example. Not a secret per se, but it tells anyone reading that I keep secrets at exactly that path, plaintext, no encryption layer.
  6. Stale local refs from an old git fetch pull/N/head:pr-N — after a force-push to main, those branches still drag pre-rewrite commits with the work email back into the local pack.
  7. Lockfile sha512 hashes containing a codename substring as part of a base64 blob → false positive of the new codename rule.

Each one is its own class. No single scanner catches all of them.

The fix — defense in depth

Five layers. Each catches its own class. There's no single source of truth, and there shouldn't be.

Layer 1. Git identity by directory

Default user.email in ~/.gitconfig is a public noreply:

[user]
  name = your-handle
  email = [email protected]

Work identity activates only inside ~/projects/work/:

[includeIf "gitdir/i:~/projects/work/"]
  path = ~/.gitconfig-work

~/.gitconfig-work (a separate file, NOT in public dotfiles):

[user]
  name = Real Name
  email = [email protected]

Effect: commit inside a work project → work author. Commit anywhere else → public author. No more "oh I forgot to switch identity".

Layer 2. Pre-commit gitleaks scan

Wired through core.hooksPath = ~/.config/git/hooks — globally, once per machine. Every existing and future repo runs the same hook, no per-repo install-hooks.sh.

The hook runs gitleaks protect --staged against two configs:

  • ~/.config/gitleaks.toml — public-safe rules. Extends built-in (~150 detectors: AWS, Stripe, GCP, GitHub, Slack, OpenAI, Anthropic, etc.) plus generic custom rules: RFC1918 ranges, sk-ant- / sk-proj- / Bearer, PEM private keys, ~/.claude/secrets/ paths. The file lives inside dotfiles, visible to anyone who clones the repo. It contains only generic patterns — no tie to a specific work domain or codename.
  • ~/.config/gitleaks.local.toml — personal rules. Lives outside dotfiles, gitignored. Contains the work domain, the home-lab codename, the personal domain. This file is itself a leak if it ends up in a public repo, hence a dedicated gitignore entry plus an installer reminder.

Hook pseudo-code:

gitleaks protect --staged --config ~/.config/gitleaks.toml         # public rules
[ -f ~/.config/gitleaks.local.toml ] && \
  gitleaks protect --staged --config ~/.config/gitleaks.local.toml # personal rules

Per-repo opt-out: git config hooks.skipLeaks true (e.g. inside a work repo where the work email in every commit is the norm).

Layer 3. Pre-push semantic review via Claude (opt-in)

Regex doesn't catch contextual leaks: "client X said …", "Project X milestone", "internal API call shape". That's semantic — keyword-based regex breaks on false positives.

The pre-push hook pipes the outgoing diff through claude --print --model haiku with a strict prompt — looking for NDA names, client codenames, real-name leaks, business logic. Cheap (haiku ~$0.0001 per push), fast (~1-3s).

Opt-in per repo: git config hooks.claudeReview true. Otherwise it doesn't run.

Layer 4. Clean up historical leaks (one-time, painful)

git filter-repo — the modern replacement for filter-branch. Run it with --mailmap or --email-callback to rewrite email metadata, and with --message-callback to strip unwanted lines from commit messages.

Mailmap example:

your-handle <[email protected]> <[email protected]>

Then git push --force --all + --tags. After that, you must:

  1. Re-add origin — git filter-repo strips the origin remote by default ("safety"); add it back with git remote add origin <url>.
  2. Delete stale remote-tracking refs — git update-ref -d refs/remotes/origin/<stale-branch> for every branch that's gone from the remote. git fetch --prune doesn't remove them if the branch had already been deleted earlier.
  3. GC — git reflog expire --expire=now --all && git gc --prune=now --aggressive. Without this, orphan commits linger locally and re-enter your tree on the next merge / cherry-pick.
  4. Force-push every branch and tag, not just default.

And don't forget: after a force-push, GitHub keeps resolving the old SHA via direct URL https://github.com/<user>/<repo>/commit/<old-sha> for ~30-90 days until cache GC. Full wipe = delete and recreate the repo.

Layer 5. Personal config files outside dotfiles

~/.config/gitleaks.local.toml, ~/.gitconfig-work, ~/.claude/secrets/*.env — never in public dotfiles. Private storage only. The installer (install.sh) on first run scaffolds a template from a public .example file into the right location with an explicit warning:

⚠  ACTION REQUIRED: edit your personal gitleaks rules
   ~/.config/gitleaks.local.toml
   ...
   To open now:
       ${EDITOR:-vi} ~/.config/gitleaks.local.toml

Without that explicit warning, future-you forgets the step and pushes the next commit without the personal guard. Learned by stepping on the rake.

Gotchas (we stepped on them)

  • filter-repo removes the origin remote. Looks like a bug, but it's a safety feature (so you can't accidentally push pre-rewrite state upstream). Re-add manually.
  • Stale remote-tracking refs aren't cleared by git fetch --prune. If you've fetched pull/N/head:pr-N, that ref stays around and drags pre-rewrite metadata. The only way out is git update-ref -d refs/remotes/origin/pr-N.
  • gitleaks [allowlist] syntax depends on scope. A top-level [allowlist] applies to every rule. A per-rule allowlist is written as [rules.allowlist] immediately after the [[rules]] block — not [[rules.allowlist]] (the array syntax dies with a decoder error).
  • Lockfile sha512 hashes contain base64 blobs of arbitrary letters. A 3-character codename will almost certainly collide with some hash of an npm transitive. Fix: add lockfile paths to the allowlist of each short-codename rule.
  • Short-codename collision with a real word. If your codename overlaps with a real word (or its substring), an allowlist can cover compound forms but not the standalone word. Accept the false positive in prose vs. a missed leak in code. Right trade-off, but worth documenting on the rule itself for the next commit author.
  • GitHub contributor cache recomputes a few minutes after a push, but full removal from the contributor page after a force-push can take up to an hour. Give it time to settle.
  • Force-push doesn't GC remote SHAs. Old commits stay reachable by direct URL for ~30-90 days. If you need an immediate wipe — delete and recreate the repo.

What's measurable

After a full pass (one day, 19 of my repos):

  • 0 mentions of the work email in commit metadata of any repo.
  • 0 mentions of the work domain in history blobs (verified via gitleaks detect).
  • 0 internal codenames across all 22 of my repos. The only hits are inside npm sha512 hashes in lockfiles, allowlisted by path.
  • 2 gitleaks configs loaded on every commit — public-safe in dotfiles + personal local outside dotfiles.
  • 2 git identities auto-switching by gitdir — work desk vs. everything else.

Takeaway

Public commits live forever. git push --force hides reachable refs from the default view, but GitHub's SHA cache keeps the old objects for ~90 more days. Once something is public, it's already an archive for anyone who cloned, forked, or got it crawled into a third-party archive. Sanitize before the push — it's not paranoia, it's the only effective moment.

Identity leaks are the cheapest class of leak to miss. Nobody planted an API key for you in plaintext. But:

  • your employer in git log (one old merge commit, once),
  • a codename in a comment (one doc you wrote yourself),
  • a LAN IP in a README (one example that seemed harmless).

Each one is small. Together they're a precise picture of where you work, what infra you run at home, and what your workflow looks like.

The defense is layers, not point rules:

  1. Identity by directory — two contexts, not one global.
  2. Pre-commit scan with two configs (public + personal local outside dotfiles).
  3. Pre-push semantic review for contextual leaks.
  4. Historical cleanup via filter-repo plus careful ref management.
  5. Personal config outside dotfiles — never the actual rules in a public repo.

Any single layer slips, another catches. And don't forget that ${EDITOR:-vi} reminder in your installer — otherwise future-you forgets the first step and ships the next commit without the guard.

P.S. — the most reliable layer

Honestly, everything above is compensation for mixing personal and work on the same machine in the first place. The most reliable defense is separate devices: a work laptop for work repos, work email, corporate VPN, NDA-everything; a personal one for your own repos, dotfiles, experiments, AI tooling.

Physical separation:

  • No shared ~/.gitconfig to juggle with includeIf — each device has its own.
  • No shared ~/.ssh/ and ~/.aws/ where you might accidentally sign a work commit with a personal key, or the reverse.
  • No shared clipboard / pasteboard / cloud sync funnelling a work-domain snippet into your personal fish config.
  • No Claude Code / Codex / Copilot with access to the work codebase and your personal repos in the same recent files at the same time.
  • Browser, password manager, messengers — separate profiles or separate apps.

Middle ground: separate OS users on one machine

If two devices is too expensive (new laptop + second monitor / mic / keyboard × NaN), there's a middle path: one MacBook, two system users. A personal user and a work user. On macOS — Fast User Switching.

What you get:

  • Separate $HOME per user → separate ~/.gitconfig, ~/.ssh/, ~/.aws/, ~/.claude/, ~/.codex/, ~/Library/Application Support/. No bleed.
  • Separate Keychain — passwords, tokens, SSH keys isolated.
  • Separate browser sessions (not profiles inside one Chrome — actually different user-level instances).
  • Separate clipboard / Raycast clipboard history.
  • FileVault + automatic lock on switch → if someone gets access to one user, the other stays behind the encryption wall.

What you give up:

  • Switching is a pain. Fast User Switch isn't instant, and every time you have to remember which user you are right now. After a couple of weeks you get tired and start "alright, I'll do this tiny task under the personal user, switching is too much friction" — and that's exactly where the whole guard collapses.
  • Double resources — two Docker Desktop instances, two IDE setups, two sets of extensions. Painful on a 16 GB RAM machine.
  • Software licenses are often per-user (Spotify, JetBrains, etc.).
  • Resume/sleep state sometimes drops at switch.

This is one step up from a single user with discipline, and one step down from two physical devices. Works well when:

  • you do work part-time or temporarily (a full work-laptop is overkill),
  • you want strict data isolation without buying hardware,
  • you have a beefy machine (32 GB+ RAM, M-series Apple Silicon) that can carry two user contexts at once.

The five layers above exist because I didn't do this in either sense — neither two-device nor two-user-account. Just one user, one machine, everything together. The layers work, but they're expensive: time, attention, periodic audits, edge cases.

Tldr — the spectrum

  1. One user + discipline (the five layers above) — cheap, demands discipline and periodic audits.
  2. Two OS users on one machine — strong data isolation, you pay in UX and switching friction.
  3. Two physical devices — gold standard, but you double your capex and your peripherals.

Higher on the scale → fewer edge cases and less audit time, but more money / friction. Lower → cheap and convenient, but one forgotten git config user.email and your employer ends up in git log.

If you're standing at that fork — climb up the scale. Future-you will thank you. xD