Git in Monorepos: Patterns That Scale When Your Repository Grows Beyond Comfortable¶
Lessons from managing a repository with 50 services, 200 developers, and a git history that could crush naive tooling.
When One Repo Is Better Than Fifty¶
I've worked on both sides: multi-repo microservices where a simple interface change required coordinating five pull requests, and monorepos where everything is one git clone away. The monorepo won for our team — atomic cross-service changes, single CI pipeline, shared tooling. But git wasn't designed for repositories this large. I had to learn patterns that make it work.
Sparse Checkout: Clone Less, Work Faster¶
In a monorepo with hundreds of services, cloning everything is wasteful. Sparse checkout lets you materialize only the directories you need:
# Clone with sparse checkout enabled
git clone --sparse https://github.com/org/monorepo.git
cd monorepo
# Add only the directories you work on
git sparse-checkout set services/auth shared/config tools/build
# Your working directory now only contains those paths
# Everything else exists in git history but isn't checked out
Adding or removing paths is instant — no network required:
# Add another service later
git sparse-checkout add services/payment
# List current sparse paths
git sparse-checkout list
I keep a team-specific sparse config file that new developers apply on day one:
# Clone and apply team-specific sparse config
git clone --sparse git@github.com:org/monorepo.git
cd monorepo
git sparse-checkout set $(cat .sparse-configs/team-platform.txt)
Partial Clone: Skip the Blob Download¶
Even with sparse checkout, git clone downloads all objects (commits, trees, blobs) for the entire history. Partial clone changes this:
# Clone without downloading file contents (blobs)
git clone --filter=blob:none https://github.com/org/monorepo.git
# Blobs are fetched on-demand when you checkout a file
# History browsing (git log) works without downloading everything
# Combine with sparse checkout for maximum efficiency
git clone --filter=blob:none --sparse https://github.com/org/monorepo.git
This reduced our clone time from 25 minutes to under 2 minutes. Developers fetch blobs lazily as they checkout files they actually need.
Shallow Clone for CI¶
CI jobs rarely need full history. Shallow clone fetches only recent commits:
# Fetch only the last 10 commits
git clone --depth=10 https://github.com/org/monorepo.git
# Fetch single branch, single commit (fastest possible)
git clone --depth=1 --single-branch --branch main https://github.com/org/monorepo.git
I use --depth=1 for build jobs and --depth=50 for jobs that need git diff or git log context.
Path-Based Ownership With CODEOWNERS¶
In a monorepo, who reviews what? The CODEOWNERS file maps paths to teams:
# .github/CODEOWNERS
/services/auth/ @team-identity
/services/payment/ @team-billing
/shared/config/ @team-platform
/tools/ @team-devex
*.proto @team-platform @team-api
Combined with branch protection rules, this ensures the right people review changes to their domain — even when someone makes a cross-cutting change.
Commit Scoping: Making History Navigable¶
With 200 developers committing to one repo, the commit log is noisy. I enforce scoped commit messages:
auth: add OAuth2 refresh token rotation
payment: fix decimal rounding in invoice calculation
config: migrate to YAML-based service discovery
build: upgrade Gradle to 8.5
This makes git log filterable:
# See only auth service commits
git log --oneline -- services/auth/
# Or filter by commit prefix
git log --oneline --grep="^auth:"
# Changes to shared code affecting my service
git log --oneline -- shared/ services/auth/
Efficient Diffing: Know What Changed¶
In CI pipelines, you need to know which services were affected by a commit to decide what to build and test:
# Files changed between last two commits
git diff --name-only HEAD~1 HEAD
# Files changed in a pull request (vs. main)
git diff --name-only origin/main...HEAD
# Extract affected service directories
git diff --name-only origin/main...HEAD | \
grep "^services/" | \
cut -d'/' -f2 | \
sort -u
I built a CI script around this pattern — only services with actual changes get built and tested. This cut our CI time from 45 minutes to 8 minutes on average.
Git LFS for Large Files¶
Monorepos often accumulate binary assets — images, compiled dependencies, test fixtures. Git LFS keeps the repo lean:
# Track large file types with LFS
git lfs track "*.jar"
git lfs track "*.whl"
git lfs track "services/*/testdata/*.bin"
# Check what LFS is tracking
git lfs ls-files
Without LFS, every clone downloads every version of every binary ever committed. With LFS, only the current version is fetched, and old versions live on the LFS server.
Branch Strategy for Monorepos¶
I've settled on trunk-based development for monorepos:
- main is always deployable
- Feature branches are short-lived (< 3 days)
- No long-lived release branches per service
- Tags mark release points:
auth/v2.3.1,payment/v1.8.0
# Tag a service release
git tag -a auth/v2.3.1 -m "Auth: add refresh token rotation"
git push origin auth/v2.3.1
# Find all releases for a service
git tag -l "auth/*" --sort=-version:refname
Key Takeaway¶
Git scales to monorepos, but only if you use the right patterns: sparse checkout for focused working directories, partial clone for fast initial setup, path-based ownership for governance, and scoped commits for navigable history. The tooling investment pays for itself in developer velocity — one repo means one source of truth, atomic changes, and no cross-repo coordination overhead.
Tags: git, monorepo, scale, sparse-checkout, performance