Episode
How GitHub Actually Migrated 100 Million Repos to a New Storage Engine
- Published
- Jun 21, 2026
- Duration seconds
- 565
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/episodes/how-github-actually-migrated-100-million-repos-to-a-new-storage-engine/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/how-github-actually-migrated-100-million-repos-to-a-new-storage-engine.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In 2024, GitHub faced an impossible problem: its 15-year-old storage backend, built on bare Git repositories, couldn't keep up with 100 million active repos, AI-generated commits, and terabyte-scale monorepos. This episode drills into how GitHub's engineering team designed and rolled out a custom storage engine called 'GitHub Storage Service' (GSS) without a single user-facing outage. We cover the fundamental shift from POSIX filesystem assumptions to object-storage-native Git, the two-year phased migration that touched every push and clone request, and the surprising performance win: 40% faster clone times for large repos. Lucas and Luna also discuss the trade-off between backward compatibility and architectural purity — and why GitHub chose to keep the Git protocol unchanged even as they ripped out the entire storage layer underneath. #GitHub #StorageEngine #Git #Infrastructure #Migration #ObjectStorage #Engineering #CTO #TechnicalLeadership #Architecture #Scalability #Monorepo #AI #BackwardCompatibility #Performance #Business #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo