What makes a database fast or slow?
In this video, we take a standard database running at 300 transactions per second and push it past one million TPS. By breaking down low-level storage mechanics, system calls, and disk synchronization, we identify the bottleneck at every stage and optimize our way around it.
Chapters & Timestamps
- 00:00 — Introduction: can we hit 1M TPS?
- 00:55 — Act 1: Anatomy of a write
- 01:17 — Benchmark: baseline (338 TPS)
- 02:00 — Memory vs. disk (RAM volatility)
- 02:25 — The
write()syscall & OS page cache - 02:49 —
fsync()& ACID durability - 03:17 — Measuring fsync latency (
ddtest) - 03:51 — Database pages & corruption risk
- 04:26 — The rollback journal
- 05:19 — The 4
fsyncs bottleneck - 06:14 — Act 2: Write-ahead log (WAL)
- 06:47 — How WAL mode works
- 07:37 — 1
fsyncper transaction - 07:47 — Checkpointing &
wal_autocheckpoint - 08:25 — Benchmark: WAL mode (1,100 TPS)
- 08:53 — The 10x–20x speedup misconception
- 09:12 — Act 3: Synchronous settings & trade-offs
- 09:40 — Benchmark:
synchronous = OFF(100k TPS) - 10:28 — Benchmark:
synchronous = NORMAL(12k TPS) - 11:30 — Why ORMs set NORMAL by default
- 11:50 — Act 4: The narrow bridge & bus analogy
- 12:42 — Group commits & batching windows
- 13:34 — Benchmark: first batching test (5,500 TPS)
- 13:50 — Scaling batch size to 28k
- 14:50 — Benchmark: 1,000,000 TPS reached!
- 15:11 — Amdahl’s law & the CPU bottleneck
- 16:17 — Time budget: disk I/O vs CPU
- 16:38 — CPU clock cycles: 5,300 cycles per tx
- 17:20 — Outro: 1M TPS achieved
- 18:03 — Real-world systems & takeaways