Disk full and WAL growth
Symptom: free space is falling fast, No space left on device in the log, or the server PANICs on a WAL write.
Never delete files from pg_wal by hand. It corrupts the database.
Find where the space went
select pg_size_pretty(sum(size)) as wal_on_disk from pg_ls_waldir();
select datname, pg_size_pretty(pg_database_size(datname)) from pg_database order by pg_database_size(datname) desc;
select datname, temp_files, pg_size_pretty(temp_bytes) as temp from pg_stat_database;
select relname, pg_size_pretty(pg_total_relation_size(oid)) from pg_class where relkind = 'r' order by pg_total_relation_size(oid) desc limit 10;
WAL is large
| Cause | Check | Fix |
|---|---|---|
| Inactive replication slot | pg_replication_slots where active is false |
Drop it, then set max_slot_wal_keep_size |
| Archiving failing | select * from pg_stat_archiver (failed_count, last_failed_time) |
Fix archive_command and its destination. Postgres keeps WAL until it archives it |
| Very write-heavy period | WAL rate in pg_stat_wal |
Expected. Size max_wal_size and the disk for it |
Data is large
- Bloat: see Vacuum and bloat.
- Queries spilling to temp files: raise
work_memfor those queries, and setlog_temp_files = 0to see them. - Old logs, dumps or other files sharing the volume.
If the disk is already full
Free space that is not Postgres data: rotated logs, core dumps, old dumps. Drop the offending slot. Then restart, and let Postgres replay its WAL on its own.
Prevent
Alert at 70% disk usage and on pg_stat_archiver.failed_count increasing. Keep the data directory on its own volume.