Replication lag and slots
Symptom: a replica serves stale data, failover would lose data, or the primary's disk fills because it is retaining WAL for a consumer.
Confirm (run on the primary)
select application_name, state, sync_state,
pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) as replay_lag_bytes,
write_lag, flush_lag, replay_lag
from pg_stat_replication;
select slot_name, slot_type, active, wal_status,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) as retained
from pg_replication_slots;
Which lag is large tells you where the delay is: write is the network, flush is the replica's disk, replay is the replica applying changes.
Causes and fixes
| Cause | Fix |
|---|---|
| Replica disk or CPU too slow | Match the primary's I/O capacity |
| Long queries on the replica conflicting with replay | Lower max_standby_streaming_delay, or accept hot_standby_feedback = on (which can bloat the primary) |
| One huge transaction | Break bulk writes into batches |
| Network bandwidth or latency | Check throughput between the hosts |
| Consumer gone but slot remains | Drop the slot (below) |
Inactive slots
A slot whose consumer is gone keeps every WAL file forever, until the disk fills. Confirm the consumer is really gone, then:
select pg_drop_replication_slot('slot_name');
Prevent
Set max_slot_wal_keep_size (PostgreSQL 13+) so a dead slot can never take the disk. Alert on wal_status = 'lost' and on retained bytes per slot.
On the replica
Connect to the replica itself to see its side of the link:
select pg_is_in_recovery(); -- true on a standby
select status, last_msg_receipt_time from pg_stat_wal_receiver; -- no row, or a status other than 'streaming', means it isn't receiving
select pg_wal_lsn_diff(pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn()) as received_not_applied;
select pg_is_wal_replay_paused();
select * from pg_stat_database_conflicts where datname = current_database(); -- queries cancelled by replay
- Judge lag by bytes, not only seconds.
pg_last_xact_replay_timestamp()keeps growing when the primary is idle, and the time-based lag columns on the primary can stall while replay is blocked by a recovery conflict. Bytes received but not applied is the honest measure. - No receiver, but
primary_conninfois set: replication is lost. Read the replica's log for why. - "requested WAL segment has already been removed": the primary discarded WAL the replica still needed. It can't catch up: rebuild it (
pg_basebackup -h primary -D /var/lib/postgresql/17/main -R -X stream), and prevent a repeat with a replication slot capped bymax_slot_wal_keep_size, a largerwal_keep_size, or arestore_commandfrom your WAL archive. - Queries cancelled "due to conflict with recovery": replay needed rows or locks a long query held. Either keep replica queries short, raise
max_standby_streaming_delay(the replica lags more), or turn onhot_standby_feedback(the primary can bloat). - Synchronous replication and commits hang: if
synchronous_standby_namesis set and that standby is gone, every commit waits. Reconnect it, or clear the setting temporarily:alter system set synchronous_standby_names = ''; select pg_reload_conf();.
The logs show what statistics can't: when it dropped and how often it reconnected. pgBadger summarizes them.