pg-checkup
Workspace
HomeReports
Learn
Start hereToo many connectionsSlow queriesLocks and blocked queriesVacuum and bloatTransaction ID wraparoundReplication lag and slotsDisk full and WAL growthOut of memoryHigh CPU and I/OStale connectionsWhen to scale
Help
File a ticketRequest a featurePricing
PrivacyTerms
Docs/Replication lag and slots
Sign inNew checkupC

Replication lag and slots

Symptom: a replica serves stale data, failover would lose data, or the primary's disk fills because it is retaining WAL for a consumer.

Confirm (run on the primary)

select application_name, state, sync_state,
       pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) as replay_lag_bytes,
       write_lag, flush_lag, replay_lag
from pg_stat_replication;

select slot_name, slot_type, active, wal_status,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) as retained
from pg_replication_slots;

Which lag is large tells you where the delay is: write is the network, flush is the replica's disk, replay is the replica applying changes.

Causes and fixes

Cause Fix
Replica disk or CPU too slow Match the primary's I/O capacity
Long queries on the replica conflicting with replay Lower max_standby_streaming_delay, or accept hot_standby_feedback = on (which can bloat the primary)
One huge transaction Break bulk writes into batches
Network bandwidth or latency Check throughput between the hosts
Consumer gone but slot remains Drop the slot (below)

Inactive slots

A slot whose consumer is gone keeps every WAL file forever, until the disk fills. Confirm the consumer is really gone, then:

select pg_drop_replication_slot('slot_name');

Prevent

Set max_slot_wal_keep_size (PostgreSQL 13+) so a dead slot can never take the disk. Alert on wal_status = 'lost' and on retained bytes per slot.

On the replica

Connect to the replica itself to see its side of the link:

select pg_is_in_recovery();                       -- true on a standby
select status, last_msg_receipt_time from pg_stat_wal_receiver;   -- no row, or a status other than 'streaming', means it isn't receiving
select pg_wal_lsn_diff(pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn()) as received_not_applied;
select pg_is_wal_replay_paused();
select * from pg_stat_database_conflicts where datname = current_database();   -- queries cancelled by replay
  • Judge lag by bytes, not only seconds. pg_last_xact_replay_timestamp() keeps growing when the primary is idle, and the time-based lag columns on the primary can stall while replay is blocked by a recovery conflict. Bytes received but not applied is the honest measure.
  • No receiver, but primary_conninfo is set: replication is lost. Read the replica's log for why.
  • "requested WAL segment has already been removed": the primary discarded WAL the replica still needed. It can't catch up: rebuild it (pg_basebackup -h primary -D /var/lib/postgresql/17/main -R -X stream), and prevent a repeat with a replication slot capped by max_slot_wal_keep_size, a larger wal_keep_size, or a restore_command from your WAL archive.
  • Queries cancelled "due to conflict with recovery": replay needed rows or locks a long query held. Either keep replica queries short, raise max_standby_streaming_delay (the replica lags more), or turn on hot_standby_feedback (the primary can bloat).
  • Synchronous replication and commits hang: if synchronous_standby_names is set and that standby is gone, every commit waits. Reconnect it, or clear the setting temporarily: alter system set synchronous_standby_names = ''; select pg_reload_conf();.

The logs show what statistics can't: when it dropped and how often it reconnected. pgBadger summarizes them.

PreviousTransaction ID wraparoundNextDisk full and WAL growth