<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Eme David Onwuka]]></title><description><![CDATA[Eme David Onwuka]]></description><link>https://emedavidonwuka.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Eme David Onwuka</title><link>https://emedavidonwuka.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sat, 10 Oct 2026 05:48:56 GMT</lastBuildDate><atom:link href="https://emedavidonwuka.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[PostgreSQL Point-in-Time Recovery with pgBackRest and Patroni: WAL, Timelines, Failures, and Recovery Verification]]></title><description><![CDATA[Point-in-time recovery (PITR) is one of the PostgreSQL recovery techniques I wanted to understand beyond the documentation and theory.
For this exercise, I used a local three-node PostgreSQL 16 enviro]]></description><link>https://emedavidonwuka.hashnode.dev/postgresql-point-in-time-recovery-with-pgbackrest-and-patroni-wal-timelines-failures-and-recovery-verification</link><guid isPermaLink="true">https://emedavidonwuka.hashnode.dev/postgresql-point-in-time-recovery-with-pgbackrest-and-patroni-wal-timelines-failures-and-recovery-verification</guid><category><![CDATA[PostgreSQL]]></category><category><![CDATA[patroni]]></category><category><![CDATA[pgbackrest]]></category><category><![CDATA[Docker]]></category><category><![CDATA[Devops]]></category><category><![CDATA[Ubuntu]]></category><dc:creator><![CDATA[Eme David Onwuka]]></dc:creator><pubDate>Tue, 06 Oct 2026 11:11:51 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ac4c0714d7cbc5d02966747/7a308dab-664a-4af3-85c1-2b7cfeeaae81.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Point-in-time recovery (PITR) is one of the PostgreSQL recovery techniques I wanted to understand beyond the documentation and theory.</p>
<p>For this exercise, I used a local three-node PostgreSQL 16 environment managed by Patroni and etcd, with pgBackRest handling backups and WAL archiving. The goal was to create a known recovery point, introduce data after that point, and then recover PostgreSQL to the earlier timestamp.</p>
<p>The exercise also became a troubleshooting exercise. The restore encountered several real operational issues, including a stale <code>postmaster.pid</code>, a non-empty PostgreSQL data directory, and the interaction between pgBackRest and a Docker-mounted PostgreSQL data volume.</p>
<p>Rather than treating those errors as obstacles to hide, I documented them as part of the recovery process. The final verification showed that the data written before the recovery target remained, the data written after the target disappeared, and PostgreSQL completed recovery on a new timeline.</p>
<p>This article documents the environment, recovery procedure, failures, diagnosis, and verification.</p>
<h2>Environment</h2>
<p>The recovery exercise was performed locally on Ubuntu 24.04.4 LTS using Docker.</p>
<p>The PostgreSQL environment consisted of three PostgreSQL 16 nodes managed by Patroni, with etcd providing the distributed configuration store. pgBackRest was used for full backups and WAL archiving.</p>
<p>The relevant components were:</p>
<ul>
<li><p><strong>PostgreSQL 16.15</strong></p>
</li>
<li><p><strong>Patroni</strong></p>
</li>
<li><p><strong>etcd</strong></p>
</li>
<li><p><strong>pgBackRest 2.59.2</strong></p>
</li>
<li><p><strong>Docker</strong></p>
</li>
<li><p><strong>Ubuntu 24.04.4 LTS</strong></p>
</li>
</ul>
<p>The cluster used the following PostgreSQL nodes:</p>
<pre><code class="language-text">pg1
pg2
pg3
</code></pre>
<p>At the beginning of the recovery exercise, <code>pg1</code> was the Patroni leader. The other two nodes were configured as replicas, although they were not healthy enough to participate in the recovery procedure at that point.</p>
<p>For this exercise, the PITR operation was performed directly against <code>pg1</code>. Patroni cluster management was paused before the restore so that automated cluster management would not interfere with the recovery operation.</p>
<h2>Establishing a Recoverable State</h2>
<p>Before testing point-in-time recovery, I first established a known-good backup and verified that WAL archiving was working.</p>
<p>I ran a pgBackRest configuration check against the <code>patroni-dev291</code> stanza:</p>
<pre><code class="language-bash">docker exec patroni-pg1 pgbackrest --stanza=patroni-dev291 check
</code></pre>
<p>The check completed successfully, including verification that a WAL segment had been archived.</p>
<p>I then created a fresh full backup:</p>
<pre><code class="language-bash">docker exec patroni-pg1 pgbackrest \
  --stanza=patroni-dev291 \
  --type=full \
  backup
</code></pre>
<p>The backup completed successfully with backup label:</p>
<pre><code class="language-text">20261006-075738F
</code></pre>
<p>The backup started at <code>2026-10-06 07:57:38 UTC</code> and completed at <code>2026-10-06 07:58:39 UTC</code>.</p>
<p>At this point, I had both a validated pgBackRest configuration and a known full backup from which recovery could be performed.</p>
<h3>Creating the recovery scenario</h3>
<p>I created a small test table and inserted a row that represented data I wanted to preserve:</p>
<pre><code class="language-sql">CREATE TABLE IF NOT EXISTS pitr_demo (
    id SERIAL PRIMARY KEY,
    event TEXT NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);

INSERT INTO pitr_demo (event)
VALUES ('BEFORE_PITR');
</code></pre>
<p>The row was created at:</p>
<pre><code class="language-text">2026-10-06 08:03:06.138954+00
</code></pre>
<p>I selected the PITR target as:</p>
<pre><code class="language-text">2026-10-06 08:06:01.308827+00
</code></pre>
<p>After the target time, I inserted another row:</p>
<pre><code class="language-sql">INSERT INTO pitr_demo (event)
VALUES ('AFTER_PITR_TARGET');
</code></pre>
<p>This row was created at:</p>
<pre><code class="language-text">2026-10-06 08:29:28.494015+00
</code></pre>
<p>The intention was straightforward: after recovery to the target timestamp, <code>BEFORE_PITR</code> should remain while <code>AFTER_PITR_TARGET</code> should no longer exist.</p>
<p>To make sure the relevant WAL was archived, I also forced a WAL switch:</p>
<pre><code class="language-sql">SELECT pg_switch_wal();
</code></pre>
<p>This returned:</p>
<pre><code class="language-text">0/110023D0
</code></pre>
<p>A subsequent <code>pgbackrest info</code> check showed that the archived WAL range had advanced, confirming that WAL containing the later activity had been archived before starting recovery.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6ac4c0714d7cbc5d02966747/3bb04d94-1aaf-455f-99f2-5ae113eaa2d9.png" alt="" style="display:block;margin:0 auto" />

<h2>Performing the Point-in-Time Recovery</h2>
<p>Before starting the restore, I checked the Patroni cluster state. <code>pg1</code> was the leader on timeline 9, while <code>pg2</code> and <code>pg3</code> were not healthy enough to participate in the recovery.</p>
<p>Because the restore was going to modify the PostgreSQL data directory, I paused Patroni cluster management first:</p>
<pre><code class="language-bash">patronictl -c /etc/patroni/patroni.yml pause
</code></pre>
<p>The command returned:</p>
<pre><code class="language-text">Success: cluster management is paused
</code></pre>
<p>I then stopped the <code>pg1</code> container so that PostgreSQL would not be running while the data directory was restored.</p>
<p>The PITR restore was performed with pgBackRest using the previously created full backup and the target timestamp:</p>
<pre><code class="language-bash">docker run --rm \
  --volumes-from patroni-pg1 \
  --entrypoint pgbackrest \
  patroni-dev291-pg1 \
  --stanza=patroni-dev291 \
  --type=time \
  --target="2026-10-06 08:06:01.308827+00" \
  --target-action=promote \
  --force \
  --log-level-console=info \
  restore
</code></pre>
<p>The restore completed successfully:</p>
<pre><code class="language-text">restore command end: completed successfully
</code></pre>
<p>The <code>--target-action=promote</code> option was important because the goal was not simply to leave PostgreSQL in recovery. I wanted the recovered instance to exit recovery and become the new primary state at the requested point in time.</p>
<p>After the restore completed, I started the <code>pg1</code> container. PostgreSQL was not immediately running, so I checked its process state and started PostgreSQL manually. PostgreSQL then performed WAL replay from the restored backup toward the requested recovery target.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6ac4c0714d7cbc5d02966747/8e4fb047-94cc-45ac-b8f2-e45403ed8bd9.png" alt="" style="display:block;margin:0 auto" />

<h2>Troubleshooting the Recovery</h2>
<p>The restore was successful, but getting there was not completely straightforward. Several errors appeared during the recovery process. I treated each one as a diagnostic step rather than simply retrying commands.</p>
<h3>1. Using the container name instead of the image name</h3>
<p>My first restore attempt used <code>patroni-pg1</code> as the Docker image name. Docker could not find that image locally and attempted to resolve it as an image reference.</p>
<p>The error was:</p>
<pre><code class="language-text">Unable to find image 'patroni-pg1:latest' locally
</code></pre>
<p>I checked the actual image associated with the PostgreSQL container and found:</p>
<pre><code class="language-text">patroni-dev291-pg1
</code></pre>
<p>I then reran the restore using the correct image name.</p>
<h3>2. A stale <code>postmaster.pid</code></h3>
<p>The next restore attempt was rejected by pgBackRest:</p>
<pre><code class="language-text">ERROR [038]: unable to restore while PostgreSQL is running
HINT: presence of postmaster.pid...
</code></pre>
<p>The PostgreSQL container itself was stopped, so I inspected the data directory from a helper container and found a stale <code>postmaster.pid</code>.</p>
<p>There was no PostgreSQL process running. The PID file was left behind from the previous PostgreSQL state.</p>
<p>I removed the stale PID file and attempted the restore again.</p>
<h3>3. The PostgreSQL data directory was not empty</h3>
<p>pgBackRest then reported:</p>
<pre><code class="language-text">ERROR [040]: unable to restore to path '/var/lib/postgresql/data' because it contains files
HINT: try using --delta
</code></pre>
<p>I inspected the data directory and confirmed that it still contained the existing PostgreSQL cluster.</p>
<p>I initially considered moving the entire data directory aside, but the directory was the root of a Docker volume mount. Attempting to rename it resulted in:</p>
<pre><code class="language-text">mv: cannot move ... Device or resource busy
</code></pre>
<p>The correct approach for this recovery exercise was to let pgBackRest cleanly overwrite the existing contents. I reran the restore with <code>--force</code>, which completed successfully and removed invalid files before restoring the backup.</p>
<h3>4. PostgreSQL was not immediately available after the restore</h3>
<p>I checked the PostgreSQL process state directly:</p>
<pre><code class="language-bash">docker exec patroni-pg1 sh -c 'pg_ctl -D /var/lib/postgresql/data status'
</code></pre>
<p>The result was:</p>
<pre><code class="language-text">pg_ctl: no server running
</code></pre>
<p>This showed that the PITR restore itself had completed, but PostgreSQL was not yet running.</p>
<p>I then started PostgreSQL manually and captured its startup output:</p>
<pre><code class="language-bash">docker exec patroni-pg1 sh -c 'pg_ctl -D /var/lib/postgresql/data -l /tmp/postgres-startup.log start'
</code></pre>
<p>The server started successfully. The startup log then showed that PostgreSQL was performing the requested point-in-time recovery, including WAL replay toward the target timestamp.</p>
<p>The recovery log reported:</p>
<pre><code class="language-text">starting point-in-time recovery to 2026-10-06 08:06:01.308827+00
</code></pre>
<p>It then stopped recovery before the later transaction and reported:</p>
<pre><code class="language-text">recovery stopping before commit of transaction 746
</code></pre>
<p>The log also confirmed the timeline change:</p>
<pre><code class="language-text">selected new timeline ID: 10
</code></pre>
<p>Finally, PostgreSQL reported:</p>
<pre><code class="language-text">archive recovery complete
database system is ready to accept connections
</code></pre>
<p>This clarified that the initial connection refusal was not a failed PITR. PostgreSQL simply needed to be started so that the recovery process could run to completion.</p>
<h2>Verifying the Recovery</h2>
<p>After PostgreSQL completed recovery, I verified both the recovery state and the data that existed after the target timestamp.</p>
<p>First, I checked whether PostgreSQL was still in recovery:</p>
<pre><code class="language-sql">SELECT pg_is_in_recovery();
</code></pre>
<p>The result was:</p>
<pre><code class="language-text">f
</code></pre>
<p>This confirmed that PostgreSQL had completed recovery and was accepting normal read-write connections.</p>
<p>I then queried the test table:</p>
<pre><code class="language-sql">SELECT * FROM pitr_demo ORDER BY id;
</code></pre>
<p>The result contained only the row created before the PITR target:</p>
<pre><code class="language-text"> id |    event
----+-------------
  1 | BEFORE_PITR
</code></pre>
<p>The <code>AFTER_PITR_TARGET</code> row was no longer present.</p>
<p>This gave me the expected recovery result:</p>
<ul>
<li><p>Data committed before the target timestamp remained.</p>
</li>
<li><p>Data committed after the target timestamp was removed.</p>
</li>
<li><p>PostgreSQL was no longer in recovery.</p>
</li>
<li><p>PostgreSQL had created a new timeline during recovery.</p>
</li>
</ul>
<p>The recovery log explicitly reported:</p>
<pre><code class="language-text">selected new timeline ID: 10
</code></pre>
<p>The cluster had been on timeline 9 before recovery, so the PITR resulted in a new timeline, timeline 10.</p>
<p>This timeline change is an important part of PostgreSQL recovery. The recovered database did not simply continue the old history. It created a new branch of WAL history from the point at which recovery stopped.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6ac4c0714d7cbc5d02966747/23447f36-edd9-4989-b7e6-f76a5ab77772.png" alt="" style="display:block;margin:0 auto" />

<h2>Patroni and Cluster Management Considerations</h2>
<p>Running PITR in a Patroni-managed PostgreSQL environment requires coordination between the recovery procedure and the cluster manager.</p>
<p>Patroni is designed to manage PostgreSQL availability and cluster state. During a manual recovery operation, allowing Patroni to continue normal cluster management could interfere with the recovery process.</p>
<p>For this exercise, I therefore paused Patroni before stopping <code>pg1</code> and restoring its data directory.</p>
<p>The recovery was performed against <code>pg1</code> rather than attempting to restore all three nodes simultaneously. The objective was to establish and verify the recovered PostgreSQL state first.</p>
<p>The other nodes were already in a failed replica state before the PITR exercise, so I did not use them as part of the recovery verification. I also did not claim that the entire three-node cluster had been successfully rebuilt as part of this exercise.</p>
<p>That distinction matters. Successfully recovering one PostgreSQL instance to a point in time is different from completing a full Patroni cluster rebuild and resynchronisation.</p>
<p>A production recovery procedure would require additional consideration for the remaining members, cluster membership, replication state, and how the recovered node should be introduced back into the HA topology.</p>
<h2>What This Exercise Taught Me</h2>
<p>This exercise reinforced several practical lessons about PostgreSQL backup and recovery.</p>
<h3>Backups are only useful when recovery is understood</h3>
<p>Creating a backup is not the same as having a reliable recovery strategy. The recovery test exposed assumptions that would have been easy to miss without actually restoring the database.</p>
<p>Running <code>pgbackrest check</code> before the exercise also provided confidence that the repository and WAL archiving were functioning before the recovery test began.</p>
<h3>WAL archiving is central to PITR</h3>
<p>The full backup provided the base state, while archived WAL allowed PostgreSQL to replay changes forward until the requested recovery target.</p>
<p>For a PITR workflow, both pieces matter. A valid base backup without the required WAL history is not sufficient to recover to an arbitrary point in time.</p>
<h3>Recovery changes the timeline</h3>
<p>The recovery created timeline 10 from the previous timeline 9.</p>
<p>This is an important PostgreSQL concept because timeline history records branching in WAL history. After PITR, the recovered database has a new history from the recovery point onward.</p>
<h3>Recovery errors contain useful information</h3>
<p>The failed restore attempts were not simply obstacles. Each error pointed toward a specific condition that needed to be understood:</p>
<ul>
<li><p>An incorrect Docker image reference meant I was invoking the wrong image.</p>
</li>
<li><p>A stale <code>postmaster.pid</code> indicated leftover PostgreSQL state.</p>
</li>
<li><p>The non-empty data directory showed that pgBackRest was protecting the existing cluster from an accidental overwrite.</p>
</li>
<li><p>The Docker volume mount explained why the data directory could not simply be renamed.</p>
</li>
<li><p>The initial connection refusal required checking PostgreSQL's actual process and recovery state instead of assuming the restore had failed.</p>
</li>
</ul>
<h3>HA management and recovery are different concerns</h3>
<p>Patroni manages PostgreSQL availability and cluster state, while pgBackRest handles backup and restore operations.</p>
<p>A recovery procedure needs to account for both systems. Pausing cluster management before modifying the PostgreSQL data directory prevented Patroni from attempting to manage the instance while the restore was being performed.</p>
<p>Most importantly, the exercise demonstrated that a successful recovery should be <strong>verified with evidence</strong>, not inferred from a successful restore command. The final checks showed the expected data state, confirmed that PostgreSQL had exited recovery, and demonstrated the creation of a new WAL timeline.</p>
<h2>Recovery Failures and Resolutions</h2>
<table>
<thead>
<tr>
<th>Failure</th>
<th>Root cause</th>
<th>Resolution</th>
</tr>
</thead>
<tbody><tr>
<td>Docker could not find <code>patroni-pg1</code></td>
<td>Container name was used as the image name</td>
<td>Identified and used the actual image <code>patroni-dev291-pg1</code></td>
</tr>
<tr>
<td>pgBackRest reported PostgreSQL as running</td>
<td>Stale <code>postmaster.pid</code> remained in the data directory</td>
<td>Verified PostgreSQL was stopped and removed the stale PID file</td>
</tr>
<tr>
<td>pgBackRest refused to restore into the data directory</td>
<td>Existing PostgreSQL cluster files were present</td>
<td>Used pgBackRest <code>--force</code> for the controlled overwrite</td>
</tr>
<tr>
<td>Data directory could not be renamed</td>
<td>Directory was the root of a Docker volume mount</td>
<td>Kept the mounted volume and allowed pgBackRest to clean and restore it</td>
</tr>
<tr>
<td>PostgreSQL initially refused connections</td>
<td>PostgreSQL had not been started after the restore</td>
<td>Checked process state, started PostgreSQL manually, and inspected the startup log</td>
</tr>
<tr>
<td>Timeline verification query failed</td>
<td><code>timeline_id</code> is not a column exposed by PostgreSQL's normal SQL tables</td>
<td>Verified the new timeline from PostgreSQL's recovery log</td>
</tr>
</tbody></table>
<p>The important lesson from these failures was that recovery troubleshooting should be evidence-driven. Each error was investigated to determine the actual system state before choosing the next action.</p>
<h2>Conclusion</h2>
<p>This PITR exercise gave me a much better understanding of PostgreSQL recovery than simply reading about the process.</p>
<p>I started with a validated pgBackRest configuration and a fresh full backup, created a known recovery scenario, archived the required WAL, and recovered PostgreSQL to a specific timestamp.</p>
<p>The recovery also exposed practical operational issues around stale PostgreSQL state, Docker volume mounts, restore safety, WAL replay, and Patroni cluster management.</p>
<p>Most importantly, I verified the result rather than assuming the restore had worked. The data created before the recovery target remained, the data created after the target disappeared, PostgreSQL exited recovery, and the recovery created timeline 10 from the previous timeline 9.</p>
<p>The next step for me is to continue building on this work by studying the PostgreSQL backup and recovery ecosystem more deeply and contributing where I can identify a genuine documentation or technical gap.</p>
<p>The lab itself is not an upstream contribution. It is the practical foundation for that next stage.</p>
<h2>Project Repository</h2>
<p>The complete lab configuration, recovery documentation, and supporting files are available on GitHub:</p>
<p><a href="https://github.com/david-onwuka/DEV-291-patroni-postgresql-disaster-recovery">https://github.com/david-onwuka/DEV-291-patroni-postgresql-disaster-recovery</a></p>
<p>The repository contains the PostgreSQL 16, Patroni, etcd, and pgBackRest lab used for this recovery exercise.</p>
<p>The lab itself is not an upstream contribution. It is a practical recovery exercise that I used to strengthen my understanding of PostgreSQL high availability, backup, WAL archiving, point-in-time recovery, and disaster recovery procedures.</p>
]]></content:encoded></item></channel></rss>