Skip to content

Replication

Replication copies a machine’s disks to another node on a schedule. It is how you get a recent copy of a machine onto a second node without buying shared storage.

That makes it the practical foundation for high availability on smaller clusters: HA restarts a machine elsewhere but does not move its disks, so something has to have put a copy there first.

  • Machines on ZFS storage. Replication does not work on LVM-Thin, which is the default on installations that did not choose ZFS.
  • A cluster of at least two nodes.
  1. Open Datacenter → Replication.
  2. Click Add.
  3. Set CT/VM ID to the machine.
  4. Set Target to the node the copy should go to.
  5. Set Schedule — how often to replicate.
  6. Optionally set Rate limit (MB/s).
  7. Leave Enabled ticked and click Create.

The first run copies everything; later runs send only what changed, so they are much faster.

The schedule is your recovery point. Failing over to a replicated copy loses everything written since the last successful run, so:

Schedule You could lose
Every 15 minutes Up to 15 minutes of work
Hourly Up to an hour
Nightly Up to a day

Frequent replication costs network and disk activity on both nodes. Pick per machine rather than applying one schedule to everything — a busy database and a print server do not deserve the same interval.

The datacenter view lists the jobs. The node and guest views add the operational detail and two buttons worth knowing:

Where What it adds
node → Replication Status, Last Sync and Duration columns, plus Log and Schedule now
guest → Replication The same, for one machine
  • Log shows the last run’s output. Look here first when a job fails.
  • Schedule now runs a job immediately — useful right before a planned failover, so the copy is as fresh as possible.

Cluster-wide defaults, including global rate limiting, live under Datacenter → Options → Replication Settings.

Replication and high availability together

Section titled “Replication and high availability together”
Shared storage Local storage + replication
Machine restarts elsewhere Yes Yes
Data at the moment of failure All of it Up to the last successful sync
Extra hardware Shared storage None

Both are valid. Replication trades a defined amount of data loss for not needing a SAN or Ceph cluster — which for many workloads is the right trade, as long as it is a decision rather than a surprise.

What you see What to do
The machine cannot be selected Its disks are not on a storage type that supports replication. Move them to ZFS.
A job keeps failing Open Log on the node or guest view for the reason. Common causes are the target node being unreachable or short of space.
The first run takes hours Expected — it copies everything. Later runs send only changes. Use a rate limit so it does not crowd out other traffic.
Replication slowed the machine down Reduce the frequency, or set a rate limit.
After failover, recent work is missing That is the design. Everything since the last successful sync is lost; shorten the schedule if that is too much.

Still stuck? Contact VM2Cloud support.