Orchestration of DR
Bringing several servers of the same customer back in the right order, with waits, notifications and checks run by the Sefthy Agent inside the machines.
When a site holds several servers, the order they come back in matters. If the business application starts while the database is still replaying its logs, the application faults and someone has to step in at the worst possible moment. With Orchestration of DR you write the restart sequence once and Sefthy runs it for you.
The feature is included in the Sefthy and Sefthy PRO plans, and configuring it needs the orchestration management permission.
Anteprima video temporaneamente non disponibile.
Apri il file direttamente →Where you build it
Open the customer Connector page. Down the page you find the Orchestration of DR card with the action sequence. Every row is a step, and you change the order by dragging the handle to the left of the step number.
The Add Action button opens the panel where you pick the type.
The four action types
- Start VM, restores and starts the virtual machine of the DR you pick in Select DR.
- Notification, sends an email with your own text to your company address. Useful to know a key step went through.
- Wait, pauses the sequence for the number of seconds you set.
- Check, verifies that the machine that just started is really ready before moving to the next step.
What a Check can verify
In Check Type you choose how to verify.
- ARP-Ping, the network reachability check. Nothing is needed on the server itself.
- Agent: service running, the Windows service or systemd unit you name is running.
- Agent: TCP port listening, the port you name accepts connections, for example 1433 for SQL Server.
- Agent: web address responding, the address you give answers with the expected status code, and optionally with an expected piece of text.
- Agent: path or file present, the file or folder you give exists.
The four checks prefixed with Agent run inside the machine that just started and need an enrolled Sefthy Agent on that DR, the program you install on the server during onboarding. If that DR has none, saving is refused and you are left with ARP-Ping and a wait.
For every check you set How many tries? and the timeout between one try and the next, and you decide whether to get an email when the check passes and when it fails.
If a check fails
By default the sequence carries on and the outcome is recorded. That is deliberate, a six month old configuration must not be able to hold up a customer restore.
If stopping is better than carrying on for that step, turn on If the check fails, would you like to pause the orchestration?. The connector page then shows the pause banner, which names the step that stopped and why, and once you have fixed it you carry on with Resume Orchestration.
How long it takes
Under the sequence the Console shows the Estimated total time to complete recovery orchestration, adding up the individual steps. It is the number to give the customer when they ask how long they will be down.
How to start it
There are two ways in, and they end up in the same place.
- From the Orchestration of DR card on the connector page, with the Initiate DR button. It appears when every DR in the group has at least one backup and none has been restored yet.
- From the page of a DR that belongs to the group. The Initiate DR modal tells you so and offers the Restore entire group (Orchestrated Disaster Recovery) switch, with the Start Orchestrated Disaster Recovery button that takes you to the connector page. Turn the switch off to restore that single server.
Either way, the modal that opens on the connector asks how the restored machines should be connected, Restore with Connector or Restore with Emergency VPN, the second only when every DR in the group is on a PRO plan. From there the whole sequence starts.