Draw the user request path
Follow a request through DNS, networking, load balancing, application, identity, database and storage. Identify components whose failure stops the whole service. Two applications using one storage system still share a failure point.
Choose the right mechanism
Restarting a VM on another host reduces hardware failure impact but takes time. An application cluster may maintain service through multiple instances. Active/passive and active/active designs have different implications for writes, sessions and conflicts.
- Place components in genuinely independent failure domains.
- Reserve capacity for the loss of a node.
- Document data consistency and network partition behaviour.
Understand quorum and fencing
A cluster must decide which members may operate when communication breaks down. Quorum helps make that decision; fencing can isolate a node to prevent unsafe concurrent writes. Design depends on technology and topology. An odd or even server count alone does not settle the issue.
Exercise failure and failback
Test host, link, storage or site loss as appropriate. Verify detection, failover, residual capacity and client reconnection. Also test the return to normal. High availability does not replace backups or recovery planning for widespread corruption.
Things to check.
- End-to-end dependencies identified
- Failure domains separated
- Quorum and degraded capacity checked
- Failover and failback tested
WHAT ABOUT YOUR CONTEXT?
Let’s find the right approach.
SysWings supports you from scoping to operations. Let’s discuss your constraints and priorities.
Cloud & hosting ↗Let’s talk about your project ↗