Repository navigation
Restarting a KVM VMs after its host fails when the management server is also unreachable #14230
Copy link
Copy link
Open
Description
Activity
🎯 Triage report
This is a feature request/design proposal for a decentralized HA/fencing mechanism that would allow surviving KVM hosts to restart VMs from a failed host even when the management server and/or database is unreachable (e.g., during a network partition or when the management server itself runs as a guest on the failed host). The author proposes options such as Corosync/Pacemaker/STONITH or a LINSTOR+drbd-reactor based approach, and references related PR #13589 and issue #14206.
📊 Assessment
Dimension Value Reasoning Type type:new-feature Proposes new distributed HA/fencing capability not currently part of CloudStack's architecture. Component component:kvm Proposal is specifically about KVM hosts/agents and libvirt domain reconstruction. Severity n/a Feature request, not a bug. Labels type:new-feature, component:kvm See reasoning above. Coding agent Not suitable This is a large architectural/design proposal (distributed quorum, fencing, agent-to-agent communication) requiring maintainer discussion and design decisions before any implementation could begin. 🔗 Similar issues
- KVM: add storage-heartbeat fencing fallback when OOBM fence fails (host stuck in Fencing on total power loss) #13589 (related) — referenced by the author as a possibly related or overlapping effort.
- KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206 (related) — referenced by the author as a possibly related HA issue.
- Redfish: Host-HA fails to mark a powered-off KVM host as Down #13376 (related) — "Redfish: Host-HA fails to mark a powered-off KVM host as Down", adjacent HA/fencing reliability issue.
💡 Notes and suggestions
- This proposal touches on core HA/fencing architecture and would benefit from discussion on the dev mailing list or a design document (per CloudStack's contribution process) before any code is written.
- Maintainers should confirm whether this overlaps with existing/planned work in PR KVM: add storage-heartbeat fencing fallback when OOBM fence fails (host stuck in Fencing on total power loss) #13589 and issue KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206 to avoid duplicated design effort.
- Key open questions for the author/maintainers: how would quorum/fencing decisions avoid split-brain with the existing HA subsystem that assumes a central management server; how would agent-to-agent proxying interact with existing security models (mutual auth between agents); what happens to VM state reconciliation once the management server reconnects.
Generated by Daily Issue Triage · sonnet50 140.8K · ◷
Add this agentic workflows to your repo
To install this agentic workflow, run
gh aw add githubnext/agentics/workflows/daily-issue-triage.md@d7c1dc4b72b00607a67caaffdcc216cb64379cf9
The required feature described as a wish
Today if a KVM host fails and the management server and/or database is unreachable at the same moment (a remote site outage, a network partition, or the management server itself running as a guest on the failing host) CloudStack cannot restart the affected VMs. Every HA and DRS decision path runs inside the management server process and depends on the database.
I propose to add a small, storage-aware cluster mechanism that can restart a VM on a surviving host using its own quorum and fencing, and hands control back cleanly once the management server returns. This might require the Agents to actually keep shadow copy of each of the VMs configuration (of just enough per-VM data to rebuild the libvirt domain: disk identifiers, NIC MACs, CPU/memory, bridge/VLAN mapping). Possibly the agents can exchange information with each other and in case one Agent looses connectivity to management server act as a proxy towards the management server?
This might be an addition/similar or even fix to the #13589 , #14206 , or rely on some other mechanisms as surviving hosts must agree "host A (and possibly the management server) is gone" without a central arbiter and must not false-positive on an ordinary network hiccup,
Options I see as viable: