Repository navigation
Redfish: Host-HA fails to mark a powered-off KVM host as Down #13376
Description
Activity
Testing on CloudStack provider/driver (nested labs) never exposes this bug, as sending the force stop VM to parent ACS lab, doesn't return error, so nested hypervisor is always marked as Down - and VM HA would begin - i.e. we could never catch this bug in virtual/ACS env/testing.
So when it comes to "ok we should fence the host" - first ask "are you dead/off" - if yes, fine continue, if not, do ipmi/redfish/cloudstack STONITH, then check to make sure host is dead - do NOT relly on status code/error returned from IPMI/Redfish/ParentACS - do explicit check:
- no unpredictable errorts
- works on top of all drivers (ipmi/redfish/cloudstack)
- does an explicit check and then proceeds.
- if no answer/check doesn't return any results - do not proceed, error out (for safety reasons)
Update 11/06: for ipmi 2.0 -sending POWER OFF to a host which is alredy OFF (shutdown or previously executed ipmi.....power off command) - does NOT report errors - exit code ($?) = 0 always
- changed the title
[-]Host-HA never marks a powered-off KVM host Down because the OOBM fence (power-off) fails against an already-off chassis — VMs only recover after the dead host is powered back on[/-][+]Redfish: Host-HA never marks a powered-off KVM host Down because the OOBM fence (power-off) fails against an already-off chassis — VMs only recover after the dead host is powered back on[/+]on Jun 11, 2026 - assigned and unassigned
on Jun 15, 2026 - changed the title
[-]Redfish: Host-HA never marks a powered-off KVM host Down because the OOBM fence (power-off) fails against an already-off chassis — VMs only recover after the dead host is powered back on[/-][+]Redfish: Host-HA fails to mark a powered-off KVM host as Down[/+]on Jun 15, 2026 - linked a pull request that will close this issueKVM HA: fence by confirming host power state (fix host stuck in Fencing when already powered off) #13377
on Jun 17, 2026 - added 2 commits that reference this issue
on Jun 19, 2026 Independent reproduction on a second site with the
ipmitooldriver (this issue and #12921 are both Redfish — this confirms the failure mode is driver-agnostic, as the issue text anticipated).Environment
- CloudStack 4.22.1.0, KVM, management clustered on 2 nodes.
- OOBM driver: ipmitool (Dell iDRAC, LAN/623). Host-HA + OOBM enabled, VM-HA enabled.
- Primary storage: SharedMountPoint, zone-wide, GFS2 (
isStorageSupportHA() == true, so the legacy investigator is not the bottleneck — same as your Linstor note).
What we observed — two distinct failure modes
Mode 1 — real power loss (both PSUs pulled), BMC therefore unreachable → OOBM fence can never complete → stuck in
Fencing:11:03:24 WARN FenceTask Exception occurred while running FenceTask ... HAFenceException: OBM service is not configured or enabled for this host <host>This repeated every ~4 s for the whole outage. As you note in "Secondary Bug #1", the message is misleading — OOBM is configured; the real cause is that the
ipmitoolpower-off against a powerless BMC throws (the BMC/iDRAC itself had no power), andKVMHAProvider.fence()'s catch-all masks it. The host only reachedFencedonce the iDRAC/BMC regained power and finished booting — at that point the still-retryingipmitoolpower-off finally succeeded — after which:11:17:38 WARN HAManagerImpl Unable to find next HA state for current HA state=[Fenced] for event=[Ineligible] ... NoTransitionException: ... from Fenced via Ineligible(repeating every ~4 s — this is the spam addressed by #13588). VMs were only restarted after the host came back — exactly the "perverse result" described here.
Mode 2 — BMC reachable (graceful iDRAC shutdown, and separately a kernel panic
echo c > /proc/sysrq-trigger): host down, storage heartbeat stale, no failover for a full 20 minutes with Host-HA enabled. (We did not have DEBUG onorg.apache.cloudstack.haenabled, so we cannot yet pin the exact FSM state for this mode — will re-run with DEBUG.)A/B control — the clearest single piece of evidence: same kernel-panic failure, with Host-HA disabled, only legacy VM-HA active:
ClusteredAgentManagerImpl Host <id> is down. Starting HA on the VMs KVMInvestigator could not find VM ... (NOT "alive? true") Fencer KVMFenceBuilder returned true HighAvailabilityManagerExtImpl HA is now restarting VM ... on Host <other>VMs back within ~5–6 minutes. So on this hardware the legacy path works fine; enabling Host-HA is what suppresses it — while the host-HA state is not
Fenced,KVMInvestigator.isVmAlive→HAManagerImpl.isVMAliveOnHost/getHostStatusFromHAConfigreportsUp/Disconnected, so VM-HA defers indefinitely.Takeaways
- Confirms the bug is not Redfish-specific — the
ipmitoolpath fails the same way when the BMC is unreachable. IPMI's "already-off returns exit 0" only helps when the BMC still has power; in a real power loss (BMC down with the host) the power-off simply throws. - The
Fenced → IneligibleNoTransition spam and the misleadingfence()catch-all ("Secondary Bug 4.2 #1") both reproduce here verbatim. - Power-state confirmation (KVM HA: fence by confirming host power state (fix host stuck in Fencing when already powered off) #13377) and a storage-heartbeat fallback (KVM: add storage-heartbeat fencing fallback when OOBM fence fails (host stuck in Fencing on total power loss) #13589) are complementary — for total power loss you need the latter: KVM HA: fence by confirming host power state (fix host stuck in Fencing when already powered off) #13377 confirms death by reading the OOBM power state, but when the BMC itself is powerless that read is unavailable too, so the host still can't be confirmed off. The only positive signal left is the cluster's view of the host's storage heartbeat via its neighbours — which is exactly what KVM: add storage-heartbeat fencing fallback when OOBM fence fails (host stuck in Fencing on total power loss) #13589 (
kvm.ha.fence.on.storage.heartbeat) uses as a fence fallback when OOBM fails. Between them they cover both "off but BMC alive" (KVM HA: fence by confirming host power state (fix host stuck in Fencing when already powered off) #13377) and "BMC dead with the host" (KVM: add storage-heartbeat fencing fallback when OOBM fence fails (host stuck in Fencing on total power loss) #13589).
We're glad to help validate the reworked #13377 on a setup it hasn't been exercised on yet:
ipmitooldriver + SharedMountPoint (GFS2) zone-wide storage, incl. the real power-loss case. Happy to share full logs.Same problem here with Cloudstack 4.22.1.0, OpenBMC RedFish.
HA is never kicked, all vms on host are seen as started, but they are down, server HA status transits between degraded and suspect.
2026-08-04 18:41:35,939 WARN [o.a.c.h.HAManagerImpl] (pool-2-thread-37:[]) (logid:) Unable to find next HA state for current HA state=[Checking] for event=[HealthCheckFailed] for host Host {"id":6,"name":"hypr2-rck3","type":"Routing","uuid":"4a5d4d67-de21-44b3-90e1-33fc2491877e"} with id 6. com.cloud.utils.fsm.NoTransitionException: Unable to transition to a new state from Checking via HealthCheckFailed at com.cloud.utils.fsm.StateMachine2.getTransition(StateMachine2.java:108) at com.cloud.utils.fsm.StateMachine2.getNextState(StateMachine2.java:94) at org.apache.cloudstack.ha.HAManagerImpl.transitionHAState(HAManagerImpl.java:151) at java.base/jdk.internal.reflect.DirectMethodHandleAccessor.invoke(DirectMethodHandleAccessor.java:103) at java.base/java.lang.reflect.Method.invoke(Method.java:580) at org.springframework.aop.support.AopUtils.invokeJoinpointUsingReflection(AopUtils.java:344) at org.springframework.aop.framework.ReflectiveMethodInvocation.invokeJoinpoint(ReflectiveMethodInvocation.java:198) at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:163) at org.springframework.aop.interceptor.ExposeInvocationInterceptor.invoke(ExposeInvocationInterceptor.java:97) at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:186) at org.springframework.aop.framework.JdkDynamicAopProxy.invoke(JdkDynamicAopProxy.java:215) at jdk.proxy3/jdk.proxy3.$Proxy418.transitionHAState(Unknown Source) at org.apache.cloudstack.ha.task.HealthCheckTask.processResult(HealthCheckTask.java:57) at org.apache.cloudstack.ha.task.BaseHATask.call(BaseHATask.java:105) at org.apache.cloudstack.ha.task.BaseHATask.call(BaseHATask.java:37) at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:317) at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1144) at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:642) at java.base/java.lang.Thread.run(Thread.java:1583)@EduFrazao Please test with the following global setting
"commands.timeout" = "CheckHealthCommand=5,CheckOnHostCommand=5 "kvm.ha.activity.check.failure.ratio"= "0.6" "kvm.ha.activity.check.interval" = "8" "kvm.ha.activity.check.max.attempts" = "5" "kvm.ha.activity.check.timeout" = "30" "kvm.ha.degraded.max.period" = "30" "kvm.ha.fence.timeout" = "30" "kvm.ha.health.check.timeout" = "30" "kvm.ha.recover.failure.threshold" = "2" "kvm.ha.recover.timeout" = "30" "kvm.ha.recover.wait.period" = "30"@EduFrazao i have tested with a mockup redfish simulator and the host ha works
we have fixed vm ha issue with host ha and will be fixed in the upcoming 4.22.2 or 4.23
@andrijapanicsb @EduFrazao , can you test 4.22-HEAD with real redfish hard-/firmware?
@EduFrazao Please test with the following global setting
"commands.timeout" = "CheckHealthCommand=5,CheckOnHostCommand=5 "kvm.ha.activity.check.failure.ratio"= "0.6" "kvm.ha.activity.check.interval" = "8" "kvm.ha.activity.check.max.attempts" = "5" "kvm.ha.activity.check.timeout" = "30" "kvm.ha.degraded.max.period" = "30" "kvm.ha.fence.timeout" = "30" "kvm.ha.health.check.timeout" = "30" "kvm.ha.recover.failure.threshold" = "2" "kvm.ha.recover.timeout" = "30" "kvm.ha.recover.wait.period" = "30"Tested. Same behavior: Host transits from degraded to suspect, VMs on it never gets started on another host (shows as running).
2026-08-05 09:43:02,969 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-9:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:43:10,951 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-6:[ctx-0bca60b7]) (logid:356cd5c5) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:43:11,018 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-11:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:43:19,003 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-b9f34119]) (logid:b0269d8e) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:43:19,074 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-13:[]) (logid:) HA state post-transition:: new state=[Degraded], old state=[Checking], for resource id=[6], status=[true], ha config state=[Degraded]. 2026-08-05 09:43:51,159 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-1:[ctx-66099631]) (logid:58756953) HA state post-transition:: new state=[Suspect], old state=[Degraded], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:43:55,191 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-aa6ba133]) (logid:10b2cbfb) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:43:55,262 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-15:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:44:03,241 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-1:[ctx-0149983e]) (logid:6ea189de) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:44:03,307 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-17:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:44:11,294 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-6:[ctx-03239e43]) (logid:2efea5dd) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:44:11,360 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-19:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:44:19,349 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-06275c87]) (logid:5b6e6999) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:44:19,415 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-21:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:44:27,395 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-1:[ctx-57d7aae0]) (logid:f6b2aad4) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:44:27,463 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-23:[]) (logid:) HA state post-transition:: new state=[Degraded], old state=[Checking], for resource id=[6], status=[true], ha config state=[Degraded]. 2026-08-05 09:44:59,542 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-6:[ctx-6245caf2]) (logid:a99dabb3) HA state post-transition:: new state=[Ineligible], old state=[Available], for resource id=[1], status=[true], ha config state=[Ineligible]. 2026-08-05 09:44:59,563 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-6:[ctx-6245caf2]) (logid:a99dabb3) HA state post-transition:: new state=[Suspect], old state=[Degraded], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:45:03,596 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-1:[ctx-5bf998c6]) (logid:ae1cf98a) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:45:03,624 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-25:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:45:11,640 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-5:[ctx-2560ddd4]) (logid:fc2a9de9) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:45:11,662 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-1:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:45:19,681 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-48587166]) (logid:5e52ba69) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:45:19,703 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-3:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:45:27,726 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-3:[ctx-ebcf71fe]) (logid:894f0560) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:45:27,753 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-5:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:45:35,767 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-5:[ctx-3b36f1aa]) (logid:8bdde64e) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:45:35,794 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-7:[]) (logid:) HA state post-transition:: new state=[Degraded], old state=[Checking], for resource id=[6], status=[true], ha config state=[Degraded]. 2026-08-05 09:46:07,902 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-d8c7a401]) (logid:c2eb4e05) HA state post-transition:: new state=[Suspect], old state=[Degraded], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:46:11,929 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-5:[ctx-2a831644]) (logid:f651b8f0) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:46:11,958 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-9:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:46:19,975 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-37ff6819]) (logid:e301711d) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:46:20,001 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-11:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:46:28,017 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-3:[ctx-da736655]) (logid:19a55180) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:46:28,038 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-13:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:46:36,058 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-5:[ctx-fcd86cf6]) (logid:d95f40d1) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:46:36,081 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-15:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:46:44,103 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-4b97d5d3]) (logid:1515b9e4) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:46:44,127 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-17:[]) (logid:) HA state post-transition:: new state=[Degraded], old state=[Checking], for resource id=[6], status=[true], ha config state=[Degraded]. 2026-08-05 09:47:16,242 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-3:[ctx-4bd2b165]) (logid:f1848cd0) HA state post-transition:: new state=[Suspect], old state=[Degraded], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:47:20,266 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-d087b4db]) (logid:bbf34e62) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:47:20,291 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-19:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:47:28,312 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-3:[ctx-15de9fb2]) (logid:f5fb894f) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:47:28,418 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-21:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:47:36,356 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-5:[ctx-2bf564ee]) (logid:7bd90672) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:47:36,380 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-23:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:47:44,397 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-4:[ctx-6d079343]) (logid:cd751ab6) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:47:44,419 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-25:[]) (logid:) HA state post-transition:: new state=[Suspect], old state=[Checking], for resource id=[6], status=[true], ha config state=[Suspect]. 2026-08-05 09:47:52,439 DEBUG [o.a.c.h.HAManagerImpl] (BackgroundTaskPollManager-3:[ctx-0676188b]) (logid:32396f39) HA state post-transition:: new state=[Checking], old state=[Suspect], for resource id=[6], status=[true], ha config state=[Checking]. 2026-08-05 09:47:52,465 DEBUG [o.a.c.h.HAManagerImpl] (pool-3-thread-1:[]) (logid:) HA state post-transition:: new state=[Degraded], old state=[Checking], for resource id=[6], status=[true], ha config state=[Degraded].@EduFrazao i have tested with a mockup redfish simulator and the host ha works
we have fixed vm ha issue with host ha and will be fixed in the upcoming 4.22.2 or 4.23
Very nice!!! Thank you very mutch.
@andrijapanicsb @EduFrazao , can you test 4.22-HEAD with real redfish hard-/firmware?
Yes! There is pre-built packages of this testing version?
@andrijapanicsb @EduFrazao , can you test 4.22-HEAD with real redfish hard-/firmware?
Yes! There is pre-built packages of this testing version?
No, the nightly builds should contain these. But I don’t think anyone in the community provides these publicly.
edit 🤦 : yes there are : http://download.cloudstack.org/testing/nightly/2026-08-05/
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNo status
TL;DR When Host-HA tries to fence an already shut down host via Redfish BMC driver - host status never moves the host into the Down state --> VM-HA never kick (VMs never get started on other hosts)
Redfish: Host-HA never marks a powered-off KVM host
Downbecause the fence (Redfish OOBM power-off) can't succeed against an already-off chassis — VM-HA only triggers once the dead host is powered back onISSUE TYPE
COMPONENT NAME
CLOUDSTACK VERSION
CONFIGURATION
VmHaEnabled).isStorageSupportHA() == true, so the legacy investigator is not the bottleneck here).OS / ENVIRONMENT
/redfish/v1/Systems/System.Embedded.1).SUMMARY
When a KVM host that has host-HA + OOBM enabled is hard powered off (e.g. forced chassis-off from the BMC console, or a real power/cable failure), CloudStack never transitions the host to
Downand therefore never restarts its VMs on other hosts. The host stays inAlert/Disconnectedindefinitely.Root cause: the host-HA state machine only declares a host dead (
HAState.Fenced→ investigatorStatus.Down) after a successful fence, and the fence is implemented as an active OOBM power-off. Against an already-off chassis that power-off cannot succeed (the BMC rejects it), so the host is pinned in theFencingstate and retried forever. The investigator mapsFencingtoStatus.Disconnected, notStatus.Down, so VM-HA is never invoked.The perverse result: the VMs are only recovered once the original (dead) host is powered back on - even during BIOS booting stage — at which point the pending power-off finally succeeds, the host transitions to
Fenced/Down, and HA restarts the VMs elsewhere. This defeats the purpose of HA.All three current branches are affected by the identical issue: the relevant code is byte-identical on
4.22andmain, and functionally identical on4.20(only a method rename and logger formatting differ). There is no4.21branch upstream. Per-element diff verification is in the CLOUDSTACK VERSION section below.STEPS TO REPRODUCE
hostA.hostAat the BMC (chassis power off / simulate power loss). The BMC itself stays reachable.hostAin CloudStack over the next 20+ minutes.EXPECTED RESULTS
Down→ VM-HA restartshostA's VMs on other hosts within a few minutes.ACTUAL RESULTS
hostAremains inAlert(host status) with the host-HA state stuck inFencing.Offthe entire time, but that knowledge is never used to declare the host down.Up(while HA state isSuspect) and thenDisconnected(while HA state isFencing) — neverDown.hostA).hostAis powered back on, the fence power-off finally succeeds → host goesDown→ VM-HA restarts the VMs on other hosts.ROOT CAUSE ANALYSIS
Decision chain (only
FencedyieldsDown)KVMInvestigator.getHostAgentStatus()→haManager.getHostStatusFromHAConfig(host)(
plugins/hypervisors/kvm/src/main/java/com/cloud/ha/KVMInvestigator.java:81)HAManagerImpl.getHostStatusFromHAConfig()maps HA state → host status(
server/src/main/java/org/apache/cloudstack/ha/HAManagerImpl.java:315):Fenced→Status.DownDegraded/Recovering/Fencing→Status.DisconnectedAvailable/Suspect/Checking/Recovered) →Status.UpAgentManagerImplonly fires theHostDownevent andscheduleRestartForVmsOnHost(...)when the investigator returnsStatus.**Down**(
engine/orchestration/src/main/java/com/cloud/agent/manager/AgentManagerImpl.java:1147,:1200).So VM-HA for an HA-eligible KVM host requires the host-HA state machine to reach
Fenced.Reaching
Fencedrequires a successful power-offFencing → FencedonEvent.Fenced(
api/src/main/java/org/apache/cloudstack/ha/HAConfig.java:139).FenceTask.processResult()only firesEvent.Fencedwhen the fence returnedtrue; otherwise it does nothing and the poll loop retriesFencingforever viaRetryFencing(
server/src/main/java/org/apache/cloudstack/ha/task/FenceTask.java:45; retry atserver/src/main/java/org/apache/cloudstack/ha/HAManagerImpl.java:724).KVMHAProvider.fence()→outOfBandManagementService.executePowerOperation(host, PowerOperation.OFF, null)and returnsresp.getSuccess()(
plugins/hypervisors/kvm/src/main/java/org/apache/cloudstack/kvm/ha/KVMHAProvider.java:87).executePowerOperation()throwsCloudRuntimeExceptionwhenever the driver response is not successful — it never returnssuccess=false(
server/src/main/java/org/apache/cloudstack/outofbandmanagement/OutOfBandManagementServiceImpl.java:432).Why the power-off fails against an already-off host (Redfish)
PowerOperation.OFF→RedfishResetCmd.GracefulShutdown(
plugins/outofbandmanagement-drivers/redfish/src/main/java/org/apache/cloudstack/outofbandmanagement/driver/redfish/RedfishWrapper.java:34).RedfishClient.executeComputerSystemReset()POSTs to.../Actions/ComputerSystem.Resetand throwsRedfishExceptionif the HTTP status is not 2XX(
utils/src/main/java/org/apache/cloudstack/utils/redfish/RedfishClient.java:300-312).GracefulShutdownis invalid because there is no running OS to shut down. 409 ∉ 2XX →RedfishException→CloudRuntimeException→HAFenceException→FenceTaskseesresult=false→ noFencedtransition → stuck inFencing.chassis power offagainst an already-off / unreachable BMC returns a non-zero exit code, judged purely by process exit status with no "already in target state" handling —IpmitoolWrapper.executeCommands()→result.isSuccess().)Net effect
The fence requires confirming an active power-off transition, but a host that is already off (precisely the case where restarting its VMs is safe) cannot be "powered off successfully." The safety mechanism deadlocks in exactly the scenario it exists to handle. VMs recover only when the dead host returns.
LOG EVIDENCE (two-MS cluster; host
kvm-host01, id:1, uuidaaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee; hostnames/IPs/VM names below are anonymised examples)OOBM STATUS poll knew the chassis was off the whole time (MS #1 log):
Investigator never returns Down —
UpwhileSuspect, thenDisconnectedwhileFencing(MS #1 log):The fence itself, on the MS node that owns the HA config (MS #2 log) — repeated every ~4s for ~20 min:
Counts over the outage: ~618 ×
409, ~308 ×HAFenceException, 930 ×Fencingstate lines,Starting HA on ... = 1(only at the very end).VM-HA only fires after the host is powered back on (MS #2 log):
(chassis Off→On detected ~15:35:05 in MS #1 log.)
SECONDARY BUGS surfaced by this incident
Misleading error message. Every fence failure logs
OOBM service is not configured or enabled for this host ..., but OOBM is configured and working. The catch-all inKVMHAProvider.fence()(plugins/hypervisors/kvm/src/main/java/org/apache/cloudstack/kvm/ha/KVMHAProvider.java:97-100) assumes any exception means "OOBM not configured," hiding the real cause (HTTP 409 / already off). This actively misdirects troubleshooting.Misleading "fencing performed" alerts. Each failed fence attempt emits
alertType=30 — "HA Fencing of host id=1 ... performed"becauseFenceTask.processResult()callssendAlert(resource, HAState.Fencing)unconditionally regardless ofresult(server/src/main/java/org/apache/cloudstack/ha/task/FenceTask.java:54). Admins receive a flood of "fencing performed" alerts while fencing is in fact failing continuously.SUGGESTED FIX (direction)
Make fencing treat "host is already off" as a successful fence, and stop hiding the real error:
KVMHAProvider.fence(), query OOBM power STATUS first; if the chassis is alreadyOff, returntrue(host is effectively fenced) instead of issuing a power-off that 409s. (A confirmed-off host is safe to declare fenced.)GracefulShutdown/ForceOffwhen already off) as success; and/or preferForceOffoverGracefulShutdownfor the HA fence path.fence()catch block to surface the actual driver error rather than "OOBM not configured."FenceTaskalerts reflect actual success/failure of the fence.NOTES
4.22.1.0.LinstorPrimaryDataStoreDriverImpl.isStorageSupportHA()returnstrue, so the legacy KVM investigator does not short-circuit; the host-HA framework path (above) is in effect.