Repository navigation
KVM HA fencing fails using Redfish with HPE ILO with misleading OBM error when Redfish reset returns HTTP 400, leaving host stuck in Fencing #12921
Description
Activity
I realised hpe redfish is not yet supported. Please reach out if you need, I have plenty of hpe servers to test with
@TimServers Cloudstack support redfish for powermanagement (oobm)
@kiranchavala It doesn't work with HPE's ILO implementation of Redfish and HPE is not listed as supported. I tested it, and it doesn't work for HA when trying to execute power reset it fails with a malformed API request to the HPE redfish API. Details above.
@TimServers Could you please provide us the management server logs
I think the oobm is not correctly configured on your kvm host
from the database please check
select * from oobm
Please find the logs that i had tested with the emulator
https://github.com/HewlettPackard/ilo-redfish-emulator
2026-04-07 06:21:52,400 DEBUG [c.c.a.ApiServlet] (qtp1390913202-19:[ctx-97ebc32c]) (logid:d8636a40) ===START=== 10.0.3.251 -- POST command=issueOutOfBandManagementPowerAction response=json action=RESET hostid=2ee21e48-324c-4ff4-a15b-25cc4dab6f15 sessionkey=S2O_d1rRWsnra_6qsJgRLI0KTOA 2026-04-07 06:21:52,400 DEBUG [c.c.a.ApiServlet] (qtp1390913202-19:[ctx-97ebc32c]) (logid:d8636a40) Two factor authentication is already verified for the user 2, so skipping 2026-04-07 06:21:52,408 DEBUG [c.c.a.ApiServer] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) CIDRs from which account 'Account [{"accountName":"admin","id":2,"uuid":"acc418c4-2137-11f1-b16a-1e00190002a8"}]' is allowed to perform API calls: 0.0.0.0/0,::/0 2026-04-07 06:21:52,411 INFO [o.a.c.a.DynamicRoleBasedAPIAccessChecker] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) Account for user id acc4d966-2137-11f1-b16a-1e00190002a8 is Root Admin or Domain Admin, all APIs are allowed. 2026-04-07 06:21:52,411 DEBUG [o.a.c.a.StaticRoleBasedAPIAccessChecker] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) RoleService is enabled. We will use it instead of StaticRoleBasedAPIAccessChecker. 2026-04-07 06:21:52,411 DEBUG [o.a.c.r.ApiRateLimitServiceImpl] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) API rate limiting is disabled. We will not use ApiRateLimitService. 2026-04-07 06:21:52,439 DEBUG [c.c.a.ApiServer] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) Retrieved cmdEventType from job info: HOST.OOBM.ACTION 2026-04-07 06:21:52,445 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) submit async job-190, details: AsyncJob {"accountId":2,"cmd":"org.apache.cloudstack.api.command.admin.outofbandmanagement.IssueOutOfBandManagementPowerActionCmd","cmdInfo":"{\"response\":\"json\",\"ctxUserId\":\"2\",\"sessionkey\":\"S2O_d1rRWsnra_6qsJgRLI0KTOA\",\"action\":\"RESET\",\"hostid\":\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\",\"httpmethod\":\"POST\",\"ctxStartEventId\":\"683\",\"ctxDetails\":\"{\\\"interface com.cloud.host.Host\\\":\\\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\\\"}\",\"ctxAccountId\":\"2\",\"cmdEventType\":\"HOST.OOBM.ACTION\"}","cmdVersion":0,"completeMsid":null,"created":null,"id":190,"initMsid":32985768264360,"instanceId":1,"instanceType":"Host","lastPolled":null,"lastUpdated":null,"processStatus":0,"removed":null,"result":null,"resultCode":0,"status":"IN_PROGRESS","userId":2,"uuid":"fde13926-0e0f-42f7-a138-ca8a2e261c59"} 2026-04-07 06:21:52,446 INFO [o.a.c.f.j.i.AsyncJobMonitor] (API-Job-Executor-2:[ctx-21440db8, job-190]) (logid:77f3723a) Add job-190 into job monitoring 2026-04-07 06:21:52,450 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl$5] (API-Job-Executor-2:[ctx-21440db8, job-190]) (logid:fde13926) Executing AsyncJob {"accountId":2,"cmd":"org.apache.cloudstack.api.command.admin.outofbandmanagement.IssueOutOfBandManagementPowerActionCmd","cmdInfo":"{\"response\":\"json\",\"ctxUserId\":\"2\",\"sessionkey\":\"S2O_d1rRWsnra_6qsJgRLI0KTOA\",\"action\":\"RESET\",\"hostid\":\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\",\"httpmethod\":\"POST\",\"ctxStartEventId\":\"683\",\"ctxDetails\":\"{\\\"interface com.cloud.host.Host\\\":\\\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\\\"}\",\"ctxAccountId\":\"2\",\"cmdEventType\":\"HOST.OOBM.ACTION\"}","cmdVersion":0,"completeMsid":null,"created":null,"id":190,"initMsid":32985768264360,"instanceId":1,"instanceType":"Host","lastPolled":null,"lastUpdated":null,"processStatus":0,"removed":null,"result":null,"resultCode":0,"status":"IN_PROGRESS","userId":2,"uuid":"fde13926-0e0f-42f7-a138-ca8a2e261c59"} 2026-04-07 06:21:52,453 DEBUG [c.c.a.ApiServlet] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) ===END=== 10.0.3.251 -- POST command=issueOutOfBandManagementPowerAction response=json action=RESET hostid=2ee21e48-324c-4ff4-a15b-25cc4dab6f15 sessionkey=S2O_d1rRWsnra_6qsJgRLI0KTOA 2026-04-07 06:21:52,499 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-2:[ctx-21440db8, job-190, ctx-01dd7951]) (logid:fde13926) Retrieved System ID '1' with request 'GET: http://10.0.35.131/redfish/v1/Systems' 2026-04-07 06:21:52,510 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-2:[ctx-21440db8, job-190, ctx-01dd7951]) (logid:fde13926) Sending ComputerSystem.Reset Command 'ForceRestart' to host '10.0.35.131' with request 'POST http://10.0.35.131/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' 2026-04-07 06:21:52,522 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (API-Job-Executor-2:[ctx-21440db8, job-190, ctx-01dd7951]) (logid:fde13926) Complete async job-190, jobStatus: SUCCEEDED, resultCode: 0, result: org.apache.cloudstack.api.response.OutOfBandManagementResponse/outofbandmanagement/{"hostid":"2ee21e48-324c-4ff4-a15b-25cc4dab6f15","powerstate":"On","enabled":"true","driver":"redfish","address":"10.0.35.131","port":"80","username":"root","password":"r*****","action":"RESET","description":"200","status":"true"}

@kiranchavala
Redfish is configured without issue - all power actions and status works fine from the UI - we only see the above errors during host ha / fencing event.You can see it's configured normally.
mysql> select * from oobm; +----+---------+---------+-------------+---------+------------+------+------------+--------------------------------------------------------------+--------------+---------------------+-----------------+ | id | host_id | enabled | power_state | driver | address | port | username | password | update_count | update_time | mgmt_server_id | +----+---------+---------+-------------+---------+------------+------+------------+--------------------------------------------------------------+--------------+---------------------+-----------------+ | 1 | 5 | 1 | On | redfish | 10.9.3.173 | 443 | CLOUDSTACK | wa1h6RRckH33xp59QJpiHEhajwv6Crl+xxxxKkyFmqGB0SV195WSsggA= | 75 | 2026-04-08 02:23:43 | 248902281439561 | | 2 | 6 | 1 | On | redfish | 10.9.3.166 | 443 | CLOUDSTACK | 1f9UAW941qi953SJ+QWxuVVrfC53BOKxxxx13kLfyUddrW3/G6Bv4DOw= | 63 | 2026-04-08 02:23:03 | 248902281439561 | +----+---------+---------+-------------+---------+------------+------+------------+--------------------------------------------------------------+--------------+---------------------+-----------------+ 2 rows in set (0.00 sec)The issue is that OOBM works fine until a HA event occurs.
During the HA fencing workflow is the only time we see any errors.
All power actions via the UI / web interface work fine.
As you can see, host never properly completes fencing.
Out-of-band management power state Off Out-of-band management driver redfish Out-of-band management address 10.9.3.173 Out-of-band management port 443 HA enabled true HA state Fencing HA provider kvmhaprovider UEFI supported true Dedicated NoAs you can see, the curl returns a 'successful' result, but within an "error" block, I think cloudstack is mis-interpreting this as a failure.
2026-04-08 02:40:31,761 WARN [o.a.c.k.h.KVMHAProvider] (pool-6-thread-8:[ctx-c6509ce8]) (logid:) OOBM service is not configured or enabled for this host Host {"id":5,"name":"kvm-6-3.servercontrol.com.au","type":"Routing","uuid":"784ce3fb-3960-4659-bd28-a0295c59663b"} error is Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.173'. The expected HTTP status code is '2XX' but it got '400'. 2026-04-08 02:40:31,761 WARN [o.a.c.h.t.FenceTask] (pool-6-thread-19:[]) (logid:) Exception occurred while running FenceTask on a resource: org.apache.cloudstack.ha.provider.HAFenceException: OBM service is not configured or enabled for this host kvm-6-3.servercontrol.com.au org.apache.cloudstack.ha.provider.HAFenceException: OBM service is not configured or enabled for this host kvm-6-3.servercontrol.com.au at org.apache.cloudstack.kvm.ha.KVMHAProvider.fence(KVMHAProvider.java:98) at org.apache.cloudstack.kvm.ha.KVMHAProvider.fence(KVMHAProvider.java:41) at org.apache.cloudstack.ha.task.FenceTask.performAction(FenceTask.java:42) at org.apache.cloudstack.ha.task.BaseHATask$1.call(BaseHATask.java:87) at org.apache.cloudstack.ha.task.BaseHATask$1.call(BaseHATask.java:84) at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264) at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) at java.base/java.lang.Thread.run(Thread.java:840) Caused by: org.apache.cloudstack.utils.redfish.RedfishException: Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.173'. The expected HTTP status code is '2XX' but it got '400'. at org.apache.cloudstack.utils.redfish.RedfishClient.executeComputerSystemReset(RedfishClient.java:312) at org.apache.cloudstack.outofbandmanagement.driver.redfish.RedfishOutOfBandManagementDriver.execute(RedfishOutOfBandManagementDriver.java:88) at org.apache.cloudstack.outofbandmanagement.driver.redfish.RedfishOutOfBandManagementDriver.execute(RedfishOutOfBandManagementDriver.java:64) at org.apache.cloudstack.outofbandmanagement.OutOfBandManagementServiceImpl.executePowerOperation(OutOfBandManagementServiceImpl.java:422) at jdk.internal.reflect.GeneratedMethodAccessor608.invoke(Unknown Source) at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.base/java.lang.reflect.Method.invoke(Method.java:569) at org.springframework.aop.support.AopUtils.invokeJoinpointUsingReflection(AopUtils.java:344) at org.springframework.aop.framework.ReflectiveMethodInvocation.invokeJoinpoint(ReflectiveMethodInvocation.java:198) at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:163) at org.apache.cloudstack.network.contrail.management.EventUtils$EventInterceptor.invoke(EventUtils.java:109) at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:175) at com.cloud.event.ActionEventInterceptor.invoke(ActionEventInterceptor.java:52) at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:175) at org.springframework.aop.interceptor.ExposeInvocationInterceptor.invoke(ExposeInvocationInterceptor.java:97) at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:186) at org.springframework.aop.framework.JdkDynamicAopProxy.invoke(JdkDynamicAopProxy.java:215) at jdk.proxy3/jdk.proxy3.$Proxy423.executePowerOperation(Unknown Source) at org.apache.cloudstack.kvm.ha.KVMHAProvider.fence(KVMHAProvider.java:90) ... 8 more 2026-04-08 02:40:31,763 WARN [c.c.a.AlertManagerImpl] (pool-6-thread-19:[]) (logid:) alertType=[30] dataCenterId=[1] podId=[2] clusterId=[null] message=[HA Fencing of host id=5, in dc id=1 performed].Here is the result of a successful curl
{"error":{"code":"iLO.0.10.ExtendedInfo","message":"See @Message.ExtendedInfo for more information.","@Message.ExtendedInfo":[{"MessageId":"Base.1.18.Success"}]}}It's a "success" return, however, may be misinterpreted by cloudstack as an error?
Here are the logs when doing power actions via the GUI, you can see it's configured fine and there are no errors - all power actions work fine. We're running the latest version of ILO.
2026-04-08 02:29:15,255 DEBUG [c.c.a.ApiServlet] (qtp1438988851-42184:[ctx-4389f345, ctx-ed4a9052]) (logid:879c23e3) ===END=== 221.121.128.73 -- POST command=issueOutOfBandManagementPowerAction response=json action=OFF hostid=784ce3fb-3960-4659-bd28-a0295c59663b sessionkey=oTvqEoVn91XqdoddtJUUKaxV6OY 2026-04-08 02:29:16,347 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-15:[ctx-8395114c, job-616, ctx-7de6e440]) (logid:b2bc657a) Retrieved System ID '1' with request 'GET: https://10.9.3.173/redfish/v1/Systems' 2026-04-08 02:29:17,775 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-15:[ctx-8395114c, job-616, ctx-7de6e440]) (logid:b2bc657a) Sending ComputerSystem.Reset Command 'GracefulShutdown' to host '10.9.3.173' with request 'POST https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' 2026-04-08 02:29:17,810 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (API-Job-Executor-15:[ctx-8395114c, job-616, ctx-7de6e440]) (logid:b2bc657a) Complete async job-616, jobStatus: SUCCEEDED, resultCode: 0, result: org.apache.cloudstack.api.response.OutOfBandManagementResponse/outofbandmanagement/{"hostid":"784ce3fb-3960-4659-bd28-a0295c59663b","powerstate":"On","enabled":"true","driver":"redfish","address":"10.9.3.173","port":"443","username":"CLOUDSTACK","password":"M*****","action":"OFF","description":"200","status":"true"} 2026-04-08 02:29:17,848 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl$5] (API-Job-Executor-15:[ctx-8395114c, job-616]) (logid:b2bc657a) Done executing org.apache.cloudstack.api.command.admin.outofbandmanagement.IssueOutOfBandManagementPowerActionCmd for job-616 2026-04-08 02:29:17,848 INFO [o.a.c.f.j.i.AsyncJobMonitor] (API-Job-Executor-15:[ctx-8395114c, job-616]) (logid:b2bc657a) Remove job-616 from job monitoring@TimServers Thanks for the update
could you upload the complete management server logs for the date 2026-04-08
@kiranchavala
Sure. no problem.Please note that I reconfigured the servers from IPMI back to redfish to assist with testing around 5 - 10 minutes before 2026-04-08 02:29:15, so you can probably ignore anything before that time.
After I reconfigured to use redfish, and tested power actions. Once the server was powerd off the first error around redfish getting a 400 error is around 02:34:41
2026-04-08 02:34:41,709 WARN [o.a.c.k.h.KVMHAProvider] (pool-5-thread-4:[ctx-445ace15]) (logid:) OOBM service is not configured or enabled for this host Host {"id":5,"name":"kvm-6-3.servercontrol.com.au","type":"Routing","uuid":"784ce3fb-3960-4659-bd28-a0295c59663b"} error is Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.173'. The expected HTTP status code is '2XX' but it got '400'.Let me know if you need any other details.
@kiranchavala I have some more details for you.
It looks like the issue is not just related to HA, but in cloudstack's REDFISH implementation of the ComputerSystem.Reset API call specifically, as well as some logic bugs.
The following tests were all done via the GUI "Issue out-of-band management power action" feature:
The behaviour seems to be tied to the current state of the host.
Scenario 1 - the host is powered ON
The following actions all work without error:
OFF RESET SOFT STATUSThe following actions will all fail with the same error:
ON CYCLEScenario 2 - the host is powered OFF
The following actions work without error:
ON STATUSThe following actions will all fail with the same error:
OFF SOFT CYCLE RESETWhen receiving an error, it's ALWAYS the same API call that's failing:
Issue out-of-band management power action (kvm-18-1.servercontrol.com.au) Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.166/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.166'. The expected HTTP status code is '2XX' but it got '400'.
So apart from the failing API call, there also look to be some logic bugs here as well.
- If the host is already ON, sending the ON action attempts to send the ComputerSystem.Reset API call and fails. Sending an ON action when the host is already ON should not reboot the system.
- If the host is already OFF, sending the OFF action attempts to send the ComputerSystem.Reset API call and fails. Sending an OFF action should never send a reset API call.
I hope this information helps, let me know if you need any more details.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDiscuss
problem
ISSUE TYPE
COMPONENT NAME
KVM, HA, Out-of-band Management, Redfish
CLOUDSTACK VERSION
4.22.x
CONFIGURATION
redfish10.9.3.166kvm-18-1.servercontrol.com.au0fb632c6-c3d4-418a-9fa8-ee21afdc9f0akvmhaproviderOS / ENVIRONMENT
SUMMARY
When a KVM host is powered off unexpectedly, CloudStack detects the host failure and enters HA fencing for the host. However, fencing fails with a misleading exception:
HAFenceException: OBM service is not configured or enabled for this hostIn reality, OBM is configured and enabled, and the host object in CloudStack confirms this. The actual underlying failure is that the Redfish reset request returns HTTP 400.
This leaves the host stuck in
HA state = Fencingand appears to block/delay HA recovery actions for VMs on that host.Additionally, manual Redfish testing from the management server confirms that the Redfish endpoint exists and accepts a valid
POSTreset request with{"ResetType":"On"}, which strongly suggests the issue is not OBM configuration but the specific fencing action CloudStack is attempting.EXPECTED RESULTS
If the host is powered off and OBM/Redfish is correctly configured, CloudStack should be able to fence the host successfully or recognize that the host is already safely powered off.
Expected behavior:
The system should not report that OBM is “not configured or enabled” when:
ACTUAL RESULTS
CloudStack detects the host failure and enters fencing, but fencing fails repeatedly.
Host details in CloudStack show OBM is enabled and configured:
Out-of-band management = trueOut-of-band management driver = redfishOut-of-band management address = 10.9.3.166Out-of-band management power state = OffHA state = FencingManagement log shows repeated fencing failures with a misleading error message:
The same pattern repeats multiple times.
Manual Redfish POST to power the system on succeeds:
Response:
versions
ACS 4.22
Ubuntu 24.04
HPE ILO 1.72 Nov 09 2025
The steps to reproduce the bug
STEPS TO REPRODUCE
Configure a KVM host with:
redfishConfirm the host is healthy and part of a KVM cluster.
Power off the host unexpectedly.
Observe the management server log and host HA state in the CloudStack UI.
Observe that the host enters
Fencingstate and fencing repeatedly fails.Test the Redfish endpoint manually from the management server:
GET /redfish/v1/Systems/1POST /redfish/v1/Systems/1/Actions/ComputerSystem.Resetwith{"ResetType":"On"}What to do about it?
Please review the KVM HA fencing flow for Redfish-backed OBM hosts.
What appears to be happening in this case:
Fencing400OBM service is not configured or enabled for this host, which is misleading because OBM is in fact enabled and functionalRequested fixes / improvements:
Fix the misleading exception text
OBM service is not configured or enabledwhen OBM is configured and the actual failure is an HTTP error returned by the BMC.Log the exact Redfish fencing request details
ResetTypebeing usedHandle already-powered-off hosts as successfully fenced where appropriate
PowerState = Off, HA should not remain stuck inFencingjust because an additional reset action is rejected by the BMC.Review the Redfish action selected by KVM HA fencing
POSTtoComputerSystem.Resetwith{"ResetType":"On"}succeeds400ResetTypeis valid for the current power state of the hostEnsure VM recovery is not blocked indefinitely by this fencing failure
In short: this looks like a Redfish fencing handling bug in KVM HA, not an OBM configuration problem.