Skip to content

KVM HA fencing fails using Redfish with HPE ILO with misleading OBM error when Redfish reset returns HTTP 400, leaving host stuck in Fencing #12921

Description

@TimServers

problem

ISSUE TYPE
  • Bug Report
COMPONENT NAME

KVM, HA, Out-of-band Management, Redfish

CLOUDSTACK VERSION

4.22.x

CONFIGURATION
  • KVM hypervisor
  • Shared primary storage on NFS
  • Host HA enabled
  • Host under test has Out-of-band management enabled
  • OBM driver: redfish
  • OBM address: 10.9.3.166
  • Host under test:
    • Name: kvm-18-1.servercontrol.com.au
    • UUID: 0fb632c6-c3d4-418a-9fa8-ee21afdc9f0a
  • HA provider: kvmhaprovider
  • sync.interval = 60
  • no ha.tag configured
OS / ENVIRONMENT
  • CloudStack management servers on Ubuntu
  • MySQL 8
  • KVM hosts on Ubuntu 24.04 / libvirt
  • BMC/iLO exposed via Redfish
  • Primary storage: NFS
SUMMARY

When a KVM host is powered off unexpectedly, CloudStack detects the host failure and enters HA fencing for the host. However, fencing fails with a misleading exception:

HAFenceException: OBM service is not configured or enabled for this host

In reality, OBM is configured and enabled, and the host object in CloudStack confirms this. The actual underlying failure is that the Redfish reset request returns HTTP 400.

This leaves the host stuck in HA state = Fencing and appears to block/delay HA recovery actions for VMs on that host.

Additionally, manual Redfish testing from the management server confirms that the Redfish endpoint exists and accepts a valid POST reset request with {"ResetType":"On"}, which strongly suggests the issue is not OBM configuration but the specific fencing action CloudStack is attempting.

EXPECTED RESULTS

If the host is powered off and OBM/Redfish is correctly configured, CloudStack should be able to fence the host successfully or recognize that the host is already safely powered off.

Expected behavior:

  1. CloudStack detects the host failure.
  2. CloudStack investigates the host and confirms it is down.
  3. CloudStack performs a successful fence action using the configured OBM provider, or treats an already-powered-off host as successfully fenced.
  4. The host exits the fencing workflow cleanly.
  5. HA recovery for affected VMs proceeds normally.

The system should not report that OBM is “not configured or enabled” when:

  • OBM is enabled on the host object, and
  • the real failure is an HTTP 400 from a Redfish action.
ACTUAL RESULTS

CloudStack detects the host failure and enters fencing, but fencing fails repeatedly.

Host details in CloudStack show OBM is enabled and configured:

  • Out-of-band management = true
  • Out-of-band management driver = redfish
  • Out-of-band management address = 10.9.3.166
  • Out-of-band management power state = Off
  • HA state = Fencing

Management log shows repeated fencing failures with a misleading error message:

2026-03-31 03:19:33,574 WARN  [o.a.c.k.h.KVMHAProvider] ... OOBM service is not configured or enabled for this host Host {"id":6,"name":"kvm-18-1.servercontrol.com.au"...} error is Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.166/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.166'. The expected HTTP status code is '2XX' but it got '400'
2026-03-31 03:19:33,575 WARN  [o.a.c.h.t.FenceTask] ... org.apache.cloudstack.ha.provider.HAFenceException: OBM service is not configured or enabled for this host kvm-18-1.servercontrol.com.au

The same pattern repeats multiple times.

Manual Redfish POST to power the system on succeeds:

curl -k -u ADMIN:xxxxxx -H 'Content-Type: application/json' -X POST https://10.
9.3.166/redfish/v1/Systems/1/Actions/ComputerSystem.Reset -d '{"ResetType":"On"}'

Response:

{"error":{"code":"iLO.0.10.ExtendedInfo","message":"See @Message.ExtendedInfo for more information.","@Message.ExtendedInfo":[{"MessageId":"Base.1.18.Success"}]}}

versions

ACS 4.22
Ubuntu 24.04
HPE ILO 1.72 Nov 09 2025

The steps to reproduce the bug

STEPS TO REPRODUCE
  1. Configure a KVM host with:

    • Host HA enabled
    • OBM enabled
    • OBM driver redfish
    • valid Redfish endpoint / credentials
  2. Confirm the host is healthy and part of a KVM cluster.

  3. Power off the host unexpectedly.

  4. Observe the management server log and host HA state in the CloudStack UI.

  5. Observe that the host enters Fencing state and fencing repeatedly fails.

  6. Test the Redfish endpoint manually from the management server:

    • GET /redfish/v1/Systems/1
    • POST /redfish/v1/Systems/1/Actions/ComputerSystem.Reset with {"ResetType":"On"}

What to do about it?

Please review the KVM HA fencing flow for Redfish-backed OBM hosts.

What appears to be happening in this case:

  • CloudStack correctly detects the KVM host failure and moves the host into Fencing
  • CloudStack then attempts a Redfish power/reset action
  • the Redfish request returns HTTP 400
  • CloudStack wraps that as OBM service is not configured or enabled for this host, which is misleading because OBM is in fact enabled and functional
  • manual Redfish power control from the management server succeeds against the same host and endpoint

Requested fixes / improvements:

  1. Fix the misleading exception text

    • Do not report OBM service is not configured or enabled when OBM is configured and the actual failure is an HTTP error returned by the BMC.
  2. Log the exact Redfish fencing request details

    • Log the ResetType being used
    • Log the response body / Redfish message ID
    • This would make debugging much easier
  3. Handle already-powered-off hosts as successfully fenced where appropriate

    • If the host is already in PowerState = Off, HA should not remain stuck in Fencing just because an additional reset action is rejected by the BMC.
  4. Review the Redfish action selected by KVM HA fencing

    • In this environment, manual POST to ComputerSystem.Reset with {"ResetType":"On"} succeeds
    • CloudStack fencing appears to be using a request that the BMC rejects with HTTP 400
    • Please verify whether the selected ResetType is valid for the current power state of the host
  5. Ensure VM recovery is not blocked indefinitely by this fencing failure

    • Once the host is confirmed down and safely powered off, HA should be able to proceed with recovery of affected VMs

In short: this looks like a Redfish fencing handling bug in KVM HA, not an OBM configuration problem.

Activity

  1. TimServers commented on Mar 31, 2026

    @TimServers
    Author

    I realised hpe redfish is not yet supported. Please reach out if you need, I have plenty of hpe servers to test with

  2. kiranchavala commented on Apr 2, 2026

    @kiranchavala
    Member

    @TimServers Cloudstack support redfish for powermanagement (oobm)

  3. TimServers commented on Apr 7, 2026

    @TimServers
    Author

    @kiranchavala It doesn't work with HPE's ILO implementation of Redfish and HPE is not listed as supported. I tested it, and it doesn't work for HA when trying to execute power reset it fails with a malformed API request to the HPE redfish API. Details above.

  4. kiranchavala commented on Apr 7, 2026

    @kiranchavala
    Member

    @TimServers Could you please provide us the management server logs

    I think the oobm is not correctly configured on your kvm host

    from the database please check

    select * from oobm

    Please find the logs that i had tested with the emulator

    https://github.com/HewlettPackard/ilo-redfish-emulator

    2026-04-07 06:21:52,400 DEBUG [c.c.a.ApiServlet] (qtp1390913202-19:[ctx-97ebc32c]) (logid:d8636a40) ===START===  10.0.3.251 -- POST
    command=issueOutOfBandManagementPowerAction
    response=json
    action=RESET
    hostid=2ee21e48-324c-4ff4-a15b-25cc4dab6f15
    sessionkey=S2O_d1rRWsnra_6qsJgRLI0KTOA
    
    2026-04-07 06:21:52,400 DEBUG [c.c.a.ApiServlet] (qtp1390913202-19:[ctx-97ebc32c]) (logid:d8636a40) Two factor authentication is already verified for the user 2, so skipping
    2026-04-07 06:21:52,408 DEBUG [c.c.a.ApiServer] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) CIDRs from which account 'Account [{"accountName":"admin","id":2,"uuid":"acc418c4-2137-11f1-b16a-1e00190002a8"}]' is allowed to perform API calls: 0.0.0.0/0,::/0
    2026-04-07 06:21:52,411 INFO  [o.a.c.a.DynamicRoleBasedAPIAccessChecker] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) Account for user id acc4d966-2137-11f1-b16a-1e00190002a8 is Root Admin or Domain Admin, all APIs are allowed.
    2026-04-07 06:21:52,411 DEBUG [o.a.c.a.StaticRoleBasedAPIAccessChecker] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) RoleService is enabled. We will use it instead of StaticRoleBasedAPIAccessChecker.
    2026-04-07 06:21:52,411 DEBUG [o.a.c.r.ApiRateLimitServiceImpl] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) API rate limiting is disabled. We will not use ApiRateLimitService.
    2026-04-07 06:21:52,439 DEBUG [c.c.a.ApiServer] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) Retrieved cmdEventType from job info: HOST.OOBM.ACTION
    2026-04-07 06:21:52,445 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) submit async job-190, details: AsyncJob {"accountId":2,"cmd":"org.apache.cloudstack.api.command.admin.outofbandmanagement.IssueOutOfBandManagementPowerActionCmd","cmdInfo":"{\"response\":\"json\",\"ctxUserId\":\"2\",\"sessionkey\":\"S2O_d1rRWsnra_6qsJgRLI0KTOA\",\"action\":\"RESET\",\"hostid\":\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\",\"httpmethod\":\"POST\",\"ctxStartEventId\":\"683\",\"ctxDetails\":\"{\\\"interface com.cloud.host.Host\\\":\\\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\\\"}\",\"ctxAccountId\":\"2\",\"cmdEventType\":\"HOST.OOBM.ACTION\"}","cmdVersion":0,"completeMsid":null,"created":null,"id":190,"initMsid":32985768264360,"instanceId":1,"instanceType":"Host","lastPolled":null,"lastUpdated":null,"processStatus":0,"removed":null,"result":null,"resultCode":0,"status":"IN_PROGRESS","userId":2,"uuid":"fde13926-0e0f-42f7-a138-ca8a2e261c59"}
    2026-04-07 06:21:52,446 INFO  [o.a.c.f.j.i.AsyncJobMonitor] (API-Job-Executor-2:[ctx-21440db8, job-190]) (logid:77f3723a) Add job-190 into job monitoring
    2026-04-07 06:21:52,450 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl$5] (API-Job-Executor-2:[ctx-21440db8, job-190]) (logid:fde13926) Executing AsyncJob {"accountId":2,"cmd":"org.apache.cloudstack.api.command.admin.outofbandmanagement.IssueOutOfBandManagementPowerActionCmd","cmdInfo":"{\"response\":\"json\",\"ctxUserId\":\"2\",\"sessionkey\":\"S2O_d1rRWsnra_6qsJgRLI0KTOA\",\"action\":\"RESET\",\"hostid\":\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\",\"httpmethod\":\"POST\",\"ctxStartEventId\":\"683\",\"ctxDetails\":\"{\\\"interface com.cloud.host.Host\\\":\\\"2ee21e48-324c-4ff4-a15b-25cc4dab6f15\\\"}\",\"ctxAccountId\":\"2\",\"cmdEventType\":\"HOST.OOBM.ACTION\"}","cmdVersion":0,"completeMsid":null,"created":null,"id":190,"initMsid":32985768264360,"instanceId":1,"instanceType":"Host","lastPolled":null,"lastUpdated":null,"processStatus":0,"removed":null,"result":null,"resultCode":0,"status":"IN_PROGRESS","userId":2,"uuid":"fde13926-0e0f-42f7-a138-ca8a2e261c59"}
    2026-04-07 06:21:52,453 DEBUG [c.c.a.ApiServlet] (qtp1390913202-19:[ctx-97ebc32c, ctx-83d31001]) (logid:d8636a40) ===END===  10.0.3.251 -- POST
    command=issueOutOfBandManagementPowerAction
    response=json
    action=RESET
    hostid=2ee21e48-324c-4ff4-a15b-25cc4dab6f15
    sessionkey=S2O_d1rRWsnra_6qsJgRLI0KTOA
    
    2026-04-07 06:21:52,499 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-2:[ctx-21440db8, job-190, ctx-01dd7951]) (logid:fde13926) Retrieved System ID '1' with request 'GET: http://10.0.35.131/redfish/v1/Systems'
    
    2026-04-07 06:21:52,510 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-2:[ctx-21440db8, job-190, ctx-01dd7951]) (logid:fde13926) Sending ComputerSystem.Reset Command 'ForceRestart' to host '10.0.35.131' with request 'POST http://10.0.35.131/redfish/v1/Systems/1/Actions/ComputerSystem.Reset'
    
    2026-04-07 06:21:52,522 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (API-Job-Executor-2:[ctx-21440db8, job-190, ctx-01dd7951]) (logid:fde13926) Complete async job-190, jobStatus: SUCCEEDED, resultCode: 0, result: org.apache.cloudstack.api.response.OutOfBandManagementResponse/outofbandmanagement/{"hostid":"2ee21e48-324c-4ff4-a15b-25cc4dab6f15","powerstate":"On","enabled":"true","driver":"redfish","address":"10.0.35.131","port":"80","username":"root","password":"r*****","action":"RESET","description":"200","status":"true"}
    
    
    
    
    Image Image
  5. TimServers commented on Apr 8, 2026

    @TimServers
    Author

    @kiranchavala
    Redfish is configured without issue - all power actions and status works fine from the UI - we only see the above errors during host ha / fencing event.

    You can see it's configured normally.

    mysql> select * from oobm;
    +----+---------+---------+-------------+---------+------------+------+------------+--------------------------------------------------------------+--------------+---------------------+-----------------+
    | id | host_id | enabled | power_state | driver  | address    | port | username   | password                                                     | update_count | update_time         | mgmt_server_id  |
    +----+---------+---------+-------------+---------+------------+------+------------+--------------------------------------------------------------+--------------+---------------------+-----------------+
    |  1 |       5 |       1 | On          | redfish | 10.9.3.173 | 443  | CLOUDSTACK | wa1h6RRckH33xp59QJpiHEhajwv6Crl+xxxxKkyFmqGB0SV195WSsggA= |           75 | 2026-04-08 02:23:43 | 248902281439561 |
    |  2 |       6 |       1 | On          | redfish | 10.9.3.166 | 443  | CLOUDSTACK | 1f9UAW941qi953SJ+QWxuVVrfC53BOKxxxx13kLfyUddrW3/G6Bv4DOw= |           63 | 2026-04-08 02:23:03 | 248902281439561 |
    +----+---------+---------+-------------+---------+------------+------+------------+--------------------------------------------------------------+--------------+---------------------+-----------------+
    2 rows in set (0.00 sec)
    

    The issue is that OOBM works fine until a HA event occurs.

    During the HA fencing workflow is the only time we see any errors.

    All power actions via the UI / web interface work fine.

    As you can see, host never properly completes fencing.

    Out-of-band management power state
    Off
    Out-of-band management driver
    redfish
    Out-of-band management address
    10.9.3.173
    Out-of-band management port
    443
    HA enabled
    true
    HA state
    Fencing
    HA provider
    kvmhaprovider
    UEFI supported
    true
    Dedicated
    No
    

    As you can see, the curl returns a 'successful' result, but within an "error" block, I think cloudstack is mis-interpreting this as a failure.

    2026-04-08 02:40:31,761 WARN  [o.a.c.k.h.KVMHAProvider] (pool-6-thread-8:[ctx-c6509ce8]) (logid:) OOBM service is not configured or enabled for this host Host {"id":5,"name":"kvm-6-3.servercontrol.com.au","type":"Routing","uuid":"784ce3fb-3960-4659-bd28-a0295c59663b"} error is Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.173'. The expected HTTP status code is '2XX' but it got '400'.
    2026-04-08 02:40:31,761 WARN  [o.a.c.h.t.FenceTask] (pool-6-thread-19:[]) (logid:) Exception occurred while running FenceTask on a resource: org.apache.cloudstack.ha.provider.HAFenceException: OBM service is not configured or enabled for this host kvm-6-3.servercontrol.com.au org.apache.cloudstack.ha.provider.HAFenceException: OBM service is not configured or enabled for this host kvm-6-3.servercontrol.com.au
            at org.apache.cloudstack.kvm.ha.KVMHAProvider.fence(KVMHAProvider.java:98)
            at org.apache.cloudstack.kvm.ha.KVMHAProvider.fence(KVMHAProvider.java:41)
            at org.apache.cloudstack.ha.task.FenceTask.performAction(FenceTask.java:42)
            at org.apache.cloudstack.ha.task.BaseHATask$1.call(BaseHATask.java:87)
            at org.apache.cloudstack.ha.task.BaseHATask$1.call(BaseHATask.java:84)
            at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264)
            at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)
            at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)
            at java.base/java.lang.Thread.run(Thread.java:840)
    Caused by: org.apache.cloudstack.utils.redfish.RedfishException: Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.173'. The expected HTTP status code is '2XX' but it got '400'.
            at org.apache.cloudstack.utils.redfish.RedfishClient.executeComputerSystemReset(RedfishClient.java:312)
            at org.apache.cloudstack.outofbandmanagement.driver.redfish.RedfishOutOfBandManagementDriver.execute(RedfishOutOfBandManagementDriver.java:88)
            at org.apache.cloudstack.outofbandmanagement.driver.redfish.RedfishOutOfBandManagementDriver.execute(RedfishOutOfBandManagementDriver.java:64)
            at org.apache.cloudstack.outofbandmanagement.OutOfBandManagementServiceImpl.executePowerOperation(OutOfBandManagementServiceImpl.java:422)
            at jdk.internal.reflect.GeneratedMethodAccessor608.invoke(Unknown Source)
            at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
            at java.base/java.lang.reflect.Method.invoke(Method.java:569)
            at org.springframework.aop.support.AopUtils.invokeJoinpointUsingReflection(AopUtils.java:344)
            at org.springframework.aop.framework.ReflectiveMethodInvocation.invokeJoinpoint(ReflectiveMethodInvocation.java:198)
            at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:163)
            at org.apache.cloudstack.network.contrail.management.EventUtils$EventInterceptor.invoke(EventUtils.java:109)
            at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:175)
            at com.cloud.event.ActionEventInterceptor.invoke(ActionEventInterceptor.java:52)
            at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:175)
            at org.springframework.aop.interceptor.ExposeInvocationInterceptor.invoke(ExposeInvocationInterceptor.java:97)
            at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:186)
            at org.springframework.aop.framework.JdkDynamicAopProxy.invoke(JdkDynamicAopProxy.java:215)
            at jdk.proxy3/jdk.proxy3.$Proxy423.executePowerOperation(Unknown Source)
            at org.apache.cloudstack.kvm.ha.KVMHAProvider.fence(KVMHAProvider.java:90)
            ... 8 more
    
    2026-04-08 02:40:31,763 WARN  [c.c.a.AlertManagerImpl] (pool-6-thread-19:[]) (logid:) alertType=[30] dataCenterId=[1] podId=[2] clusterId=[null] message=[HA Fencing of host id=5, in dc id=1 performed].
    
    
    

    Here is the result of a successful curl

    {"error":{"code":"iLO.0.10.ExtendedInfo","message":"See @Message.ExtendedInfo for more information.","@Message.ExtendedInfo":[{"MessageId":"Base.1.18.Success"}]}}
    

    It's a "success" return, however, may be misinterpreted by cloudstack as an error?

    Here are the logs when doing power actions via the GUI, you can see it's configured fine and there are no errors - all power actions work fine. We're running the latest version of ILO.

    2026-04-08 02:29:15,255 DEBUG [c.c.a.ApiServlet] (qtp1438988851-42184:[ctx-4389f345, ctx-ed4a9052]) (logid:879c23e3) ===END===  221.121.128.73 -- POST
    command=issueOutOfBandManagementPowerAction
    response=json
    action=OFF
    hostid=784ce3fb-3960-4659-bd28-a0295c59663b
    sessionkey=oTvqEoVn91XqdoddtJUUKaxV6OY
    
    2026-04-08 02:29:16,347 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-15:[ctx-8395114c, job-616, ctx-7de6e440]) (logid:b2bc657a) Retrieved System ID '1' with request 'GET: https://10.9.3.173/redfish/v1/Systems'
    
    2026-04-08 02:29:17,775 DEBUG [o.a.c.u.r.RedfishClient] (API-Job-Executor-15:[ctx-8395114c, job-616, ctx-7de6e440]) (logid:b2bc657a) Sending ComputerSystem.Reset Command 'GracefulShutdown' to host '10.9.3.173' with request 'POST https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset'
    
    2026-04-08 02:29:17,810 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (API-Job-Executor-15:[ctx-8395114c, job-616, ctx-7de6e440]) (logid:b2bc657a) Complete async job-616, jobStatus: SUCCEEDED, resultCode: 0, result: org.apache.cloudstack.api.response.OutOfBandManagementResponse/outofbandmanagement/{"hostid":"784ce3fb-3960-4659-bd28-a0295c59663b","powerstate":"On","enabled":"true","driver":"redfish","address":"10.9.3.173","port":"443","username":"CLOUDSTACK","password":"M*****","action":"OFF","description":"200","status":"true"}
    
    2026-04-08 02:29:17,848 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl$5] (API-Job-Executor-15:[ctx-8395114c, job-616]) (logid:b2bc657a) Done executing org.apache.cloudstack.api.command.admin.outofbandmanagement.IssueOutOfBandManagementPowerActionCmd for job-616
    
    2026-04-08 02:29:17,848 INFO  [o.a.c.f.j.i.AsyncJobMonitor] (API-Job-Executor-15:[ctx-8395114c, job-616]) (logid:b2bc657a) Remove job-616 from job monitoring
    
  6. kiranchavala commented on Apr 8, 2026

    @kiranchavala
    Member

    @TimServers Thanks for the update

    could you upload the complete management server logs for the date 2026-04-08

  7. TimServers commented on Apr 8, 2026

    @TimServers
    Author

    @kiranchavala
    Sure. no problem.

    Please note that I reconfigured the servers from IPMI back to redfish to assist with testing around 5 - 10 minutes before 2026-04-08 02:29:15, so you can probably ignore anything before that time.

    After I reconfigured to use redfish, and tested power actions. Once the server was powerd off the first error around redfish getting a 400 error is around 02:34:41

    2026-04-08 02:34:41,709 WARN  [o.a.c.k.h.KVMHAProvider] (pool-5-thread-4:[ctx-445ace15]) (logid:) OOBM service is not configured or enabled for this host Host {"id":5,"name":"kvm-6-3.servercontrol.com.au","type":"Routing","uuid":"784ce3fb-3960-4659-bd28-a0295c59663b"} error is Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.173/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.173'. The expected HTTP status code is '2XX' but it got '400'.    
    
    

    management-server.log.tar.gz

    Let me know if you need any other details.

  8. TimServers commented on Apr 16, 2026

    @TimServers
    Author

    @kiranchavala I have some more details for you.

    It looks like the issue is not just related to HA, but in cloudstack's REDFISH implementation of the ComputerSystem.Reset API call specifically, as well as some logic bugs.

    The following tests were all done via the GUI "Issue out-of-band management power action" feature:

    Image

    The behaviour seems to be tied to the current state of the host.

    Scenario 1 - the host is powered ON

    The following actions all work without error:

    OFF
    RESET
    SOFT
    STATUS
    

    The following actions will all fail with the same error:

    ON
    CYCLE
    

    Scenario 2 - the host is powered OFF

    The following actions work without error:

    ON
    STATUS
    

    The following actions will all fail with the same error:

    OFF 
    SOFT
    CYCLE
    RESET
    

    When receiving an error, it's ALWAYS the same API call that's failing:

    Issue out-of-band management power action
    (kvm-18-1.servercontrol.com.au) Failed to execute System power command for host by performing 'POST' request on URL 'https://10.9.3.166/redfish/v1/Systems/1/Actions/ComputerSystem.Reset' and host address '10.9.3.166'. The expected HTTP status code is '2XX' but it got '400'.
    
    Image

    So apart from the failing API call, there also look to be some logic bugs here as well.

    1. If the host is already ON, sending the ON action attempts to send the ComputerSystem.Reset API call and fails. Sending an ON action when the host is already ON should not reboot the system.
    2. If the host is already OFF, sending the OFF action attempts to send the ComputerSystem.Reset API call and fails. Sending an OFF action should never send a reset API call.

    I hope this information helps, let me know if you need any more details.

  9. added this to the 4.22.2 milestone on May 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions