Skip to content

KVM: destroy with vm.destroy.forcestop leaves the instance running when its host is briefly disconnected #14232

Description

@bhouse-nexthop
ISSUE TYPE
  • Bug Report
COMPONENT NAME
engine-orchestration, destroyVirtualMachine, KVM
CLOUDSTACK VERSION
4.22
CONFIGURATION

vm.destroy.forcestop=true. KVM hosts. Any rolling restart of the agents or the management servers while instances are being destroyed.

OS / ENVIRONMENT

KVM / libvirt.

SUMMARY

With vm.destroy.forcestop=true, destroying an instance whose host is briefly disconnected releases the instance's NICs, IP addresses and volumes without stopping it. The domain keeps running on the host with no record in CloudStack. Its IP is handed to the next instance, so two live instances answer for the same address, and its root volume stays in Destroy because the delete fails while the domain holds the image.

Same end state as #14206, different trigger. #14206 is a stale power report during start; this one is the destroy path itself.

STEPS TO REPRODUCE
  1. Set vm.destroy.forcestop=true.
  2. Restart cloudstack-agent on a KVM host (or a management server, which disconnects the agents connected to it while they rebalance).
  3. While the host is Disconnected, destroy a running instance on it with expunge=true.

Any client that destroys instances routinely hits this on every rolling upgrade.

EXPECTED RESULTS

The destroy fails, the instance stays Running, and the caller retries once the host is back. Or the destroy waits for the host.

ACTUAL RESULTS

From a real occurrence on 4.22, during a rolling package upgrade:

13:39:02 WARN  Unable to stop VM instance {"id":3645123,...,"state":"Stopping"} due to [AgentUnavailableException:
               Resource [Host:110] is unreachable: Host 110: Host with specified id is not in the right state: Disconnected]
13:39:02 WARN  Unable to actually stop VM instance {"id":3645123,...} but continue with release because it's a force stop
13:39:02 DEBUG VM instance {"id":3645123,...} is stopped on the host.  Proceeding to release resource held.
13:39:02 DEBUG Successfully released network resources for the VM ...
13:39:15 DEBUG Expunged VM instance {"id":3645123,...}
13:41:58 WARN  Host reports 9 instance(s) that do not exist in CloudStack DB, they are running unmanaged. host: hv104, instances: [i-625-3645123-VM, ...]

Across two rolling upgrades, 31 instances on 6 hosts were left running this way, and new instances were given their addresses within hours.

CAUSE

UserVmManagerImpl.destroyVm(DestroyVMCmd), VirtualMachineManagerImpl.destroy() and VirtualMachineManagerImpl.advanceExpunge() all stop the instance with cleanUpEvenIfUnableToStop = vm.destroy.forcestop. In advanceStop(), a forced stop that gets AgentUnavailableException or OperationTimedoutException releases the resources and marks the instance stopped:

} catch (AgentUnavailableException | OperationTimedoutException e) {
    logger.warn("Unable to stop {} due to [{}].", ...);
} finally {
    if (!stopped) {
        if (!cleanUpEvenIfUnableToStop) { ... throw ... }
        else { logger.warn("Unable to actually stop {} but continue with release because it's a force stop", vm); ... }
    }
}
releaseVmResources(profile, cleanUpEvenIfUnableToStop);

A forced stop means "the host cannot tell us, treat the instance as stopped". That is right when the host is gone (Down, Removed). It is wrong when the host is only unreachable for a while, whatever its status says: Disconnected, Connecting, Alert or Rebalancing, and also Up while a crashed management server's hosts have not yet been taken over or a disconnect investigation is inconclusive. Those hosts are expected back with their domains still running.

vm.destroy.forcestop is a global setting applied to every destroy, not a statement by the caller that it knows the host is gone. It should not bypass that distinction.

IMPACT
  • an instance runs unmanaged and invisible
  • its IP is reassigned, giving an address conflict between two live instances
  • its root volume stays in Destroy and the storage cleanup fails on it every run

Activity

  1. added a commit that references this issue on Sep 23, 2026
    017c41e
  2. github-actions commented on Sep 24, 2026

    @github-actions

    🎯 Triage report

    When vm.destroy.forcestop=true and a KVM host is only briefly Disconnected (e.g. during a rolling agent/management-server restart), destroying a running instance releases its NICs/IPs/volumes and marks it stopped without the domain actually being stopped. The domain keeps running unmanaged, its IP gets reassigned to a new instance, and its root volume is stuck in Destroy. This is a well-documented, code-level analysis with logs from a real occurrence and a clear root cause identified in VirtualMachineManagerImpl.advanceStop()/destroy().

    📊 Assessment

    Dimension Value Reasoning
    Type type:bug Clear defect causing data/state inconsistency and resource leaks.
    Component component:kvm Issue is specific to KVM/libvirt hosts and the KVM power-report/destroy path.
    Severity n/a (not applied) Impact is significant (duplicate IPs, stranded storage) similar to #14206 (Major), but left for maintainer judgement given the specific vm.destroy.forcestop precondition.
    Labels type:bug, component:kvm Conservative labeling; no speculative labels added.
    Coding agent Needs more info Root cause and code paths are well identified, but the correct fix requires a design decision on how advanceStop()/destroy() should distinguish "host is gone" vs "host is temporarily unreachable" across all affected host states (Disconnected, Connecting, Alert, Rebalancing), which needs maintainer input before an agent should attempt an implementation.

    🔗 Similar issues

    💡 Notes and suggestions
    • Since KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206 and this issue share the same underlying pattern (forced/assumed "stopped" state without host confirmation leading to orphaned domains), a maintainer may want to track them together or reference one from the other, and consider whether a shared fix (e.g., only allowing forced release when host state is Down/Removed, not merely unreachable) can address both.
    • Reproduction requires a host transitioning through Disconnected during an agent/management-server restart while a destroy with expunge=true is in flight — this is timing-sensitive and may be easier to validate via targeted unit/integration tests around VirtualMachineManagerImpl.advanceStop() and releaseVmResources() than live rolling-upgrade reproduction.
    • Suggested next step: confirm with the reporter/maintainers which host states besides Disconnected (e.g., Connecting, Alert, Rebalancing) should also block the forced release path, since the issue proposes several.

    Generated by Daily Issue Triage · sonnet50 81.1K · ◷

    Add this agentic workflows to your repo

    To install this agentic workflow, run

    gh aw add githubnext/agentics/workflows/daily-issue-triage.md@d7c1dc4b72b00607a67caaffdcc216cb64379cf9
    
  3. added this to the 4.22.2 milestone on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions