Repository navigation
KVM: destroy with vm.destroy.forcestop leaves the instance running when its host is briefly disconnected #14232
Copy link
Copy link
Open
Labels
Milestone
Description
Activity
- added a commit that references this issue
on Sep 23, 2026 🎯 Triage report
When
vm.destroy.forcestop=trueand a KVM host is only brieflyDisconnected(e.g. during a rolling agent/management-server restart), destroying a running instance releases its NICs/IPs/volumes and marks it stopped without the domain actually being stopped. The domain keeps running unmanaged, its IP gets reassigned to a new instance, and its root volume is stuck inDestroy. This is a well-documented, code-level analysis with logs from a real occurrence and a clear root cause identified inVirtualMachineManagerImpl.advanceStop()/destroy().📊 Assessment
Dimension Value Reasoning Type type:bugClear defect causing data/state inconsistency and resource leaks. Component component:kvmIssue is specific to KVM/libvirt hosts and the KVM power-report/destroy path. Severity n/a (not applied) Impact is significant (duplicate IPs, stranded storage) similar to #14206 (Major), but left for maintainer judgement given the specific vm.destroy.forcestopprecondition.Labels type:bug, component:kvm Conservative labeling; no speculative labels added. Coding agent Needs more info Root cause and code paths are well identified, but the correct fix requires a design decision on how advanceStop()/destroy()should distinguish "host is gone" vs "host is temporarily unreachable" across all affected host states (Disconnected,Connecting,Alert,Rebalancing), which needs maintainer input before an agent should attempt an implementation.🔗 Similar issues
- KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206 (related) — same end-state (unmanaged running instance, IP reuse, stranded volume) but triggered via a stale out-of-band power report during VM start, rather than the destroy path with
vm.destroy.forcestop. The reporter explicitly notes this is the same failure mode via a different trigger.
💡 Notes and suggestions
- Since KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206 and this issue share the same underlying pattern (forced/assumed "stopped" state without host confirmation leading to orphaned domains), a maintainer may want to track them together or reference one from the other, and consider whether a shared fix (e.g., only allowing forced release when host state is
Down/Removed, not merely unreachable) can address both. - Reproduction requires a host transitioning through
Disconnectedduring an agent/management-server restart while a destroy withexpunge=trueis in flight — this is timing-sensitive and may be easier to validate via targeted unit/integration tests aroundVirtualMachineManagerImpl.advanceStop()andreleaseVmResources()than live rolling-upgrade reproduction. - Suggested next step: confirm with the reporter/maintainers which host states besides
Disconnected(e.g.,Connecting,Alert,Rebalancing) should also block the forced release path, since the issue proposes several.
Generated by Daily Issue Triage · sonnet50 81.1K · ◷
Add this agentic workflows to your repo
To install this agentic workflow, run
gh aw add githubnext/agentics/workflows/daily-issue-triage.md@d7c1dc4b72b00607a67caaffdcc216cb64379cf9- KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206 (related) — same end-state (unmanaged running instance, IP reuse, stranded volume) but triggered via a stale out-of-band power report during VM start, rather than the destroy path with
ISSUE TYPE
COMPONENT NAME
CLOUDSTACK VERSION
CONFIGURATION
vm.destroy.forcestop=true. KVM hosts. Any rolling restart of the agents or the management servers while instances are being destroyed.OS / ENVIRONMENT
KVM / libvirt.
SUMMARY
With
vm.destroy.forcestop=true, destroying an instance whose host is briefly disconnected releases the instance's NICs, IP addresses and volumes without stopping it. The domain keeps running on the host with no record in CloudStack. Its IP is handed to the next instance, so two live instances answer for the same address, and its root volume stays inDestroybecause the delete fails while the domain holds the image.Same end state as #14206, different trigger. #14206 is a stale power report during start; this one is the destroy path itself.
STEPS TO REPRODUCE
vm.destroy.forcestop=true.cloudstack-agenton a KVM host (or a management server, which disconnects the agents connected to it while they rebalance).Disconnected, destroy a running instance on it withexpunge=true.Any client that destroys instances routinely hits this on every rolling upgrade.
EXPECTED RESULTS
The destroy fails, the instance stays
Running, and the caller retries once the host is back. Or the destroy waits for the host.ACTUAL RESULTS
From a real occurrence on 4.22, during a rolling package upgrade:
Across two rolling upgrades, 31 instances on 6 hosts were left running this way, and new instances were given their addresses within hours.
CAUSE
UserVmManagerImpl.destroyVm(DestroyVMCmd),VirtualMachineManagerImpl.destroy()andVirtualMachineManagerImpl.advanceExpunge()all stop the instance withcleanUpEvenIfUnableToStop = vm.destroy.forcestop. InadvanceStop(), a forced stop that getsAgentUnavailableExceptionorOperationTimedoutExceptionreleases the resources and marks the instance stopped:A forced stop means "the host cannot tell us, treat the instance as stopped". That is right when the host is gone (
Down,Removed). It is wrong when the host is only unreachable for a while, whatever its status says:Disconnected,Connecting,AlertorRebalancing, and alsoUpwhile a crashed management server's hosts have not yet been taken over or a disconnect investigation is inconclusive. Those hosts are expected back with their domains still running.vm.destroy.forcestopis a global setting applied to every destroy, not a statement by the caller that it knows the host is gone. It should not bypass that distinction.IMPACT
Destroyand the storage cleanup fails on it every run