Repository navigation
KVM: instance left running on the host with no record in CloudStack after an out-of-band power report during start #14206
Description
Activity
🎯 Triage report
The reporter documents a race condition on KVM hosts where an out-of-band
PingCommand(triggered by an unrelated instance crashing/shutting down on the same host) causes CloudStack to mistakenly mark a different, still-starting instance asPowerReportMissing. This releases the instance's IP/NICs without ever sending aStopCommand, so the domain keeps running unmanaged on the hypervisor while its IP is reassigned to another instance and its volume becomes stranded in a pending-delete state. The report includes a real-world timeline, relevant code excerpts (VirtualMachinePowerStateSyncImpl,VirtualMachineManagerImpl,LibvirtDomainListener), and a clear root-cause breakdown across five contributing code paths.📊 Assessment
Dimension Value Reasoning Type type:bug Concrete defect with reproducible sequence and root-cause code analysis Component component:kvm Confirmed KVM/libvirt-specific interaction (out-of-band ping from LibvirtDomainListener)Severity Severity:Major Can cause IP address conflicts between live instances and stranded/undeletable primary storage; requires a specific timing window so not universally triggered, but has real operational impact Labels type:bug, component:kvm, Severity:Major Based on documented technical root cause Coding agent Needs more info Root cause is clearly identified, but the fix touches five separate code paths (power-state sync, missing-report handling, expunge/advanceStop early-return, power-on-report-for-destroyed handling, unknown-instance logging) and requires a design decision on how much of the behavior to change (e.g., is forcesemantics change safe for other callers?) before an automated agent could safely implement it🔗 Similar issues
No close duplicates were found. Related KVM/HA issues surfaced by search (e.g., #12921, #13605, #14178, #13376) concern out-of-band-management (OOBM)/host-HA fencing scenarios, not this VM power-state race condition, so they are not flagged as related.
💡 Notes and suggestions
- Maintainers should confirm intended semantics of
force/outOfBandinPingCommand: it appears designed only to expedite a reported stop, not to bypass the graceful period for a missing report. ApplyingforcetoprocessMissingVmReport()conflates "instance reported stopped" with "instance absent from a report," which is the core defect per the reporter's analysis. - Consider whether
handlePowerOffReportWithNoPendingJobsOnVM()should send aStopCommandeven in thePowerReportMissingbranch before releasing resources, to avoid orphaned running domains. advanceStop()'s early return forStopped/Destroyed/Expunging/Errorstates meansvm.destroy.forcestopnever gets a chance to contact the host — worth checking if that early-return should be conditioned on whether the host still might have an active domain.- Given RBD/Ceph is called out as the only reason the extra disk isn't deleted from under a running guest, this may deserve cross-checking against other primary storage backends where the same protection might not exist.
- Recommend a maintainer with engine-orchestration expertise validate the proposed fix approach before any implementation work begins, given the five interacting code paths and behavioral trade-offs involved.
Generated by Daily Issue Triage · sonnet50 93.5K · ◷
Add this agentic workflows to your repo
To install this agentic workflow, run
gh aw add githubnext/agentics/workflows/daily-issue-triage.md@d7c1dc4b72b00607a67caaffdcc216cb64379cf9- Maintainers should confirm intended semantics of
- linked a pull request that will close this issueengine: do not force the destroy stop while the host may reconnect #14233
on Oct 7, 2026
ISSUE TYPE
COMPONENT NAME
CLOUDSTACK VERSION
CONFIGURATION
KVM hosts. Any host where instances stop or crash while another instance is being deployed.
OS / ENVIRONMENT
KVM / libvirt.
SUMMARY
An instance can end up running on a KVM host with no record in CloudStack. Its IP address is released and later handed to another instance, so two instances answer for the same address. The root volume is marked for deletion but cannot be deleted while the domain holds it, so the storage is stranded too.
The trigger is the out-of-band ping the KVM agent sends when another instance on the same host shuts down or crashes.
STEPS TO REPRODUCE
Needs a host with instance churn during a deploy.
StartCommandis still in flight, have another instance on the same host shut itself down or crash.LibvirtDomainListener.onLifecycleChange, which callsAgent.triggerUpdate()and sends aPingCommandwithoutOfBand=true.The window that does the damage is small: the report has to be collected before the new domain exists but processed after the start job finishes. It does not fire on every deploy, but it recurs.
EXPECTED RESULTS
The deploy succeeds, or it fails and the host is told to stop the instance.
An instance absent from a report collected before it started is not treated as missing.
ACTUAL RESULTS
Sequence, from a real occurrence (times shortened):
StartCommandsentforceStartAnswerarrives, instance -> RunningPowerReportMissingadvanceStop()returns at once because state is ErrorUnable to find matched VM in CloudStack DB. name: i-2-3-VMThe domain is still running the whole time.
CAUSE
Five separate things line up.
1.
forceskips the graceful period for a missing instance.VirtualMachinePowerStateSyncImpl.processMissingVmReport():forcecomes fromPingCommand.outOfBand, set only byAgent.triggerUpdate(), called only fromLibvirtDomainListeneron a self-shutdown or crash. It was added so a reported stop takes effect immediately. Applying it to the absence of a report is the defect: a report says nothing about an instance that did not exist yet when it was collected.The KVM report lists only powered-on domains, so an instance that is still starting is simply absent.
The log line makes this hard to see. It prints "has passed graceful period" even when
forceshort-circuited the check, so it reports an elapsed time far below the graceful period as having passed it.2. The missing-report branch releases resources without stopping anything.
VirtualMachineManagerImpl.handlePowerOffReportWithNoPendingJobsOnVM():The IP is freed while the instance is still using it.
3. Expunge never contacts the host for an instance in Error or Stopped.
advanceStop()returns before it looks at the host id:vm.destroy.forcestopdoes not help, the early return happens first.4. A power-on report for a destroyed instance is only logged.
handlePowerOnReportWithNoPendingJobsOnVM()hascase Destroyed: case Expunging:log andbreak. CloudStack knows the host and the instance name at that moment and does nothing with them.5. Unknown instances are only logged at debug.
convertVmStateReport()writes one debug line per unknown instance per report and nothing acts on it, so the condition can persist unnoticed indefinitely.IMPACT
Destroyand cannot be deleted while the domain holds it, so primary storage is stranded