Skip to content

proxmox

Run the Hypervisor Migration as a Project, Not as a Long Weekend

Defining success criteria upfront roughly doubles a project's success rate. Here is what that looks like applied to a VMware to Proxmox migration.


6/21/2026 · No. 20 · 5 min read

The infrastructure team has done the arithmetic, the renewal quote has arrived, and somebody has said the words “how hard can it be, it’s just VMs.” This is the moment a platform migration either becomes a project or becomes a story people tell at the next job.

The two disciplines this site covers, project management and infrastructure certification, meet almost nowhere as directly as here. So it is worth spending a post on the overlap, because the failure mode is well documented and entirely avoidable.

The finding that should shape the plan

PMI’s 2025 Pulse of the Profession cites earlier PMI research with a blunt conclusion: projects that define success criteria upfront and put a well-established performance measurement system in place have success rates nearly two times higher.

Read that again with a migration in mind. Not “projects with better engineers.” Not “projects with more budget.” Projects where somebody wrote down what finished looks like before starting.

Now ask what “done” means for a hypervisor migration. Most teams, if pressed, would say “all the VMs are running on the new platform.” That definition passes on the day of the last cutover and fails three weeks later when nobody can restore a backup, the monitoring is still pointed at the old cluster, and the one application owner who was on leave discovers their latency doubled.

The same report notes that only 18% of project professionals demonstrated both high skill proficiency and strong use of roadblock mitigation tactics. The gap between knowing the practice and applying it is where most of the damage happens.

Success criteria that survive contact

A defensible definition of done for this kind of migration has more than one line in it. Something closer to:

  • Every workload runs on the new platform at agreed performance thresholds, measured, not asserted.
  • Backups run on the new platform and a restore has been tested for each workload class.
  • Monitoring and alerting cover the new platform, and the old alerts are retired rather than muted.
  • Runbooks are rewritten for the new platform, including the failure procedures nobody has needed yet.
  • The team can perform routine operations without the migration consultant.
  • The old platform is decommissioned and its licences are actually cancelled.

That last one is the most commonly missed, and it is the item the whole business case rested on. A migration that leaves the old cluster running “just in case” for another renewal cycle has delivered a cost increase.

The risks worth writing down

Risk registers get a bad reputation because they are so often theatre. This is a case where the exercise pays, because the risks are known in advance and several have cheap mitigations.

Storage rework is the schedule risk. As covered in what transfers from vSphere to Proxmox, the jump from VMFS to ZFS or Ceph is the largest conceptual change in the migration. Estimating it as “reconfigure storage” is how a three-month plan becomes a six-month one.

Skills are the delivery risk. Proxmox VE is Debian underneath and does little to hide it. If the team is fluent in a vSphere client and uncomfortable at a shell, that is a training line item, not an attitude problem. Budget it before somebody discovers it at two in the morning.

Third-party tooling is the surprise risk. Backup products, monitoring agents, and automation written against vSphere APIs do not transfer. Inventory every integration before the first migration wave, not after.

Rollback capability is the risk everyone assumes away. For every wave, know how to get back. There is a window during which a migrated workload can return to the source cluster cheaply, and that window closes when you reclaim the storage. Write down when it closes.

Sequence the waves by blast radius, not by convenience

The instinct is to migrate the easy things first, which usually means the things nobody cares about. That is right for the wrong reason. The value of an early wave is not that it is easy, it is that it teaches you what your runbook got wrong while the consequences are small.

A sequence that tends to work:

  1. Something disposable that you own entirely. Confirm the process end to end, including a restore.
  2. Something real but tolerant, with a known owner who will tell you honestly how it went.
  3. The bulk of the estate, in waves sized so a bad wave is recoverable within one maintenance window.
  4. The workloads with the most political weight, last, when the process is boring.

Between waves, hold a short retrospective and change the runbook. A migration that reaches wave four using the same runbook it started with has not been paying attention.

The part that is not technical

Stakeholder management is the skill this project consumes most, which will not surprise anyone who has read the PMI data. Application owners do not experience a migration as a cost saving. They experience it as risk imposed on them by somebody else’s budget, and they are not wrong.

The thing that makes it go smoothly is unglamorous: tell each owner what will happen to their workload, when, what the rollback is, and how they can verify it themselves afterward. Give them a way to say “this got slower” that results in someone looking, rather than being told the graphs are fine.

Do that and the migration is a project. Skip it and the migration is a series of arguments held after midnight.

Sources: PMI Pulse of the Profession 2025.

Related