Back

CR-2026-02 · Done · 0 downtime

Moving APOLO off VMware without turning anything off

Scope
7 machines, VMware ESXi → Proxmox
Period
Aug 2025 – Feb 2026
Result
0 downtime, $0 licensing. Average node load 55% → 35% on equivalent workloads
Status
Done
Evidence
Production

Summary

I moved seven machines at APOLO from VMware ESXi to Proxmox with zero downtime. The virtualization licensing cost went to $0, so APOLO no longer pays for its hypervisor, and the researchers who share its clusters didn't lose a service along the way. Then I tuned kernel and operating-system parameters, and average node load on equivalent workloads dropped from 55% to 35%.

APOLO is the scientific computing center where I did my internship, and this migration was my internship project. My work covered two clusters, APOLO 2 and APOLO 3, and the services that keep them usable, like FreeIPA, which handles logins to the clusters, and a central Prometheus collector that feeds a Grafana server for monitoring.

The move itself was one step of several. First I evaluated Proxmox as the new platform and prepared the migration. Then I converted the machines from ESXi to Proxmox while the services kept running. Once everything was on Proxmox, I tuned the kernel and OS parameters and checked the result.

In the same environment I configured the Proxmox firewall and UFW to restrict unwanted traffic. I also kept the virtual-machine backups in a protected location on the NAS. If a host or a service fails, those backups are the way back.

How it was measured

The load figure comes from production. Average node load was compared on the two hypervisor environments, the old ESXi one and the new Proxmox one, running equivalent workloads: 55% before, 35% after. That's 20 percentage points, about 36% less load than where it started.

What I like about this comparison is that both sides ran the same kind of work. The platform underneath changed, and so did the tuning, and I count the 20 points as one result of the two together.

0%25%50%75%100%55%35%BeforeAfter

Timeline notes

When the internship ended in Feb 2026, I stayed at APOLO as a Supercomputing Research Assistant, with the same core responsibilities, until Jul 2026. So Proxmox, the platform I had evaluated, migrated to and tuned, stayed my responsibility for five more months.

I think that counts as a fair test: the same person who evaluated Proxmox, migrated to it and tuned it was also the one who had to keep it running for the next five months.

Timeline

Aug 2025Started at APOLO as Supercomputing Analyst, an internship. The migration was my project.

  1. Evaluated Proxmox as the new platform and prepared the move
  2. Moved the 7 machines from VMware ESXi to Proxmox, with no downtime
  3. Tuned kernel and operating-system parameters
  4. Average node load on equivalent production workloads, both hypervisors: 55% → 35%

Feb 2026Internship ended. I stayed at APOLO as Supercomputing Research Assistant until Jul 2026

Stack

Proxmox · VMware ESXi · Linux kernel tuning · UFW · NAS

Next change review: ASC26 in Wuxi: my part was the infrastructure Next change review: Message middleware in C++, from replication to AWS load tests Next change review: GRID-EAFIT: energy prediction models, containerized and on AWS