DEV Community

Cover image for Day 52: Undo Rolls Forward, and the Disk Already Mounted Is the One Not to Trust
Nnamdi Felix Ibe
Nnamdi Felix Ibe

Posted on AI-assisted

Day 52: Undo Rolls Forward, and the Disk Already Mounted Is the One Not to Trust

Day 52 of DevOps, Day 2 of Azure. Both tasks today were about commands that do something a little different from what their names suggest.

kubectl rollout undo sounds like a rewind. It is not one; it rolls forward to an old template. And az vm create sounds like it creates a VM. It does, along with a network, a firewall rule, a public address and one more disk than I asked for, and that extra disk is the one not to keep anything on.

One Kubernetes task, one Azure task. Roll a Deployment back to its previous version, then create an Azure VM with a specific image, size and disk type. The tasks come from the KodeKloud Engineer platform.

Undo rolls forward

kubectl rollout undo deployment nginx-deployment
kubectl rollout status deployment nginx-deployment
Enter fullscreen mode Exit fullscreen mode

Yesterday's rolling update kept the old ReplicaSet on purpose. Today's task is what it was kept for: a release has a bug, so go back to the previous revision.

Under the hood, kubectl's rollback code takes the pod template stored in the older ReplicaSet and patches it into the Deployment's .spec.template. A changed template is exactly what starts a rollout, as yesterday's post covered, so an undo is a new rollout with the same rolling behaviour as any other update. The Kubernetes docs put it briefly: each rollback updates the revision of the Deployment.

Two consequences are worth knowing.

Only the template goes back. The docs say "only the Deployment's Pod template part is rolled back". If someone scaled the Deployment or changed its strategy since that revision, those changes stay.

And the history has a limit. Revisions are the old ReplicaSets, and .spec.revisionHistoryLimit keeps ten of them by default. Set it to zero, and in the docs' words, "a new Deployment rollout cannot be undone".

One habit I am adding is kubectl rollout history before the undo, to see what there is to go back to. Its CHANGE-CAUSE column shows <none> unless you set the kubernetes.io/change-cause annotation, and the old --record flag that used to fill it is deprecated. A history where every revision says <none> tells you nothing at the moment you most need it.

The point I would underline, though, is this one. Undo fixes the cluster, not the source of truth. If the buggy image is still in a manifest file, the next kubectl apply of that file rolls the bug straight back out. The rollback is only finished when the file, or the commit, goes back too.

One command, a lot of resources

The Azure task was a single VM: Ubuntu 24.04, Standard_B1s, a 30 GB Standard HDD disk, SSH access.

az vm create does much more than its name says. Alongside the VM, it created a virtual network and subnet, a network security group with an SSH rule, a public IP address, a network interface, an OS disk and a data disk. Microsoft describes this as the CLI using default values to create any required supporting resources.

The catch is on the way out. Microsoft's page on deleting VMs says that by default, the disks, NICs and public IPs associated with a VM are kept when the VM is deleted. One command builds it all, and the matching delete leaves most of it behind, where a leftover disk or public IP is still billed.

Coming from AWS, the concept with no direct equivalent is the resource group. Every Azure resource lives in exactly one, and deleting the group deletes everything in it. In a lab that is the cleanest teardown Azure offers.

A naming trap while I am here: "Standard HDD" in the portal is Standard_LRS in the CLI, and Standard SSD is StandardSSD_LRS. One token apart, a different storage medium. LRS is not the disk type at all. It is the redundancy: locally redundant storage, three copies within one data centre.

Three disks, and one to leave alone

lsblk on the new VM showed three disks, not two:

sda       30G  disk
├─sda1    29G  part /
sdb        4G  disk
└─sdb1     4G  part /mnt
sdc       30G  disk
Enter fullscreen mode Exit fullscreen mode

sda is the OS disk, and sdc is the data disk, attached but raw. sdb is the one I did not ask for. It is the temporary disk, and its size comes with the VM size, 4 GiB on a B1s. It arrives formatted and mounted at /mnt.

It looks like free storage, ready to use. Microsoft says data on it "might be lost during a maintenance event, when you redeploy a VM, or when you stop the VM", though it does survive a normal restart. The Linux VM FAQ puts it more bluntly: "Don't use the temporary disk, mounted under (/mnt) to store data."

The nearest AWS equivalent is instance store, and the rule is the same: nothing you need to keep goes there.

The data disk arriving raw is normal, the same as an extra EBS volume. Making it usable means partitioning, formatting, mounting, and an /etc/fstab entry by UUID with nofail. Microsoft recommends the UUID because device paths like /dev/sdc1 are not persistent and change on reboot.

The reason I had wrong

The lab host runs as root, and az vm create takes its default admin username from the local account. root is on Microsoft's list of disallowed VM usernames, along with admin, administrator, user and test. So my notes said the create would fail unless I set a username myself.

The CLI reference says otherwise. The default is the current OS username, and "If the default value is system reserved, then default value will be set to azureuser." Running as root, the CLI would have picked azureuser on its own. What does get rejected is a reserved name you pass explicitly.

I still set --admin-username azureuser, and I would again, because it makes the result independent of who runs the command. But the reason in my notes was wrong, and a wrong reason is worth correcting even when the command it produced was right.

Read what the command did

An undo that rolls forward. A create that builds a small network, and a delete that leaves most of it standing. A disk that is mounted and ready, and not safe to use. In each case the name was a fair summary and a poor specification. The rollout history, the resource list and the documentation had the real answer.

So here is the Day 52 question. The last time you rolled something back, did you also change the file that would have deployed it again?

Day 52 down. Forty-eight to go.

Top comments (5)

Collapse
 
merbayerp profile image
Mustafa ERBAY •

Great one again, Nnamdi 😄

“Undo fixes the cluster, not the source of truth” is probably the most important line here. A successful kubectl rollout undo can give you that beautiful green feeling… right until GitOps looks at the cluster and says, “Cute. Anyway…” and deploys the broken state again. 😂

And Azure handing you /mnt already mounted feels almost suspiciously friendly: “Here, have some free storage!”

Azure, five minutes later: “You weren’t keeping anything important there, right?” 😅

To answer your question: for me, a rollback isn’t really finished until the desired state is corrected too. Otherwise you didn’t fix the deployment — you just borrowed some time from the next deployment. 😄

Day 52 down. Keep them coming! 🚀

Collapse
 
ndcodes profile image
Nnamdi Felix Ibe •

Thank you! You know what you said about borrowing sometime for the next deployment is actually a debt with a deployment-shaped repayment date.

The GitOps version is worse than it looks, too. Without a reconciler, the bug comes back on the next apply, which is a moment someone chose. With one, the cluster gets corrected automatically, so the rollback gets undone by the thing whose entire job is to be right. The tool works perfectly, and the outcome is a fresh outage.

That's what I keep running into. The dangerous component is rarely the broken one. It's the correct one holding a stale definition of correct.

And Azure's /mnt has an even better detail. It survives a reboot, which is exactly enough reliability to teach you the wrong lesson. If it lost data every restart, nobody would ever trust it. Instead it works fine until the maintenance event nobody scheduled.

Thank you for reading my article.

Day 53 in a few hours. 😄

Collapse
 
merbayerp profile image
Mustafa ERBAY •

Exactly 😂 “deployment-shaped repayment date” is a much better way to put it.

And that GitOps case is beautifully evil: nothing is broken. The reconciler is healthy, automation is working perfectly, and the system confidently recreates the outage because yesterday’s truth is still today’s desired state. 😄

Also /mnt surviving reboots is such a trap. Azure basically gives you just enough evidence to build confidence before eventually teaching the lesson the expensive way. 😂

Looking forward to Day 53, my friend. At this rate I’m starting to expect every innocent-looking command to have a plot twist. 😄

Thread Thread
 
ndcodes profile image
Nnamdi Felix Ibe • • Edited

Hahaha😀, really a plot twist indeed. You got me cracking up. And I'm not being clever. Every one of them is just the docs saying something slightly different from the command name.

"Yesterday's truth is still today's desired state." That one's going in my notes. 😄

Day 53 tomorrow. 👊

Some comments may only be visible to logged-in visitors. Sign in to view all comments.