Adventures in data recovery - Proxmox and LVM
I tend to always want to go buy the shiny new gadget. That, when looking for a homelab, tends to meaning dropping quite a good buck. After being made fun from a couple of peers for buying new HW for my PFsense router, rather than fashioning one out of scraps, I chose to buy a second-hand thinkcentre for my new Proxmox server - space is now a constraint, and my old tower server was too big for the new house.
Just my luck - borught in February, and the SSD failed catastrophically in August. No issue on the really important data, it was on another HDD via NFS share (whew, I’m glad I chose that setup) but I had no backup of the VMs and didn’t fancy re-setting up all the NFS and accounts manually. Learnt some nifty commands while at it, so this is the write up of my adventrure.
What not to do.
Everyone says “make a copy asap and work on the copy!” but that seemed like a chore and an exageration: I just wanted some data. Mount it quickly, chroot, grab the backups, 20 minute adventure.
Yeah. While on my first attempt it did allow me to reach the boot menu and drop into Proxmox rescue mode, it degenerated real quick. Next boot, I got dropped into the grub shell, the root partition no longer bootable, readable or mountable.
So yeah. I can attest it does get worse quickly now. Don’t think another warning would have stopped me, but leaving it here for posterity.
The tricky issue: LVM
Proxmox uses LVM, and apparently at default settings it keeps the vmlinuz file inside the lvm volumes, rather than in the boot partition as sane setups do. The SSD failure corrupted the LVM data, so I could neither boot, mount or chroot into the stuff. No VGs found, no LV, no PV, no nothing.
LVM is a black magic fuckery that sits between your filesystem and the rest of the system and makes it appear as fluid, resizable volumes, which can actually span across multiple disks and so on. In my case there was only one disk, but Proxmox uses it to handle VM disks and storage pools.
This is my second encounter with LVM and I can only say it confirmed itself as my archenemy. It has only ever caused me problems.
Anyway, the LVM metadata was still mostly readable with strings and other tricks, so I held hope, got up to get the big HDD I use as a NAS, and went to flex my data recovery stale skills.
Data recovery
I ran a standard Proxmox install with default partition layout on an SSD, so the rescue recipe was the straightforward:
ddrescue -d -r3 /dev/nvme0n1 /mnt/hdd/recovery/pve-full.img /mnt/hdd/recovery/pve-full.mapplus, three days worth of patience for the 512GB image to be read and written. Yeah. I did kill the recovery once the forward pass was done, as it threathened to take a week more just to recheck failed sectors. If the provess gets interrupted, the same command just restarts from where it left off thanks to the map file, so I could always go back.
The nifty tricks I learnt for this occasion actually relate to qemu: given the image, and given I was going to attempt repairs and therefore modify it, I learnt how to create a writable overlay and mount it, so that if anything went wrong the original image stayed untouched. It’s subjectively more elegant than making a copy of a 512 GB file, and objectively WAY quicker.
# Creating the overlay
qemu-img create -f qcow2 -b /mnt/hdd/recovery/pve-full.img -F raw /mnt/hdd/recovery/work.qcow2
# In my case, the nbd kernel module was not loaded and it's needed for the next command, so I'm loading it:
sudo modprobe nbd max_part=16
# Mount the overlay
sudo qemu-nbd --connect=/dev/nbd0 /mnt/hdd/recovery/work.qcow2
sudo partprobe /dev/nbd0Now, I could work on /dev/nbd0 safely. Qemu is really nifty, and I should really tinker with it more.
LVM repair
The great success - differently from the original disk, the image recovered by ddrescue did show some volume groups!
versus@neverland > sudo vgchange -ay pve
Check of pool pve/data failed (status:64). Manual repair required!
2 logical volume(s) in volume group "pve" now activeWell, partial success. The actually important volume required manual repair. Only root + swap were active:
versus@neverland > sudo lvs -a -o +lv_size,data_percent,lv_active pve
LV VG Attr LSize Pool LSize Active
data pve twi---tz-- <348.82g <348.82g
[data_tdata] pve Twi------- <348.82g <348.82g
[data_tmeta] pve ewi------- <3.56g <3.56g
[lvol0_pmspare] pve ewi------- <3.56g <3.56g
root pve -wi-a----- 96.00g 96.00g active
swap pve -wi-a----- 8.00g 8.00g active
vm-100-disk-0 pve Vwi---tz-- 64.00g data 64.00g
vm-101-disk-0 pve Vwi---tz-- 4.00m data 4.00m
vm-101-disk-1 pve Vwi---tz-- 32.00g data 32.00gAnd running the automated repair command gave the same error as vgchange:
versus@neverland /m/h/recovery> sudo lvconvert --repair pve/data
no compatible roots found
Repair of thin metadata volume of thin pool pve/data failed (status:64). Manual repair required!Status 64 & manual repair required mean needing to run thin_chack manually. All pools in LVM have apparently two hidden subvolumes, one of which (data_tmeta) holds the mapping metadata. This is the part that needed to be first activated, then repaired.
# Activating the hidden volume:
versus@neverland > sudo lvchange -ay pve/data_tmeta
Do you want to activate component LV in read-only mode? [y/n]: y
Allowing activation of component LV.
# Getting the actual diagnosis
versus@neverland > sudo thin_check /dev/mapper/pve-data_tmeta
TRANSACTION_ID=17
METADATA_FREE_BLOCKS=928677
29 nodes in data mapping tree contain errors
0 io errors, 29 checksum errors
Thin device 1 has 17 errors and is missing 3352 mappings, while expected 320070
Thin device 3 has 12 errors and is missing 2235 mappings, while expected 405402
Check of mappings failedWhile this, on sight, kinda discouraged me, the mapping number is actually quite low - 1% data loss or less per machine. To fix this, I made a spare logical volume, which needed to be at least as big as data_tmeta, and attempted the repair onto the new volume.
# Find out the size of the LV:
versus@neverland > sudo lvs -o lv_size pve/data_tmeta
LSize
<3.56g
versus@neverland > sudo lvcreate -n data_meta_new -L 4g pve
Logical volume "data_meta_new" created.Now, here I thought I’d simply run sudo thin_repair -i /dev/mapper/pve-data_tmeta -o /dev/pve/data_meta_new, but for some reason, despite thin_check having no problems finding the root, it failed with error no compatible roots found. The same went for thin_dump, the last resort to dump metadata.
This is where the real fun began.
Back to the roots
Now: all the tools I’ve used so far for lvm are contained in the thin-provisioning-tools package, which is old enough to have, in its history, a full rewrite in Rust. Some older tools don’t even exist anymore in the rust rewrite. Guess what I needed?
I found some quick and to-the-point compile instruction with the precise tag to check out in this medium post. It fails to compile on some more recent compilers, specifically inside a thin_show_duplicates tool, so I edited that command out with sed. I also compiled against the original repo rather than the tool linked there.
git clone https://github.com/jthornber/thin-provisioning-tools.git
cd thin-provisioning-tools/
git checkout v0.9.0
# First sed command
sed -i -e '/thin_show_duplicates.cc/d' \
-e '/thin_show_metadata.cc/d' \
-e '/thin_scan.cc/d' \
Makefile.in
# Second sed command
sed -i -e '/thin_scan_cmd/d' \
-e '/thin_generate_damage_cmd/d' \
-e '/thin_generate_metadata_cmd/d' \
-e '/thin_generate_mappings_cmd/d' \
-e '/thin_show_duplicates_cmd/d' \
-e '/thin_show_metadata_cmd/d' \
-e '/thin_journal_cmd/d' \
thin-provisioning/commands.cc
autoconf
./configure --enable-dev-tools
make bin/pdata_toolsBefore working I made a copy of the metadata volume with dd and operated on that:
dd if=/dev/mapper/pve-data_tmeta of=/mnt/hdd/recovery/tmeta.bin bs=1M
./pdata_tools thin_ll_dump ./tmeta-work.bin -o lldump.xmlYay! I have metadata!
LVM archaeology
So, why is this magical tool not included if it just poofed out my metadata? The answer is: it didn’t. The normal dump simply stops at errors; this one extracted all it could, and just dumped everything else as an orphan node.
I haven’t included any explanations on LVM because I don’t understand it that well myself, but here is what I got to learn on this matter: the data I’m trying to rebuild is a tree. The root node is called superblock, and it’s block number 0. It contains a node for each device, and those contain other nodes.
You can see it quite well from the head of my lldump.xml file:
<superblock blocknr="0" data_mapping_root="529" device_details_root="543">
<device dev_id="1">
<node blocknr="533" flags="1" key_begin="0" key_end="951062" nr_entries="8" value_size="8"/>
</device>
<device dev_id="2">
<node blocknr="18805" flags="2" key_begin="0" key_end="8" nr_entries="9" value_size="8"/>
</device>
<device dev_id="3">
<node blocknr="435" flags="1" key_begin="0" key_end="485705" nr_entries="11" value_size="8"/>
</device>
</superblock>
<orphans>
<node blocknr="77737" flags="1" key_begin="0" key_end="246" nr_entries="3" value_size="8"/>
<node blocknr="77687" flags="1" key_begin="0" key_end="246" nr_entries="3" value_size="8"/>
<node blocknr="77693" flags="1" key_begin="0" key_end="937898" nr_entries="5" value_size="8"/>… and so on for 92603 orphan nodes.
Now, the tool only found one node for each device and dumped the rest as orphans. The good news? It does find the root node for each device tree!
Now let’s see if thin_ll_restore does its job, by tunning it from the metadata and against the spare volume I created earlier:
sudo ./pdata_tools thin_ll_restore -E ./tmeta-work.bin -i lldump.xml -o /dev/pve/data_meta_new
sudo ./pdata_tools thin_check /dev/pve/data_meta_new
examining superblock
TRANSACTION_ID=17
METADATA_FREE_BLOCKS=1048575
examining devices tree
examining mapping tree
checking space map countsIs this… did I do it? Setting the new volume as metadata for the actual pve/data:
versus@neverland > sudo lvconvert --thinpool pve/data --poolmetadata pve/data_meta_new
# disable the old metadata node we had force-activated earlier:
versus@neverland > sudo lvchange -an pve/data_tmeta
versus@neverland > sudo vgchange -ay pve
7 logical volume(s) in volume group "pve" now activeVictory, at last!
The actual backups
Now. Technically it should be possible for me to boot all of this up, do the backups the Proxmox way (vzdump) and go on my merry way.
Kinda. There are still some errors laying around the filesystem, and I can access the data I need in other ways. So here is what I grabbed
The VM configs
Mounting the root volume, I grabbed the config files for the VMs. they’re stored in a sql database so I went about extracting them:
mount -o ro /dev/pve/root /mnt/host/
cp /mnt/hostvar/lib/pve-cluster/config.db /mnt/hdd/recovery/backups/
# 100 is my VM id - repeat as needed for the VMs to save
sqlite3 config.db ".mode list" ".headers off" "SELECT writefile('100.conf', data) FROM tree WHERE name = '100.conf';"This was particularly important to me cause of a USB pasthrough feeding the disk to the NAS VM, which I couldn’t be bothered re-doing. Looking at the bash history it was probably just these two commands:
pct set 100 -mp0 /dev/sda,mp=/media/nas
qm set 100 -scsi2 /dev/disk/by-id/ata-WDC_WD100EFGX-68CPLN0_WD-B203DLNJBut I don’t quite remember if they were both needed or if I’d done something via UI, so better have both options.
The actual VMs
I used ddrescue, but being a ddrescue’d image there wasn’t much point to it really:
# for each VM and VM disk
sudo ddrescue -d /dev/pve/vm-100-disk-0 /mnt/hdd/recovery/backups/vm-100-disk-0.imgThe two images can be mounted with kpartx -a -v -r and checked with the appropriate filesystem check tool.
Conclusions
In my case, the OMV VM only had light errors, and will probably be bootable easily. The Fedora IoT one took heavier hits: all the recovery in the world cannot bring back undeadable bytes. It tracks with some random failures I’d been seeing. Being only the k8s host, I’ll just err on the side of caution and re-install it.
From the OMV one, I also grabbed the config files themselves for OMV and the NFS shares - which were my real goal, since I didn’t wanna rebuild them by hand. All in all, a good success, and I’ve skilled up a bunch on several tools.
Long live qemu, down with LVM!