I moved a 167.9 GiB directory off a full disk and freed zero bytes. It also consumed 172 GiB on the destination.
That tree was hardlinked. The media library name and the seeding payload name pointed at the same inodes. The import created the library entry with ln instead of a copy. rsync --remove-source-files copied the data and deleted one of the two names. The inode kept its other name. Its link count stayed above zero, so every byte stayed on the original disk.
The tell
This is the part I’d want to know. The operation reports success. The progress bar is honest, because bytes really are being written.
dfon the source stays flat. Over 90 seconds of a large move, the delta was exactly 0.iostat -dxshowsw/s = 0.00on the source disk. Reads climb, writes are nil. A real move writes to the source too, because deleting files updates metadata.
Put those two together and the source isn’t losing anything. Kill the job there. The damage so far is a duplicate, not a loss.
Survey before you move
One line would have saved me. It’s a link-count census, not a size census:
find <dir> -type f -printf '%n %s\n'
%n is the link count. Anything above 1 has a second name somewhere. A mover that ignores that duplicates instead of relocating. On my array the split was stark:
| Tree | Size | Files with link count > 1 |
|---|---|---|
| Movies | 292.7 GiB | 19 of 20 |
| Shows | 871.2 GiB | 388 of 388 |
| Staging, music, downloads | — | 0 |
So about 620 GiB of that library can’t be moved one-sidedly. Those trees look identical in a file browser and behave nothing alike.
Here’s the rule I apply now. A bulk mover has to abort on any tree with link counts above 1. It shouldn’t treat them as ordinary files.
The repair
I didn’t want to re-transfer anything. The bytes were already correct in two places. I wanted the link back and the surplus copy gone.
- For each moved file, find its surviving twin by exact basename plus exact size. 102 of 102 matched. Size alone is too weak and path is useless after a move. Together they were unambiguous across the whole set.
ln -fthe library name back to the twin’s inode.- Assert with
stat -c %ithat both paths now report the same inode number. - Delete the stray copy on the destination disk.
Step 3 is what makes the rest safe. Skip it and a mismatched pair means deleting the only copy. Deleting on a match is one program. Deleting on a verified identity is another.
The bonus 234.7 GiB
While cleaning up I hit the adjacent classic. A file gets deleted from the filesystem. A long-running process still holds it open, so the space never comes back. Everybody says restart the process or reboot. You don’t have to:
truncate -s 0 /proc/<pid>/fd/<n>
That truncates the open descriptor in place. 234.7 GiB came back instantly with the consuming service still up. Guard it first. Check the descriptor’s size against what you expect to reclaim. /proc/<pid>/fd/ is full of descriptors you do not want to zero.
One more, and it cost me a session. pkill -f <pattern> over SSH kills your own shell. The remote command line contains the pattern you’re matching. So pkill -f rebalance matches the bash -c that invoked it. My session died with no output. The target script kept running. Exact inverse of what I asked for. Use explicit PIDs, or bracket the pattern so it can’t match the invocation.