Expanding the lab NAS that holds Milo-Ark and the rest of the house data: DS1823xs+ from an aging 8×8 TB RAID 5 pack toward 8×24 TB RAID 6 in the chassis, with DX517 expansion as park/bulk. This post is the storage rebuild. Model inventory stays on the Milo-Ark page.
Pool 13×24 TB RAID 5 · optimizing
Free chassis3×24 TB (bays 4–6)
Emptybays 7–8 · +2×24 TB soon
Pool 37×8 TB bulk · Healthy
Unverifiedcleared (HDD_db)
32 GBRAM arrives tomorrow
August 4 afternoon. Cutover held: HB restore OK, PARK emptied and removed, bulk data back on main. Chassis has 6×24 TB seated (3 in Pool 1 RAID 5 still optimizing; 3 free in bays 4–6). Bays 7–8 empty pending two more IronWolf Pros. Expansion: Pool 3 = seven 8 TB bulk (all Healthy after Synology_HDD_db). One free 8 TB on DX517-1. Next: wait Pool 1 Healthy → Add Drive 4–6 → RAID 6 (OK on 8 GB); install 32 GB ECC tomorrow before growing past 108 TB / full 8-wide. Final: chassis 8×24 RAID 6 + expanders 8×8 RAID 6.
Current drive map (2026-08-04)
Chassis: 1–3 = Pool 1 (main); 4–6 = free 24 TB ready to add; 7–8 empty. DX517-1/2 hold 8 TB bulk (Pool 3). M.2 = SSD cache.
Location
Drives
Role
Chassis 1–3
3×24 TB IronWolf Pro
Pool 1 RAID 5 (optimizing)
Chassis 4–6
3×24 TB IronWolf Pro
Free — add to Pool 1 when Healthy
Chassis 7–8
Empty
+2×24 TB in ~1 week → 8-wide RAID 6
M.2
2× Samsung 990 PRO 2 TB
SSD Cache Group 1 on Volume 1 (already live)
DX517-1 bay 3
1×8 TB
Free (keep off main 24 TB pool)
DX517-1/2
7×8 TB
Pool 3 bulk · Healthy
Final target layout
Location
Drives
Array
Role
DS1823xs+ chassis
8×24 TB IronWolf Pro
RAID 6
Primary (Milo-Ark + critical) · needs 32 GB RAM for single vol >108 TB
DX517-1 + DX517-2
8×8 TB IronWolf
RAID 6
Bulk / secondary across both expanders
M.2
2×990 PRO
SSD cache
Live on Volume 1 (Cache Group 1)
Path: chassis 3→6→8 ×24 TB (RAID 5 then RAID 6); expansion Pool 3 already 7×8 TB RAID 6 — add the free DX517-1 8 TB to reach 8×8 RAID 6. Do not mix 8 TB members into the 24 TB chassis pool.
Unverified drives fix (2026-08-04)
Why this exists: Synology ships a drive “compatibility” database that treats many perfectly good CMR NAS disks (and most third-party NVMe) as second-class. You get red Unverified badges, “At risk” pool chrome, and nag dialogs — even when SMART is clean and the array is fine. It is partly support-policy theater and partly a push toward their own (expensive) validated SKUs. DSM is still a strong appliance UI; the drive politics are the part that makes people hate the company.
On this box: chassis 24 TB Pool 1 members went Healthy after create. Expansion 8 TB on DX517 (Pool 3) stayed Unverified until we fixed the DB. Cosmetic until it isn’t — some wizards get pickier over time, and “At risk” is noise you should not live with.
What Synology shows: pool At risk + red Unverified on IronWolf-class disks that are otherwise fine. Compatibility theater, not a SMART failure.
What does not fully fix it
Step
Result here
support_disk_compatibility="no" in /etc/synoinfo.confand/etc.defaults/synoinfo.conf
Unblocks some create/use paths; often leaves the red badge
UI refresh / reboot alone
Insufficient if the model is missing from DSM’s drive DB
Ignoring it forever
Works until the next nag, support call, or picky DSM update
What worked: 007revad Synology_HDD_db
Community script that injects your real HDD/SSD/NVMe models (and firmware strings) into DSM’s compatible-drive databases for the NAS and expansion units, then triggers a compatibility recheck. We used v3.6.137.
Enable SSH; use an admin on the Terminal allow-list; sudo -i (root SSH login is off by default — fine).
Prefer a quiet window (not mid multi‑TB move). Resync/optimize can keep running.
On the NAS:
cd /volume1
mkdir -p /volume1/_admin_scripts && cd /volume1/_admin_scripts
curl -fsSL -o hdd_db.tgz \
https://github.com/007revad/Synology_HDD_db/archive/refs/tags/v3.6.137.tar.gz
tar xzf hdd_db.tgz
cd Synology_HDD_db-3.6.137
chmod +x syno_hdd_db.sh
# -n = prevent DSM from auto-overwriting the drive DB later (recommended)
./syno_hdd_db.sh -n
grep support_disk_compatibility /etc/synoinfo.conf /etc.defaults/synoinfo.conf
Our run on DS1823xs+ · DSM 7.3.2-86009-4 reported models found:
ST24000NT002-3N1101 (24 TB) → host DB + dx517_v7.db
Script also: backed up synoinfo.conf, disabled drive-DB auto updates (-n), ran DSM disk compatibility recheck successfully. It may flip compatibility flags around after DB inject — that is normal. Hard-refresh Storage Manager (re-login if needed). With M.2 present, reboot if badges remain; we cleared without a mandatory reboot after refresh.
Result
No more Unverified warnings. DX517 Pool 3 8 TB members show Healthy. Chassis 24 TB already Healthy.
Rollback
cd /volume1/_admin_scripts/Synology_HDD_db-3.6.137
./syno_hdd_db.sh --restore
Hard rules
Do not Initialize / Remove a pool to “clear Unverified.”
Unverified ≠ bad SMART — still check health separately.
Avoid -f / force flags unless -n failed and you know why.
After major DSM upgrades, badges can return — re-run the script or re-check confs.
This is a community tool. It fixed real friction; it is not Synology-supported. Prefer it over living with lying UI, and over buying only “approved” labels at a tax.
Synology gripe (earned): Selling a RAID appliance that gaslights standard IronWolf / IronWolf Pro disks as unverified — while charging enterprise money for the badge — is a shitty business move. DSM remains excellent ops UI; the compatibility moat is why people keep one foot out the door toward TrueNAS/UGREEN. We stayed for the appliance; we refuse to pretend the nag is “safety.” Same energy as the volume-crash path that pointed at a support ticket while mdadm still showed a clean [UUU] array — we mounted it ourselves.
Goal
Piece
Plan
Chassis
8×24 TB RAID 6 in chassis; 8×8 TB RAID 6 on DX517-1+2 bulk
Usable (8×24 R6)
~6 data disks → on the order of ~130+ TB class
Old 8 TB members
Move to DX517 shelves as bulk/secondary
Why not live re-drive
Full-member RAID 5 rebuilds on 24 TB are multi-day degraded cycles; offload/rebuild preferred
Why RAID 6 not R5+hot spare
Same ~usable capacity; dual parity covers second failure during long rebuilds. Write gap vs R5 is modest on 10GbE for bulk work.
Hardware on hand
Item
Role
DS1823xs+ (192.168.1.21)
Primary · 10GbE · production until wipe
DS1019+ (192.168.1.23)
Hyper Backup Vault target (1GbE)
DX517 ×1–2
eSATA expansion; PARK / later bulk
6×24 TB now (8 planned)
Interim path: park + chassis, then expand; RAID 5→6 conversion OK on DSM
DS1823xs+ and DX517 on the bench.Lab rack: local compute and NAS.IronWolf Pro 24 TB members in trays (serials/QR redacted).
Mistake #0: more storage than stock RAM allows
We ordered the full 8×24 TB chassis rebuild (and DX517 park) while the DS1823xs+ still had stock 8 GB ECC. On this model DSM caps a single volume at 108 TB unless the box has 32 GB RAM (200 TB ceiling with 32 GB).
8×24 RAID 6 usable lands in the ~130+ TB class — over the 108 TB gate on 8 GB.
Without the RAM upgrade, you either split volumes (worse archive UX) or cannot use the capacity you paid for as one big pool.
Fix in flight: OWC 2×16 GB ECC SODIMM ordered → replace the stock 8 GB stick for a clean 32 GB (2 slots, max).
Install RAM after the array is stable (or at least not mid-backup); third-party ECC is fine if returnable — official SKU is D4ES03-16G class.
Lesson: check DSM volume-size vs memory tables before buying a wall of 24 TB drives. Spindles without RAM headroom are half a plan.
Mistake #1: Entire System + PARK (≈1.5 days)
What we did wrong. Pre-wipe safety used Hyper Backup Entire System from the 1823 to the 1019. That mode backs up every online volume, not “whatever is left on main.”
We had already built a temporary PARK RAID on a DX517 and offloaded multi-TB trees main→PARK to shrink the rebuild. Entire System then backed up main + PARK + system.
Main “used” looked like ~4–6 TB after cull.
Vault on 1019 climbed to ~8 TB and kept growing.
UI % bar was nonsense (e.g. bytes nearly done while % showed teens).
Roughly a day and a half of 1GbE wall-clock burned before we admitted the scope bug.
1019 free space got uncomfortably tight.
Fix: discard that job, unmount PARK so it is offline to DSM, start a new Entire System task (1823Fullbackupminuspark) with only main online. Optional cleaner path next time: Data backup of main shares only (exclude PARK by selection).
Lesson (print it): offloading to another volume on the same NAS does not shrink an Entire System backup. Sideways copy ≠ out of scope. Unmount/deactivate park volumes before ES, or use a folder-scoped Data task.
Mistake #2: Safe Eject vs unmount
To exclude PARK from the next Entire System job we used Safe Eject on the park pool. That is the feature for removing an expansion shelf cleanly — not “pause this volume for backup.”
Pool disappeared from Storage Manager; DX517 disks still visible as Unverified.
mdstat showed only main md2 (8-disk RAID 5). PARK was not assembled.
mdadm --examine still showed clean RAID 5 superblocks on sata5/6/7p3 (UUID 158fdbe7:…, AAA, ~44.7 TiB) — data intact.
Recovery path (official): reboot if needed → Storage Manager → Available Pool / Detected → Online Assemble. Never Create Pool / Remove on those disks.
Manual mdadm --assemble is a known fallback; Online Assemble is preferred on a live DSM box.
Going forward: hide PARK with Unmount volume only, or leave it mounted and use a Data backup that excludes PARK shares. Reserve Safe Eject for actually pulling the shelf.
Mistake #3: crash the park volume while trying to “just unmount”
After the big move, main was only ~200 GB. We still wanted PARK offline so Entire System would stay small. DSM 7 did not show a clear Unmount on Volume 2; thrashing pool/volume menus (and residual Safe Eject muscle memory) left Volume 2 Crashed in the UI.
mdstat: md4 PARK RAID 5 still [3/3] [UUU] — not a disk failure.
LVM: vg2/volume_2 still active on md4.
btrfs: same UUID on /dev/mapper/vg2-volume_2 and cachedev_0 (SSD cache path).
Only /volume1 was mounted; PARK data invisible to File Station until remount.
Proof: RO mount of cachedev_0 showed park/miloshare/{forge,Milo-Ark,p24-backup-*}.
Fix: mount -t btrfs … /dev/mapper/cachedev_0 /volume2 then reload scemd. df: volume2 ~24 TB used.
Lesson: UI “Crashed” ≠ dead array. Check mdstat + lvs + blkid before Remove/Repair/Create. And if main is already ~200 GB, stop trying to hide PARK for Entire System — use a Data backup of main + .dss instead.
DSM’s happy path after the crash was basically open a support ticket. We skipped the queue: proved md4 [UUU], mounted cachedev_0, data intact in minutes. Support is fine when you’re stuck; it is not required to remount a healthy RAID Synology scored as theater.
Mistake #4: one copy while reshuffling the archive
During this rebuild window we are living too close to the edge on redundancy. In particular, Time Machine / machine backups for the laptop, M4 Max, M3 Ultra, and M5 Max were nuked off the NAS to free space and shrink the migration surface.
PARK holds offloads; main holds residual; 1019 is meant to hold an Entire System of main — but until ES is Successful, several trees exist in only one place.
Killing workstation backups removes the easy rollback if a Mac or a bad copy corrupts something mid-move.
Hyper Backup of the NAS is not a substitute for endpoint backups; endpoint backups are not a substitute for a verified NAS backup before wipe.
Lesson: free space by deleting replaceable caches first; keep at least one verified second copy of anything unique before a chassis wipe. Re-enable TM/machine backups (with quotas) as soon as the new array is Healthy — do not wait for “someday when Ark is perfect.”
Leave PARK mounted (bulk already there). Export fresh .dss.
Hyper Backup Data task: main shares only (~200 GB) → DS1019+. Optional Applications. Wait for Successful.
Spot-check restore paths / version list.
Only then: delete old main pool; install 24 TB members in chassis.
Build RAID (start RAID 5 if <4 disks in chassis, else RAID 6); restore; cut over shares.
Bring remaining 24 TB in; change type / expand to full chassis RAID 6.
Old iron → DX517 bulk. Remount PARK only as needed to drain or retire it.
Install 32 GB ECC before expecting a single volume above the 108 TB DSM gate.
Migration state
Step
Status
ES #1 (included PARK)
Discarded after ~1.5 days
Safe Eject → Online Assemble
Recovered (morning)
Main → PARK ~3.78 TB move
Done (~280 MB/s peak)
Volume 2 “Crashed” / unmount mess
Recovered — mount cachedev_0 → /volume2
Main residual
~198 GB
PARK bulk
~24 TB used (forge / Milo-Ark / p24-backup …)
Data HB main-only → 1019
Done + integrity OK
Wipe old Pool 1 / chassis swap
Done (3×24 TB in)
New pool RAID 5 (3×24)
Created · optimizing
HB restore (config + shares + apps)
Succeeded
PARK deleted · data on main
Done
Unverified / HDD_db
Cleared
Pool 1 optimize
In progress (blocks Add Drive)
Add bays 4–6 → RAID 6 (6×24)
After Healthy · OK on 8 GB RAM
32 GB ECC install
Tomorrow — before >108 TB / 8-wide
+2×24 TB → bays 7–8 · 8×24 R6
~1 week
Endpoint backups (Mistake #4)
Re-enable with quotas
Side notes
Unverified IronWolf badges: support_disk_compatibility="no" on live + defaults confs (doesn’t always clear the label; pool use still OK). Optional later: community drive-DB script after backup idle.
M.2 SSD cache: built-in slots; 2× Samsung 990 PRO already configured as SSD Cache Group 1 on Volume 1 — no M2D20 card needed.
SSH: admin user must be on Terminal allow-list; root has no direct SSH login.
Safe Eject hides a pool until Online Assemble.
Volume Crashed but md UUU: check lvs/blkid; PARK btrfs may be cachedev_0 — mount to /volume2 before any Remove.
Config: Data HB ≠ full settings; always export .dss before wipe.