Three Days of Dead Records and What They Actually Meant
Three Days of Dead Records and What They Actually Meant
The DNS migration was scheduled for 2 a.m. on a Wednesday in late August, which in Singapore means the air conditioning in the data center hallway is fighting a losing battle against the humidity that seeps in every time the loading bay door opens. A technician who worked that night recalled the way the server room’s floor tiles felt slightly sticky underfoot — a detail that mattered more than most would expect, because it meant someone had been in and out of the cooled space more times than usual that shift, and the climate control was struggling to catch up.
The company in question ran about 1,200 client accounts on a shared cPanel environment. The migration was to a custom-built dashboard that their in-house team had spent fourteen months developing. The theory was sound: better resource isolation, faster provisioning, a control panel that didn’t look like it had been designed in 2003. The reality involved three days of broken DNS records and a support queue that at one point stretched to 47 unresolved tickets, most of them from clients who had woken up to find their websites pointing at nothing.
A Hostel Booking Platform in Langkawi
The first mistake was timing, but not in the way most postmortems frame it. The team had chosen August because it was supposed to be a slow month for their client base — mostly small e-commerce stores and local business directories in Singapore and Malaysia. What they hadn’t accounted for was that August is also peak season for several of their clients who ran travel-related sites. One of them ran a hostel booking platform in Langkawi that did 60 percent of its annual revenue between June and September. That client’s DNS broke at 3:14 a.m. and wasn’t restored until 11:20 a.m. The support ticket logged a loss of 22 confirmed bookings during those eight hours.
The second issue was more structural. The custom dashboard handled DNS differently than cPanel did. In cPanel, DNS zones were managed as flat files with predictable naming conventions. The custom system used a database-backed approach with a caching layer that the team had tuned for performance but not for the edge case of mass migration. When 1,200 accounts were imported in a single batch process, the caching layer effectively treated every new zone as identical to every other one. It started serving stale records from the first batch to clients in later batches — a client who had migrated at 2:30 a.m. might resolve to the IP of a client who had migrated at 2:15 a.m., or to no IP at all.
“We saw the logs and thought it was a propagation delay,” said the lead systems engineer, who asked not to be named. “It took about an hour to realize the records were being served correctly but to the wrong clients.”
The Temperature Climbed by Three Degrees
By 6 a.m., with the support queue climbing past 20 tickets, the team had two options. One was to roll back the entire migration to cPanel, which would restore DNS for everyone but would also mean restarting the entire process from scratch — a timeline the company’s management had already told clients would be measured in days, not weeks. The other was to build a fix on the fly and deploy it without taking the system offline.
A hosting engineer who worked the overnight shift described the decision this way: “Rolling back felt like admitting failure, but not rolling back felt like we were gambling with client data. Neither option was great.”
What actually happened was a hybrid solution that nobody had planned for. The team wrote a script that bypassed the caching layer entirely for DNS queries, routing them directly to the database. It worked, but it also meant that every DNS query hit the database directly — a design that would never survive normal traffic loads. They ran it for 72 hours while rebuilding the caching layer from scratch. During that time, database load spiked to 340 percent of its normal peak. The server room’s temperature climbed by three degrees Celsius over the course of the first afternoon.
The fix made it through on the third day, at 4:50 p.m. on a Friday. By then, 14 clients had requested migration to other providers. One of them, a travel agency in Johor Bahru, had lost what their owner described as “enough business to cover the quarter.” The hosting company offered three months of free service to affected clients. Some accepted. Some didn’t.
The Wrong Glue Records
One particular detail that most postmortems would miss: the DNS failures weren’t uniform. They affected A records more than MX records, and they affected clients whose domains had been registered through the company’s own registrar more than clients who had brought their own domains. This wasn’t a technical accident — it was a consequence of how the import script handled nameserver delegation.
Domains managed internally had their nameservers updated automatically during the migration. Domains managed externally required manual updates that the migration script didn’t trigger. The result was a split where roughly 300 clients correctly showed the new nameservers but the records behind them were wrong, while another 200 clients never got the nameserver update at all and continued pointing at the old server — which was still running but had been taken out of production.
“That was the part that confused everyone the longest,” said a support technician who handled the escalation queue. “You’d have two clients sitting next to each other in the same office, both on our network, and one site worked and the other didn’t, and the difference was just where they’d registered the domain two years earlier.”
A forum post on a Singapore hosting community board from that week captured the frustration well — it was written by someone who had spent three hours on the phone with support, only to realize the issue was that his domain registrar in the United States hadn’t propagated the new nameserver details because the migration team had sent the wrong glue records. The thread got 14 replies, most of them from other clients reporting similar stories.
Eight Minutes Per Client, Fourteen Hours of It
Between Thursday morning and Friday evening, two engineers manually corrected DNS records for 90 of the most critical client accounts. Each one required logging into the database, identifying the correct zone data, cross-referencing it against a cPanel backup that was still stored on a server in the same rack, and then writing the corrected records by hand. The process took roughly eight minutes per client. Over the course of 14 hours, the two engineers corrected about two-thirds of the affected accounts this way. The remaining third were handled by the automated fix that went live on Friday afternoon.
The manual process had an unintended side effect. Because the engineers were working from cPanel backups that were three days old, any DNS changes that clients had made between the backup and the migration were lost. One client — a photo studio in Chinatown — had changed their MX records to point to a new email provider on the Tuesday before the migration. That change was overwritten when the manual fix was applied. The studio didn’t notice for another four days, by which point they’d missed about a week’s worth of email inquiries. They left the hosting provider within two months.
The Low-Priority Warning
The company eventually published a postmortem, but it was the kind that names the technical root cause and skips the messy parts. Among the details it omitted: the fact that the lead engineer had warned about the caching layer issue during a code review six weeks before the migration, and that the warning had been filed as “low priority” because the test environment didn’t replicate production traffic volumes. Or that the migration had been pushed forward by two weeks because the company’s CEO had told a client at a networking event that the new dashboard would be live by September 1.
A former employee who left the company four months after the migration described the aftermath this way: “The dashboard actually worked well after that. But trust took longer to rebuild than the code took to fix. Some clients never really came back from it.”
For the team that lived through those three days, the memory is less about the technical fix and more about the humidity in that hallway, the database load graphs climbing into yellow then red, the phone ringing at 2 a.m. and again at 3 a.m. and again at 4 a.m. The travel agency that left eventually found a provider that still used cPanel. The photo studio in Chinatown moved to a cloud platform and never looked back. The company itself kept running, with a dashboard that worked exactly as designed — for the clients who stayed.
📷 Photos: Stephen Phillips – Hostreviews.co.uk (Unsplash), Quilia (Unsplash)