Servers offline

03 June 2016 9:56 AM

Three servers are offline. It appears to be a problem at the datacentre, which is being investigated.

Affected servers:

Daenerys
Lannister
Email 4

10:03 AM Hetzner has confirmed a networking error in their Samrand DC which is being worked on.

10:05 AM Hetzner seems to have fixed the problem, all servers are now accessible.

Daenerys inaccessible

2 May 2016 5:38 PM Daenerys (197.189.230.226) is currently inaccessible due to what appears to be a network problem at Seacom / Hetzner. I’m still waiting for feedback from Hetzner.

5:45 PM Daenerys is accessible again, but connections are slow. Still waiting for feedback from Hetzner.

5:56 PM Inaccessible again. Still waiting for Hetzner to respond.

6:02 PM No feedback from Hetzner yet (email or phone) but they have updated their “Network Notices” page with the cause of the problem: a DDOS attack:

  • DDOS – Network connectivity – Johannesburg data centre (2016-05-02, ATTENDING)

    Start: 2016-05-02 5:53:48 SAST
    Resolved: TBA
    Status: attending
    Point of impact: Truserv and Co-Location customers
    Symptoms: Truserv and Co-location customers will have intermittent connectivity to their servers.
    Cause of problem: DDos
    Estimated time of repair: TBA
    Attending: Hetzner Engineers

Source: Hetzner Network Notices

6:20 PM Daenerys appears to be 1oo% accessible now.

Intl. connectivity: CloudFlare

Clients using CloudFlare for their website might be affected by an outage affecting two undersea cables and CloudFlare.

CloudFlare status: https://www.cloudflarestatus.com

More information: http://www.techcentral.co.za/seacom-wacs-problems-hit-sa-internet/62649/

Arryn: offline

8:50PM There is a problem in the Hetzner Datacenter, which they are working on. Arryn is offline until they get it fixed, which hopefully will not be too long.

9:30PM Hetzner has fixed the problem – I’m waiting for a RFO to find out exactly what their problem was.

Update: Hetzner says that a “network cable was loose”.

Arryn: apache problem

11:30AM All sites on Arryn are displaying a 500 server error. We’re working on the problem.

11:40 We’ve found the problem, and are working on a fix

11:45 Server is being rebooted after the fix.

11:57 Server is up after reboot, but the problem remains. Emergency ticket has been submitted to CloudLinux, who have logged into the server to fix the problem with their software.

12:01 The problem has been fixed. Apologies for the downtime.

Brienne: server offline

2:40AM Scheduled maintenance was completed on Brienne at 9:10PM on Friday 28 November, but the server failed to reboot. Hetzner’s datacenter technicians are still working on the problem, almost 6 hours later. If they fail to get the server up and running, I will ask them to remove the harddrives so that we can move all the data to another server.

3:45AM the Hetzner technicians have been unable to boot the server, even after swapping out multiple bits of hardware. We will attache the hard drives to another server, and start copying the data to another one of our servers. Unfortunately, this will take many hours, during which time the sites formerly hosted on Brienne will be offline.

4:55AM The technicians have attached the hard drives to another server, and we have started copying the data.