Quote from: Danihel Txechescu on October 23, 2023, 09:30:53 PMAzul.
The server underwent an unexpected service discontinuity today, 2023-10-23, at about 19:28 Talossan Time, which lasted for about 1 hour, until 20:26 Talossan Time. The affected services were only the web server, which currently hosts Wittenberg, the main web space, the Wiki, and other web services; email remained operational throughout the entire time; backups are still safe.
The nature of this disruption was a hard drive getting filled up to 100%. This should never happen; why this happened was because it was not actively monitored. To prevent such an problem from occurring again, a dedicated lookup/monitor on hard disk utilization has now been put in place.
I apologize for the inconvenience.
Danihel Txechescu
MinTech
Today the server had exactly the same kind of problem, even with monitoring (though not as active during the weekend). I can only think of a recent upgrade to Wordpress that's changed things with the database service.
The cause behind these outages is the system losing its available scratch space, due to an aggressive saving of binary logs. While these are a good safety net, we have other kinds of database backups, so I will have this feature disabled entirely as it's causing too much pain.
Apologies again for this outage. This should not happen again once this feature is disabled. We'll see in two weeks' time.
Danihel Txechescu


