From a site that crashed on campaign day to planned scaling
- Situation
- During the two big campaign periods of the year, the site stopped responding in the first hour of peak traffic. Server capacity had already been increased, yet the problem kept recurring.
- Approach
- Zabbix metrics showed the bottleneck was not CPU but the database connection pool. Web, cache and database tiers were moved onto separate servers and Redis was deployed with replication. Capacity increases were scheduled to apply three days before each campaign.
- Outcome
- The site ran without interruption in the following campaign period, and because resources were scaled back afterwards the extra capacity never became a permanent cost.