24/7 support with an SLA
An on-call engineer at night, at weekends and over holidays. Thirty-minute P1 response written into a contract, not promised in an email.
What's included
- Round-the-clock cover, holidays included
- 30-minute P1 response at any hour
- SLA with remedies for breach
- Escalation matrix and second responder
- Incident reviews and postmortems
- NDA and security team sign-off
- Monthly availability reporting
Round-the-clock support isn't for everyone. If your site serves an audience that works nine to five and an overnight outage costs you a couple of missed enquiries, paying for a night shift makes no sense — a standard plan covers it. But some projects treat night exactly like day: services handling payments, applications with users across time zones, B2B platforms that have committed to uptime in their own customer contracts.
The difference between a round-the-clock plan and a standard one isn't the amount of work. It's who picks up the phone and when. Monitoring runs continuously on every plan and will faithfully record a failure at three in the morning. The question is whether that alert wakes a human being or waits politely for office hours.
How we work
- Onboarding and a criticality map. We establish what specifically constitutes a first-level incident. It differs by project: payments for one, authentication for another, an overnight data feed to a partner for a third.
- On-call rotation. Engineers go into a rota with a second line for when the first responder is unreachable. One person with no backup isn't twenty-four-hour cover, it's a hope.
- Escalation matrix. Written down: who responds, when the second engineer joins, at what point we wake you, and which of your people we call if the problem sits on your side.
- Alerts worth trusting. Thresholds are calibrated against your traffic. A couple of false alarms at 3am and the on-call engineer starts ignoring notifications — which is worse than having none.
- Response and recovery. P1 is picked up within thirty minutes at any hour. The first job is restoring service, with a workaround if necessary; root cause analysis comes afterwards.
- Postmortem. After every significant incident, a short written review: what happened, why, and what changes so it doesn't repeat.
What you get
Assurance that an overnight failure won't sit untouched until morning. Response times are committed in an SLA carrying remedies for breach, rather than agreed in a thread somewhere. There's a rota and a defined escalation path, so one engineer's holiday or illness doesn't leave you uncovered.
Then there's what corporate clients usually require: work under NDA, access approval with your information security team, connections restricted to agreed addresses, and a written incident history for internal audit. Where requirements go beyond the standard, or several projects are involved, we assemble bespoke terms.
If round-the-clock cover is more than you need, the lighter plans on the technical support page respond during business hours at a considerably lower price. For scheduled routine work unconnected to incidents there's the maintenance retainer. Monitoring can also be set up on its own — see server monitoring setup — without any on-call component.
Timeline
Onboarding takes a week or more: we have to learn the project, stand up or reconfigure monitoring, calibrate thresholds and agree the escalation matrix. Going on call sooner would be pointless, because an engineer who doesn't know the system won't fix it at 3am. The contract runs monthly with automatic renewal, and the SLA is signed as a separate schedule.
A typical scenario
Here's a representative case: at three in the morning a disk fills up and the service stops recording orders, while the site itself still loads normally. Monitoring sees the error rate climb and pages the on-call engineer. They respond within thirty minutes, free space, restart the failed workers and restore order intake before the morning peak. Working out why logs grew faster than usual goes into scheduled work, and the report records the incident as closed within SLA. On a business-hours plan the same failure would have meant eight hours without orders.