The preventive maintenance routine that catches hardware failures before they become outages — and the schedule that actually gets followed.
Most server failures aren't sudden. A disk degrades for weeks before it fails outright. A power supply runs hotter and hotter before it finally gives out. A RAID array loses a drive that nobody notices because the alert email went to an inbox no one checks anymore. The pattern in almost every "unexpected" outage we get called in for is the same: the warning signs were there, but nothing was watching for them.
A maintenance checklist only works if it's short enough to actually get done every month, not a 40-item spreadsheet everyone ignores after the second week. Here's the version we hold our own support contracts to.
Every item on this list earns its place by catching a specific, common failure mode. A longer checklist looks more thorough on paper, but in practice it's the first thing skipped when a team is busy — and a maintenance routine that gets skipped provides exactly the same protection as no routine at all.
If your team is already stretched thin, the honest next question is whether this needs to run in-house at all, or whether it's a better fit for a support contract with defined SLAs and someone else holding the checklist.
Infrastructure Support contracts cover this checklist and more, under a defined SLA.