Matt Scott

Your uptime percentage isn't calendar math

Every uptime percentage calculator assumes continuous 24/7 coverage. Real monitoring dashboards don't work that way: paused monitors and unmonitored gaps get excluded, not counted, which is why the two numbers rarely match.

An uptime percentage calculator takes a target like 99.9% and converts it into a downtime budget: 43 minutes and change over a 30-day month. That's correct arithmetic. It's also not how the percentage on your actual monitoring dashboard gets calculated, because a calculator assumes continuous coverage for the whole period and your dashboard doesn't have that.

The formula everyone gives you

Uptime percentage is: (monitored time minus downtime) divided by monitored time, times 100. Feed 99.9% into any of the SLA calculators built for this and you get the standard table: about 43 minutes a month, 8.75 hours a year. That math is fine as far as it goes. It goes wrong the moment you try to reconcile it with a number your monitoring tool actually reported, because the calculator's total time and your dashboard's monitored time aren't the same quantity.

What actually gets left out of the denominator

A calculator has no concept of a paused monitor, a site added mid-month, or a check that failed from one location but passed from every other. A real monitoring platform does, and it has to decide what to do with that time. WebPixie's own answer is explicit: time before a monitor existed, or while it was paused, is excluded from the total rather than counted as either up or down, and, where cross-location verification is available, a verification check that fails from one location but passes from the others is recorded and excluded rather than counted as downtime, so a brief routing flicker doesn't distort the figure. Both are described directly in WebPixie's own uptime percentage FAQ.

That's a real design decision, not a rounding error, and it's one every serious monitoring tool has to make the same way: the alternative is treating time nobody actually checked as if it were confirmed up, which is a worse kind of dishonesty than excluding it.

Where the two numbers stop agreeing

Take a monitor paused for 3 days mid-month for planned maintenance, with 20 real minutes of downtime somewhere in the remaining 27 days. Your dashboard divides that 20 minutes against the 27 days it actually watched: 38,880 monitored minutes, which comes out to 99.949%. Feed the same 20 minutes of downtime into a generic calculator that assumes the full 30-day month was covered, and you get 99.954%, a number that looks slightly cleaner because it's quietly crediting 3 days of coverage that never happened.

Neither number is wrong. They're answers to different questions: one is asking what happened during the time you actually watched, the other is assuming you watched the whole period. The gap here is small because the example is small; it gets larger the longer a monitor sits paused or the more location-based verification excludes.

Why this is the right tradeoff anyway

The alternative to excluding unmonitored time is worse in both directions. Count it as down and a routine maintenance window tanks your SLA number for something nobody could have observed either way. Count it as up and you're reporting confidence about a period you have zero data for. Excluding it is the only option that keeps the percentage meaning what it claims to mean: the share of observed time your site actually responded, nothing about the time nobody was watching.

  • A monitor paused for maintenance doesn't drag your SLA down, and it shouldn't get credited as clean uptime either
  • A monitor added on day 15 of the month reports against 15 days of data, not 30
  • A check that fails from one region but passes from the rest gets excluded rather than counted as an outage, so route flakiness doesn't read as your site being down

Why the comparison between two tools' numbers gets messy

This is also why an advertised '99.99% uptime' from one monitoring vendor and a '99.95%' from another aren't automatically a fair comparison. If one excludes unmonitored gaps and the other quietly counts them as up, or one runs cross-location verification before confirming an outage and the other doesn't, the two percentages were computed by different rules before either number ever reached a decimal point. Before comparing two uptime figures, the actual question isn't which one is higher, it's whether both tools excluded the same kinds of time the same way. Most vendor pages don't spell that out, which is part of why practitioners on forums like Hacker News routinely argue about whether a stated nines figure means anything at all. The practical fix is to ask the vendor directly whether paused time, unmonitored gaps, and cross-location retries are excluded from the denominator or counted as uptime, before treating two numbers as comparable.

Where a calculator is still the right tool

None of this makes the calculator useless. Converting a target percentage into a plain downtime budget, in minutes per day or hours per year, is exactly the question a calculator answers well, and it's the right first step for deciding what target to even negotiate into an SLA. See WebPixie's own uptime and downtime calculator for that conversion, including the reverse direction: turning an observed outage back into the percentage it costs you. Just don't expect that number to match your dashboard's reported uptime down to the decimal once real gaps and exclusions are in the picture; they're measuring related but different things.

The actual decision that matters more than either number, picking a check interval that catches problems fast enough to matter, is a separate question. See how to pick a check interval by what failure mode you're actually trying to catch.