Dev.to WebDev 🛠 Dev 👁 0 📖 3 min read

What is Uptime Monitoring? A Practical Guide for Developers

Uptime monitoring means automatically checking whether your website, API, or service is accessible — and alerting you when it's not. But the full picture is more nuanced than that. Here's what you actually need to know a

Uptime monitoring means automatically checking whether your website, API, or service is accessible — and alerting you when it's not. But the full picture is more nuanced than that. Here's what you actually need to know as a developer or DevOps engineer.

The Simple Definition

Uptime monitoring is the practice of sending regular automated requests to your service and verifying that it responds correctly. When it doesn't, you get an alert.

Most uptime monitors work like this:

  1. Every minute (or every 30 seconds), make an HTTP GET request to your URL
  2. Check that the response status is 200 (or your expected code)
  3. Optionally check that the response body contains expected content
  4. If any check fails, wait for one more confirmation to rule out transient network issues
  5. Fire an alert via email, Slack, or webhook

Why Uptime Monitoring Matters

The cost of undiscovered downtime

Without uptime monitoring, you find out about downtime one of two ways:

  • A user tells you
  • You happen to visit your own site

Both are terrible. Users who experience downtime and have to tell you about it are less likely to come back. Every minute of undetected downtime has a measurable cost.

What 99.9% uptime actually means

Uptime % Downtime per year Downtime per month
99% 3.65 days 7.2 hours
99.9% 8.7 hours 43.8 minutes
99.95% 4.4 hours 21.9 minutes
99.99% 52.6 minutes 4.4 minutes
99.999% 5.26 minutes 26 seconds

Most hosted apps on managed infrastructure (AWS, GCP, Heroku) target 99.9% uptime. To achieve that, you need to know when failures happen.

Types of Uptime Monitoring

1. HTTP/HTTPS Monitoring

The most common type. The monitor sends an HTTP GET request to your URL and checks:

  • Status code (usually expecting 200)
  • Response time (should be under a threshold, e.g., 3 seconds)
  • Response body (optional keyword match)

2. TCP Port Monitoring

For services that are not HTTP — databases, SMTP servers, custom protocol servers. The monitor checks that a TCP connection can be established on a specific port.

3. Ping Monitoring

ICMP ping to a host. Verifies the host is responding at the network layer, independent of any application layer.

4. SSL Certificate Monitoring

Checks your HTTPS certificate's expiry date and alerts you before it expires. Certificate expiry is a top cause of preventable outages.

5. Heartbeat/Cron Monitoring

Inverted monitoring: your process (cron job, background worker) sends a ping to the monitoring service on every successful run. If the service does not receive a ping within the expected window, it alerts you.

Single-Region vs Multi-Region Monitoring

This is the most important concept for production monitoring.

Single-region monitoring checks your service from one location. If that location experiences a network issue, you get a false alert even though your service is fine.

Multi-region monitoring checks from multiple global locations simultaneously. It only alerts when multiple locations agree the service is down. This dramatically reduces false positives.

Vigilmon uses multi-region consensus monitoring: checks run from multiple regions, and an alert only fires when multiple regions confirm the failure.

Setting Up Your First Uptime Monitor

Step 1: Add a health check endpoint

Add a dedicated /health endpoint:

// Express.js
app.get('/health', (req, res) => {
  res.json({ status: 'ok' });
});
# Flask
@app.route('/health')
def health():
    return {'status': 'ok'}, 200

Step 2: Choose check frequency

  • 1 minute — standard for production services
  • 30 seconds — for high-availability, latency-sensitive services
  • 5 minutes — acceptable for less critical services

Step 3: Configure alerts

At minimum:

  • Email — for async notification
  • Slack — for immediate team visibility

What to Monitor

Asset What to check Frequency
Main website /health returns 200 1 min
API endpoints Critical paths return expected status 1 min
Background workers Heartbeat ping from the worker Per job cycle
SSL certificates Days until expiry Daily
TCP services Port accessible 1 min

Common Mistakes in Uptime Monitoring

Monitoring from only one region. A single-region failure looks like your service is down when it is not.

Monitoring the homepage instead of a health endpoint. Homepages include third-party scripts, CDN assets, and dynamic queries that can trigger false alerts.

Not testing the alert path. Add a monitor for a URL you control, then take it offline briefly to confirm you actually receive the alert.

Setting check intervals too long. 5-minute intervals mean up to 5 minutes of undetected downtime.

Not monitoring SSL. Expired SSL is the most common preventable outage type.

Getting Started

Vigilmon provides free multi-region uptime monitoring with:

  • HTTP, TCP, and ping monitoring
  • SSL certificate expiry alerts
  • Multi-region consensus (no false alerts from single-region blips)
  • Instant alerts via email, Slack, or webhook
  • Public status pages

Setup takes about 2 minutes: add your URL, set your alert channel, and you are monitoring.

Try Vigilmon free: vigilmon.online

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.