uptime-monitor
A small single-container Python 3 uptime monitor.
- Monitors a configurable list of HTTP/HTTPS endpoints.
- Each endpoint independently specifies: scheme (http/https), port, the DNS server used to resolve it, and the expected HTTP status code.
- DNS resolution never uses the system/router-provided resolver — every lookup is sent explicitly and only to the DNS server configured for that endpoint (see "How DNS enforcement works" below).
- Checks are spread evenly across
interval_secondsrather than all firing at once: with N endpoints, one is checked everyinterval_seconds / Nseconds, rotating through all of them (see "Scheduling" below). - If any endpoint fails, all endpoints are immediately re-checked once, instead of waiting for their normal rotation slot.
- Writes every check result to an InfluxDB 2.x bucket (
uptime) over HTTPS, with certificate verification disabled for InfluxDB specifically (self-signed cert). - Everything — endpoint list, InfluxDB URL/token/org/bucket, intervals — lives in one YAML file that's bind-mounted into the container. Edit it and restart the container; no image rebuild required.
Setup
cp config.example.yaml config.yaml
# edit config.yaml: fill in your InfluxDB url/token/org, and your endpoints
docker compose up -d --build
Editing endpoints / InfluxDB settings later
vim config.yaml
docker compose restart uptime-monitor
No rebuild needed — config.yaml is bind-mounted (docker-compose.yml),
never baked into the image (see .dockerignore).
Config reference (config.yaml)
influx:
url: "https://influxdb.example.local:8086"
token: "..." # InfluxDB 2.x API token with write access
org: "..."
bucket: "uptime"
verify_ssl: false # keep false: influx uses a self-signed cert
check:
interval_seconds: 600 # each endpoint checked once per this many seconds,
# spread evenly across N endpoints (see Scheduling below)
timeout_seconds: 10 # default per-check timeout
immediate_recheck_delay_seconds: 5 # pause before the failure-triggered recheck
endpoints:
- name: "my-service"
host: "service.example.com"
scheme: "https"
port: 443
dns_server: "1.1.1.1" # required, per-endpoint — no default DNS is ever used
expected_status: 200
path: "/" # optional, default "/"
timeout_seconds: 5 # optional, overrides check.timeout_seconds
verify_tls: true # optional, set false if THIS endpoint has a self-signed cert
Add/remove/edit entries under endpoints: freely.
How DNS enforcement works
For each check:
- A fresh
dnspythonresolver is created withconfigure=False(so it never reads/etc/resolv.conf) and pointed at only the endpoint'sdns_server. - That resolver performs the A-record lookup.
- The resulting IP is pinned for the single outgoing HTTP(S) request via
a scoped patch of
socket.getaddrinfo— sorequests/urllib3 connects directly to that IP instead of resolving the hostname itself. The hostname in the URL is left untouched, so theHostheader and TLS SNI/certificate checks behave exactly as they normally would.
At no point is the container's own (router-provided) DNS resolver consulted for endpoint hostnames.
Literal IPs in host: if host is already an IPv4/IPv6 address (e.g.
192.168.0.2), DNS resolution is skipped entirely — the monitor connects
to it directly, and dns_server is optional/ignored for that endpoint.
(Without this, a literal IP would be sent as a DNS query to whatever
dns_server you configured, which returns NXDOMAIN — it is not a valid
hostname — and the check would always fail.)
What gets written to InfluxDB
Measurement uptime_check, one point per endpoint per check:
- tags:
endpoint,host,scheme,port - fields:
success(0/1),expected_status,status_code(if a response was received),response_time_ms,resolved_ip,error(if any)
Example Flux query:
from(bucket: "uptime")
|> range(start: -24h)
|> filter(fn: (r) => r._measurement == "uptime_check")
|> filter(fn: (r) => r._field == "success")
Scheduling
Endpoints are not all checked together and then slept on as a batch.
Each endpoint gets its own slot, spread evenly across interval_seconds:
with N endpoints configured, one is checked every
interval_seconds / N seconds, rotating through all of them. So
interval_seconds: 3600 with 6 endpoints means one endpoint is checked
every 10 minutes — each individual endpoint is still only checked once
per hour, but the 6 checks are spread across that hour instead of firing
all at once, then waiting an hour.
Example with 3 endpoints and interval_seconds: 600 (so ~200s apart):
t=0s endpoint A checked
t=200s endpoint B checked
t=400s endpoint C checked
t=600s endpoint A checked again (600s after its previous check)
t=800s endpoint B checked again
...
Failure / immediate recheck behavior
- If any single endpoint check fails, all endpoints are checked once,
immediately, after
immediate_recheck_delay_seconds— instead of waiting for their normal rotation slot. - The rotation is then re-spread evenly from that point, so it doesn't bunch endpoints back up onto the same slot, and resumes normally.
- If the outage continues, that endpoint's next regular slot will detect it again and trigger another single immediate recheck of everyone — it does not hammer the endpoint in a tight loop.
Local dev (without Docker)
pip install -r requirements.txt
CONFIG_PATH=./config.yaml python -m app.main