How an incident is opened
1
A check fails
A scheduled check fails from one probe location. No incident yet: it could be a transient problem.
2
Confirmation checks run
About one second after each failure, the check is repeated from a different probe location, up to three attempts in total.
3
Consensus is reached
An incident opens only when several probe locations agree that the target is failing. See Probe locations for where checks run.If a confirmation check succeeds, the failure is treated as transient and no incident is opened.
4
Notifications go out
Every verified notification destination assigned to the monitor that has the “Monitor goes down” event enabled is alerted.
This applies whichever locations you select on a monitor. Selected locations decide where regular checks run; confirmation checks prefer them but use the other locations when needed, so the rule holds even with a single selected location.
How an incident is resolved
Recovery mirrors detection: several locations must see the target succeed before the incident resolves. This prevents flapping, where a service bounces between failing and passing and creates alert fatigue. When the incident resolves, destinations with the “Monitor recovers” event enabled are notified. You can also resolve an incident manually from its page or with the API.Incident types
Each incident has atype that describes the cause:
Slow response
Setslow_response_threshold_ms on an HTTP or keyword monitor to be told when responses are too slow.
- A slow-response incident is separate from a downtime incident: a monitor can be up and still have an open slow-response incident.
- It resolves when response time stays below 80% of the threshold for 3 consecutive checks. With a 2,000 ms threshold, that is under 1,600 ms for three checks in a row.
SSL certificate problems
Certificate expiry warnings (30, 15, 7 and 1 days before expiry, whichever you selected) are sent to your notification destinations and do not create an incident. An expired or invalid certificate does open anssl_error incident, which resolves automatically once a valid certificate is seen. See Setting Up Alerts.
Statuses and severity
Allowed transitions:
open to acknowledged or resolved; acknowledged to resolved or closed; resolved to closed.
Severity is one of critical, major, minor or warning.
Reading an incident
Open Incidents in the sidebar and select an incident. The page shows:
The incident list can be filtered by status (
open, acknowledged, resolved, closed), severity, and search text.
Work with incidents through the API
List open incidents:Reduce false alarms
Set a realistic timeout
Set a realistic timeout
The maximum timeout is 60 seconds, and it must be shorter than the check interval. Use 10 to 15 seconds for fast APIs and up to 60 seconds for slow services. Timeouts that are too short cause failures that are not real.
Match the expected status codes
Match the expected status codes
The default accepts only
200. If your endpoint returns 201, 204 or a redirect status you rely on, add those codes.Allow UptimeIO through your firewall
Allow UptimeIO through your firewall
If a firewall or rate limiter blocks probes, allow this user agent:
Next steps
Setting Up Alerts
Choose who is notified
Reading Metrics
Uptime and response time
Creating Monitors
Monitor settings
Incidents API
Automate incident handling