The Alert Nobody Read
Monitoring that catches an outage and delivers the alert to an inbox nobody opens has still failed. How to route uptime alerts by who has to act, when a ticket beats a message, and how to prove a channel works before you need it.
The API went down at two in the afternoon on a Tuesday. The monitoring caught it within a minute
and sent an alert, exactly as configured, to ops@. ops@ was a shared inbox with four thousand
unread messages, and the alert sorted itself neatly between a newsletter and a GitHub digest.
The team found out the way teams usually do: a customer posted about it. The awkward line in the postmortem was not "monitoring failed". It was "monitoring worked, and nobody saw it".
Detection is not delivery
An alert has two halves. The first is noticing that something is wrong. The second is getting that fact in front of a person who can do something about it, fast enough to matter.
Teams spend nearly all their effort on the first half. They compare check intervals, argue about timeouts and tune confirmation windows. Then they send the result to wherever was easiest to set up on the first day, which is almost always email.
An alert that arrives where nobody is looking is not an alert. It is a record of what you would have known.
Why the shared inbox fails
Email is a perfectly good channel for some things, and a shared email address is a poor place for anything urgent:
- Nobody owns it. Everyone assumes someone else is reading it, and on a busy day everyone is right.
- Filters swallow it. Somebody sets up a rule to tidy the inbox, and the alerts go with the rest of the automated mail.
- It looks like everything else. An outage alert and a weekly report arrive with the same weight, in the same list, with the same notification. Often the notification is switched off, because work email on a phone is a lot of noise.
- Recovery is invisible. If the recovery notice lands in the same pile, nobody knows whether the incident they half-noticed is still going.
Route by who has to act
The useful question is not "which channels do we have?" It is "who has to do something when this fires, and where are they right now?" Each type of channel answers that differently.
Personal channels reach a person. Email and Telegram in Barkme belong to one user. Nobody else on the team sees them, not even the account owner, and each teammate sets up their own. They suit the person who has to act. A Telegram message on a phone is much harder to miss than an email, which is why solo founders tend to live there.
A shared room reaches the team. A Slack channel means everyone sees the same alert at the same moment, and the conversation about it happens in one thread. Someone saying "I'm on it" is useful in its own right, because it stops three people investigating the same thing.
A ticket is a record. Linear and Trello turn an incident into a piece of work with an owner and a history. That matters when the cause needs looking into, not just the symptom.
A webhook is for your own tooling. A paging system, an internal dashboard, a chat bot, a spreadsheet. Anything you run that accepts JSON.
Most teams need more than one of these. The mistake is sending everything everywhere, which is how you get alert fatigue. Pick the channel by who has to act, not by what is available.
A ticket that stays open on purpose
When an incident opens, Barkme's Linear channel creates an issue in the team you choose, and the Trello channel creates a card in the list you choose. The monitor, the URL, the time it was detected, the error and a probable cause are all in it, so nobody has to copy anything across from a chat message.
When the site recovers, Barkme does not close the ticket. It adds a comment with the recovery time and the total downtime, and leaves the issue open.
That is deliberate. "The site is responding again" and "we understand what happened" are different states. An issue that closes itself when the symptom goes away files the incident as handled before anyone has looked at why the database ran out of connections. Leaving it open means someone closes it on purpose, after the cause has been investigated. If you want tickets only as a mirror of uptime, closing them yourself costs one click. Reopening an investigation nobody started costs a lot more.
Only incidents create tickets. SSL and domain expiry warnings, and blacklist alerts, go to email, Telegram, Slack and webhooks, because a reminder that a certificate expires in thirty days is not a bug.
Webhooks for everything else
If your tooling is not on the list, the custom webhook sends every incident, expiry warning and blacklist alert to your own endpoint as JSON. Each request carries two headers:
X-Barkme-Timestamp, the Unix time the request was sent.X-Barkme-Signature, an HMAC-SHA256 of the timestamp, a dot, and the raw request body, keyed with the channel's signing secret and written as hex.
Barkme generates the signing secret when you create the channel. Verify every request before you trust it, and reject old timestamps so a captured request cannot be replayed:
$timestamp = (string) $request->header('X-Barkme-Timestamp');
$signature = (string) $request->header('X-Barkme-Signature');
$expected = hash_hmac('sha256', $timestamp.'.'.$request->getContent(), $signingSecret);
if (! hash_equals($expected, $signature) || abs(time() - (int) $timestamp) > 300) {
abort(401);
}
The same check in Node:
const crypto = require('node:crypto');
function isValidBarkmeRequest(rawBody, timestamp, signature, signingSecret) {
const expected = crypto
.createHmac('sha256', signingSecret)
.update(`${timestamp}.${rawBody}`)
.digest('hex');
const fresh = Math.abs(Date.now() / 1000 - Number(timestamp)) <= 300;
return (
fresh &&
typeof signature === 'string' &&
signature.length === expected.length &&
crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(signature))
);
}
Two details catch almost everyone. First, compute the signature over the raw body, exactly as
it arrived. If your framework parses the JSON and you re-encode it, the bytes change and the
signature never matches. Second, compare with a constant-time function like hash_equals or
timingSafeEqual, not ==, so the comparison does not leak how many characters matched.
Prove it before you need it
The worst time to find out a channel does not work is during an outage.
Every Barkme channel has to pass a real test message before it receives alerts. The test goes through the actual pipe to the actual destination, so "configured" and "working" mean the same thing. Send another test whenever something changes: a Slack channel renamed, a Telegram chat recreated, a webhook endpoint redeployed.
Channels also break on their own. If Slack revokes a webhook, Barkme marks the channel as disconnected on the dashboard. You see a problem that needs fixing, not a channel that quietly stopped delivering.
And when a channel is too loud for a while, during a migration or a noisy staging week, mute it rather than delete it. Muting keeps the configuration and pauses delivery. A deleted channel has to be recreated by someone who remembers it existed, usually the day after the outage it would have caught.
A routing table for a five-person team
Here is a setup I would start from for a small product team. Adjust the names, keep the shape:
| Event | Where it goes | Why |
|---|---|---|
| Production incident | Slack #incidents and a Linear issue |
Everyone sees it, and the investigation has a home |
| Production incident, out of hours | Telegram for whoever is responsible that week | One person gets a phone notification |
| Staging incident | A separate, muted-by-default Slack channel | Visible when you look, never wakes anyone |
| SSL or domain expiry warning | Email to whoever owns DNS and billing, plus Slack | Needs a renewal, not an emergency |
| Blacklist listing | Slack and the email of whoever runs mail | Needs an investigation and a delisting request |
| Recovery | The same place as the alert | Whoever saw the problem sees it end |
On Barkme's plans, Free sends email only, Founder adds Slack and Telegram, and Pro adds Linear, Trello and the custom webhook. The pricing section has the details.
A checklist
- Write down who has to act for each kind of alert, then choose the channel.
- Never rely on a shared inbox for anything urgent.
- Put incidents where the team can see them together, and give investigations a ticket.
- Send expiry warnings to whoever can actually renew the thing.
- Verify webhook signatures over the raw body, with a timestamp check.
- Test every channel after every change.
- Mute noisy channels. Don't delete them.
None of this makes your site more reliable. It makes the monitoring you already have worth something, which is the difference between knowing first and learning from a customer.
For keeping alerts rare enough that people still read them, see why your uptime monitor cries wolf. For what the alerts are telling you once they arrive, see the status code field guide. How it works follows an incident from the first failed check to the recovery notice, and the F.A.Q covers channel setup in detail.
Barkme watches your sites around the clock and barks the moment one goes down. Free plan, no card required.
Keep reading
More from the blog
Checkout Down Is Not Site Down
A shop can have a perfectly healthy homepage and a cart that has been failing since Friday. Which URLs to monitor on an online store, what an uptime check can and cannot prove, and the quiet failures that cost sales without taking anything down.
What to Put Behind a /health Endpoint
A health route is the URL your uptime monitor trusts more than any other. What a shallow and a deep check should each touch, when to return 503, how to keep it fast and safe to expose, and a small example you can adapt.
SPF, DKIM and DMARC, Explained
Three DNS records decide whether your mail is trusted or quietly binned. What SPF, DKIM and DMARC each check, how alignment ties them together, the mistakes that break them, and how to check your own domain.