Practice

The Alert Nobody Read

Monitoring that catches an outage and delivers the alert to an inbox nobody opens has still failed. How to route uptime alerts by who has to act, when a ticket beats a message, and how to prove a channel works before you need it.

The API went down at two in the afternoon on a Tuesday. The monitoring caught it within a minute and sent an alert, exactly as configured, to ops@. ops@ was a shared inbox with four thousand unread messages, and the alert sorted itself neatly between a newsletter and a GitHub digest.

The team found out the way teams usually do: a customer posted about it. The awkward line in the postmortem was not "monitoring failed". It was "monitoring worked, and nobody saw it".

Detection is not delivery

An alert has two halves. The first is noticing that something is wrong. The second is getting that fact in front of a person who can do something about it, fast enough to matter.

Teams spend nearly all their effort on the first half. They compare check intervals, argue about timeouts and tune confirmation windows. Then they send the result to wherever was easiest to set up on the first day, which is almost always email.

An alert that arrives where nobody is looking is not an alert. It is a record of what you would have known.

Why the shared inbox fails

Email is a perfectly good channel for some things, and a shared email address is a poor place for anything urgent:

  • Nobody owns it. Everyone assumes someone else is reading it, and on a busy day everyone is right.
  • Filters swallow it. Somebody sets up a rule to tidy the inbox, and the alerts go with the rest of the automated mail.
  • It looks like everything else. An outage alert and a weekly report arrive with the same weight, in the same list, with the same notification. Often the notification is switched off, because work email on a phone is a lot of noise.
  • Recovery is invisible. If the recovery notice lands in the same pile, nobody knows whether the incident they half-noticed is still going.

Route by who has to act

The useful question is not "which channels do we have?" It is "who has to do something when this fires, and where are they right now?" Each type of channel answers that differently.

Personal channels reach a person. Email and Telegram in Barkme belong to one user. Nobody else on the team sees them, not even the account owner, and each teammate sets up their own. They suit the person who has to act. A Telegram message on a phone is much harder to miss than an email, which is why solo founders tend to live there.

A shared room reaches the team. A Slack channel means everyone sees the same alert at the same moment, and the conversation about it happens in one thread. Someone saying "I'm on it" is useful in its own right, because it stops three people investigating the same thing.

A ticket is a record. Linear and Trello turn an incident into a piece of work with an owner and a history. That matters when the cause needs looking into, not just the symptom.

A webhook is for your own tooling. A paging system, an internal dashboard, a chat bot, a spreadsheet. Anything you run that accepts JSON.

Most teams need more than one of these. The mistake is sending everything everywhere, which is how you get alert fatigue. Pick the channel by who has to act, not by what is available.

A ticket that stays open on purpose

When an incident opens, Barkme's Linear channel creates an issue in the team you choose, and the Trello channel creates a card in the list you choose. The monitor, the URL, the time it was detected, the error and a probable cause are all in it, so nobody has to copy anything across from a chat message.

When the site recovers, Barkme does not close the ticket. It adds a comment with the recovery time and the total downtime, and leaves the issue open.

That is deliberate. "The site is responding again" and "we understand what happened" are different states. An issue that closes itself when the symptom goes away files the incident as handled before anyone has looked at why the database ran out of connections. Leaving it open means someone closes it on purpose, after the cause has been investigated. If you want tickets only as a mirror of uptime, closing them yourself costs one click. Reopening an investigation nobody started costs a lot more.

Only incidents create tickets. SSL and domain expiry warnings, and blacklist alerts, go to email, Telegram, Slack and webhooks, because a reminder that a certificate expires in thirty days is not a bug.

Webhooks for everything else

If your tooling is not on the list, the custom webhook sends every incident, expiry warning and blacklist alert to your own endpoint as JSON. Each request carries two headers:

  • X-Barkme-Timestamp, the Unix time the request was sent.
  • X-Barkme-Signature, an HMAC-SHA256 of the timestamp, a dot, and the raw request body, keyed with the channel's signing secret and written as hex.

Barkme generates the signing secret when you create the channel. Verify every request before you trust it, and reject old timestamps so a captured request cannot be replayed:

$timestamp = (string) $request->header('X-Barkme-Timestamp');
$signature = (string) $request->header('X-Barkme-Signature');

$expected = hash_hmac('sha256', $timestamp.'.'.$request->getContent(), $signingSecret);

if (! hash_equals($expected, $signature) || abs(time() - (int) $timestamp) > 300) {
    abort(401);
}

The same check in Node:

const crypto = require('node:crypto');

function isValidBarkmeRequest(rawBody, timestamp, signature, signingSecret) {
  const expected = crypto
    .createHmac('sha256', signingSecret)
    .update(`${timestamp}.${rawBody}`)
    .digest('hex');

  const fresh = Math.abs(Date.now() / 1000 - Number(timestamp)) <= 300;

  return (
    fresh &&
    typeof signature === 'string' &&
    signature.length === expected.length &&
    crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(signature))
  );
}

Two details catch almost everyone. First, compute the signature over the raw body, exactly as it arrived. If your framework parses the JSON and you re-encode it, the bytes change and the signature never matches. Second, compare with a constant-time function like hash_equals or timingSafeEqual, not ==, so the comparison does not leak how many characters matched.

Prove it before you need it

The worst time to find out a channel does not work is during an outage.

Every Barkme channel has to pass a real test message before it receives alerts. The test goes through the actual pipe to the actual destination, so "configured" and "working" mean the same thing. Send another test whenever something changes: a Slack channel renamed, a Telegram chat recreated, a webhook endpoint redeployed.

Channels also break on their own. If Slack revokes a webhook, Barkme marks the channel as disconnected on the dashboard. You see a problem that needs fixing, not a channel that quietly stopped delivering.

And when a channel is too loud for a while, during a migration or a noisy staging week, mute it rather than delete it. Muting keeps the configuration and pauses delivery. A deleted channel has to be recreated by someone who remembers it existed, usually the day after the outage it would have caught.

A routing table for a five-person team

Here is a setup I would start from for a small product team. Adjust the names, keep the shape:

Event Where it goes Why
Production incident Slack #incidents and a Linear issue Everyone sees it, and the investigation has a home
Production incident, out of hours Telegram for whoever is responsible that week One person gets a phone notification
Staging incident A separate, muted-by-default Slack channel Visible when you look, never wakes anyone
SSL or domain expiry warning Email to whoever owns DNS and billing, plus Slack Needs a renewal, not an emergency
Blacklist listing Slack and the email of whoever runs mail Needs an investigation and a delisting request
Recovery The same place as the alert Whoever saw the problem sees it end

On Barkme's plans, Free sends email only, Founder adds Slack and Telegram, and Pro adds Linear, Trello and the custom webhook. The pricing section has the details.

A checklist

  1. Write down who has to act for each kind of alert, then choose the channel.
  2. Never rely on a shared inbox for anything urgent.
  3. Put incidents where the team can see them together, and give investigations a ticket.
  4. Send expiry warnings to whoever can actually renew the thing.
  5. Verify webhook signatures over the raw body, with a timestamp check.
  6. Test every channel after every change.
  7. Mute noisy channels. Don't delete them.

None of this makes your site more reliable. It makes the monitoring you already have worth something, which is the difference between knowing first and learning from a customer.

For keeping alerts rare enough that people still read them, see why your uptime monitor cries wolf. For what the alerts are telling you once they arrive, see the status code field guide. How it works follows an incident from the first failed check to the recovery notice, and the F.A.Q covers channel setup in detail.

Put this into practice

Barkme watches your sites around the clock and barks the moment one goes down. Free plan, no card required.

Try Barkme free