What to Put Behind a /health Endpoint
A health route is the URL your uptime monitor trusts more than any other. What a shallow and a deep check should each touch, when to return 503, how to keep it fast and safe to expose, and a small example you can adapt.
In the post on false alarms I suggested pointing your monitor at a small, dedicated health route instead of the homepage. That advice is only as good as the route. Build it badly and it either reports green while the site is broken, or pages you every time a third party has a slow afternoon.
This post covers what a health endpoint should check, what it should leave alone, and how to answer so a monitor understands you.
Two different questions
A health endpoint can answer one of two questions, and mixing them up causes most of the trouble.
Is the process alive? This is a liveness check. It returns 200 as long as the application
can answer a request at all, and touches nothing else. Orchestrators like Kubernetes use it to
decide whether to restart a container. It should be almost impossible to fail, because a failure
means "kill me and start again".
Can the application do its job? This is a readiness check, sometimes called a deep check. It also verifies the things every real request depends on, such as the database and the cache. A failure means "this instance should not receive traffic right now".
Keep them separate. If a database outage makes your liveness check fail, the orchestrator restarts every instance in a loop, which fixes nothing and adds a cold-start storm to the outage.
For external uptime monitoring you usually want the second question, asked carefully. A monitor that only checks liveness will happily report green while every page returns an error because the database is gone.
What a deep check should touch
The rule I use: check the things you own, without which every request fails.
- The primary database. A trivial query such as
SELECT 1proves the connection pool works and the server answers. Do not run a real query against a real table. You want to know the database is reachable, not how fast your slowest report is today. - The cache or session store, if the application cannot serve pages without it. A
PINGor a read of a known key is enough. - Storage the application writes to on every request, if there is any. Most applications do not have any, and that is fine.
- The queue, measured the right way. A queue that accepts jobs but has no workers looks perfectly healthy to a ping. Checking the age of the oldest waiting job tells you far more, but it usually belongs in a separate, slower check, because a backed-up queue rarely means the site is down.
Each dependency gets its own short timeout, well inside the monitor's timeout, so one hung
connection produces a clean 503 instead of a request that never finishes.
What it must leave alone
Third-party APIs. It is tempting to check the payment gateway, the email provider and the search service while you are at it. Don't. Their outage becomes your pager alert about their outage, and a monitor checking every 30 seconds means you call each of them 2,880 times a day to ask how they are. Watch their status pages instead, or monitor the pages of your site that depend on them as separate monitors, so the alert names the thing that broke.
Anything slow. A health check that takes three seconds on a good day will time out on a bad one and page you about latency you already knew about.
Anything that writes. A health endpoint gets called thousands of times a day by monitors, load balancers and, eventually, scanners. It must be safe to call from anywhere, as often as anyone likes.
Answer with the status code
Most uptime monitors, Barkme included, judge a check by its HTTP status code, not by the response
body. A 200 carrying {"status": "down"} reads as perfectly healthy. So:
- Return
200when everything the check covers is fine. - Return
503 Service Unavailablewhen it is not. That is the code that means "temporarily unable to serve", and load balancers understand it too.
For Barkme, a 4xx or 5xx, a timeout, and a DNS or TLS failure all count as a failed check. So
the status code is the whole message, and the body is only for people.
Be careful with degraded states. If search is down but everything else works, the site is up. Only fail the check for what is worth waking someone for, and report the rest in the body or in your logs.
Keep it fast
The whole check should finish in well under a second. A few things help:
- Run the checks in parallel if there is more than one, so the total time is the slowest dependency, not the sum of them all.
- Cache the result for a few seconds. Several monitors and a load balancer may all be asking the same question. A five to ten second cache turns that into one real check. Keep the cache short, so a failure shows up almost at once.
- Set short connection timeouts on the connections the check uses. A database driver's default connect timeout is often thirty seconds or more, which is longer than your monitor will wait.
Safe to expose
A health endpoint has to be reachable by an outside monitor, which means it is reachable by everyone. Design for that:
- No authentication. A monitor that needs a login is a monitor that fails every time a session expires.
- Leak nothing. No version numbers, hostnames, dependency names, connection strings or stack
traces. An attacker who finds
/healthshould learn only that it answers. If you want a detailed diagnostic page, put it behind authentication or on an internal network. - No redirects. The route should answer directly, not bounce to a login page, to HTTPS or to a version with a trailing slash. Point the monitor at the exact final URL.
- No caching in front of it. Send
Cache-Control: no-store, and make sure your CDN honours it. A health route cached at the edge keeps returning the last200long after the origin behind it has died. It is one of the most convincing ways to report a dead site as healthy. - Keep it out of search results with a
noindexheader or arobots.txtentry, if it lives on a public hostname.
Do not let your own defences block it
Rate limiters, web application firewalls and tools like fail2ban are built to punish exactly what a monitor does: the same request, from the same address, forever. When one of them blocks your monitor, you get an outage alert for a site that is fine for everyone else.
Exempt the health route from rate limiting. Barkme's checks come from a fixed prober address, shown
under Settings and Team as "Monitoring Prober", and every check carries a Referer header starting
with Barkme-, so you can allowlist either one. There is a worked example in
the fail2ban post.
A small example
Here is the shape of it in Laravel. The same structure works in any framework:
use Illuminate\Support\Facades\Cache;
use Illuminate\Support\Facades\DB;
use Illuminate\Support\Facades\Log;
use Illuminate\Support\Facades\Route;
Route::get('/health', function () {
try {
DB::select('select 1');
Cache::store('redis')->get('health-probe');
$healthy = true;
} catch (Throwable $e) {
Log::warning('Health check failed', ['error' => $e->getMessage()]);
$healthy = false;
}
return response()
->json(['status' => $healthy ? 'ok' : 'fail'], $healthy ? 200 : 503)
->header('Cache-Control', 'no-store');
});
The failure detail goes to the log, where you can read it, and not into the response, where everyone else can. Register the route outside any rate-limited middleware group.
The equivalent in Express, with a hard two-second ceiling:
app.get('/health', async (req, res) => {
res.set('Cache-Control', 'no-store');
const timeout = new Promise((_, reject) =>
setTimeout(() => reject(new Error('health check timed out')), 2000),
);
try {
await Promise.race([Promise.all([db.query('SELECT 1'), redis.ping()]), timeout]);
res.status(200).json({ status: 'ok' });
} catch (err) {
console.warn('Health check failed', err);
res.status(503).json({ status: 'fail' });
}
});
A note on Laravel's built-in /up
New applications on Laravel 11 and later register a /up route in bootstrap/app.php. Out of the box it only proves
the framework boots. It fires a DiagnosingHealth event, and returns a 500 if any listener throws,
so you can add your own checks by listening for that event.
Two things to know about it. It returns an HTML page, not JSON, which does not matter to a monitor
reading the status code. And it is excluded from maintenance mode, so it keeps answering 200
while the rest of the site shows the maintenance page after php artisan down. That is handy if
planned maintenance should not page anyone, and surprising if you expected the monitor to notice.
A checklist
- One route, one purpose. Keep liveness and readiness apart.
- Check the database and cache you own. Leave third parties out.
200for healthy,503for not. Don't rely on the body.- A short timeout on every dependency, and a total under a second.
- No auth, no redirects, no caching, no internal details in the response.
- Exempt it from rate limits and allowlist your monitor.
- Monitor it, and monitor one real page as well, on a slower interval, to catch what the health route cannot see.
That last point matters. A health route tells you the application can work. It cannot tell you that the template on the checkout page is broken. Watching one real page alongside it covers the gap, and the interval post covers how often to check each one. When the health route does fail, the status code field guide covers where to look first. The how it works page shows what Barkme does with a failed check, from the first retry to the recovery notice.
Barkme watches your sites around the clock and barks the moment one goes down. Free plan, no card required.
Keep reading
More from the blog
Checkout Down Is Not Site Down
A shop can have a perfectly healthy homepage and a cart that has been failing since Friday. Which URLs to monitor on an online store, what an uptime check can and cannot prove, and the quiet failures that cost sales without taking anything down.
SPF, DKIM and DMARC, Explained
Three DNS records decide whether your mail is trusted or quietly binned. What SPF, DKIM and DMARC each check, how alignment ties them together, the mistakes that break them, and how to check your own domain.
The Alert Nobody Read
Monitoring that catches an outage and delivers the alert to an inbox nobody opens has still failed. How to route uptime alerts by who has to act, when a ticket beats a message, and how to prove a channel works before you need it.