The 404, the 500 and the Empty State: Designing for Failure
A practitioner’s guide to 404 page design, 500 pages and empty state design: real status codes, static fallbacks, caching headers and how to test failure
A client rang us about a checkout that “did nothing”. The button spun, then went back to idle, and the customer tried again. Four times. What was actually happening: the POST returned a 500 from a PHP fatal error in a payment plugin, the frontend caught the rejected promise and silently reset the button state, and nothing was logged client-side. Four duplicate charge attempts, zero error messages, and a support ticket that said “the site is broken” with no further detail.
That’s the thing about failure states. They’re the screens nobody designs, nobody QAs and everybody eventually sees. A 404, a 500 and an empty table look superficially similar (sparse page, some text, maybe an illustration) but they’re three completely different problems: a wrong address, a broken engine, and a correct page with nothing in it yet.
Most teams spend their failure budget on the 404 illustration and nothing on status codes, caching headers, logging or recovery paths. That’s backwards. The mascot is the least important part of error page ux.
- A 500 page must be pre-rendered static HTML served by the web server or CDN, not a template rendered by the application that just died. If your 500 page needs PHP, Node or the database, you don’t have a 500 page.
- Soft 404s (returning HTTP 200 with “page not found” copy) are the most common SEO-damaging mistake in SPAs and headless setups. Google decides on its own what to do with those URLs, and it usually guesses wrong.
- There are at least six distinct empty states: first run, no results after filtering, user-cleared, permission-denied, load error and offline. Shipping one generic “Nothing here” component for all six is why dashboards feel unfinished.
- Error pages need
Cache-Control: no-store. We’ve seen a CDN cache a 503 maintenance page at the edge for 10 minutes after the site came back. - Failure paths need deliberate tests: curl status audits, DevTools offline mode, request interception in Playwright, and stopping a service in staging to see what the user actually sees.
Three failures, three different jobs
Before you touch any markup, get clear on who caused the problem and who can fix it. That dictates the copy, the actions and the status code.
- 404: the user or a referring link got the address wrong, or you deleted something. The user can fix it with help. Your job is navigation.
- 500: you broke it. The user cannot fix it. Your job is honesty, a retry path, and an identifier they can quote to support.
- Empty state: nothing is broken at all. The system is working perfectly and has nothing to show. Your job is teaching the next action.
Conflating the last two is the classic bug. A table that says “No results found” when the API returned a 503 is a lie that sends the user hunting for a filter they never applied. Separate the branches at the data layer, not in the component.

404 page design: recovery, not apology
The useful question for 404 page design isn’t “how do we make this charming” but “what was this person probably looking for, and how many clicks until they have it”. On content sites and stores, three elements earn their place: a search field that actually searches, links to the five to eight highest-intent destinations, and the main navigation intact. The illustration is optional.
Two things we do on every build. First, log 404s with the referrer and the requested path, then read the log weekly for the first month after launch. On a 2,000 page migration, roughly a third of the 404 hits in week one came from three bad internal links in a footer partial. You cannot redirect what you don’t measure.
Second, if the URL looks like it had a sibling, say so. A request for /blog/how-to-fix-wordpress-cron/ that no longer exists should offer the closest live match rather than a generic sorry. Fuzzy-matching the slug against your post index is maybe 20 lines of code and converts far better than a search box nobody uses.
Keep the design inside your existing system. A 404 that looks like a different website makes people think they’ve been redirected somewhere hostile. Canvas ships 404 variants that inherit the same header, footer and type scale as the rest of the template. The page should feel like your site having a bad day, not a stranger’s site.
The 500 page has to survive the thing that broke
Here’s the test. SSH into your server, stop PHP-FPM (or kill the Node process), then load the site. If you see nginx’s default grey “502 Bad Gateway”, your branded 500 page was a template rendered by the application, and the application is the thing that’s down.
Pre-render it. Build a single self-contained HTML file with inlined CSS, no third-party fonts, no analytics, no JS, under about 15KB, and serve it from disk:
# nginx: serve a static fallback when the upstream app is unreachable
error_page 500 502 503 504 /50x.html;
location = /50x.html {
root /var/www/fallback; # plain files, never proxied to the app
internal;
# An outage page cached at the edge outlives the outage
add_header Cache-Control "no-store, must-revalidate" always;
}
Copy matters more here than anywhere. “Something went wrong” tells the user nothing. Tell them three things: it’s our fault, we know about it, and here’s what to do next. If you have a status page, link it. If the failure is in a single feature rather than the whole site, say which feature. People will otherwise assume their data is gone.
Include a request ID. A short correlation ID printed on the page and attached to the server-side log entry turns a useless ticket (“it crashed”) into a 30 second lookup. Generate it at the edge so it exists even when the app can’t produce one.
One hard rule: never put a stack trace, SQL fragment or file path on a production 500 page. In WordPress that means WPDEBUGDISPLAY false and WPDEBUGLOG true, with the log outside the web root. We still find debug.log publicly readable at /wp-content/debug.log on live sites, which is a free database schema dump for anyone curious.
Status codes, caching and the soft 404 trap
Status codes are part of the design because they change behaviour in crawlers, CDNs, service workers and monitoring tools.
- 404 for “not found, might come back”. 410 for “deliberately deleted, forget it”. Google drops 410 URLs faster in practice, which is handy after pruning thousands of expired listings.
- 503 with a
Retry-Afterheader for planned maintenance. Never 200, never 404. A 200 maintenance page can get your whole site reindexed as “we’ll be back soon”. - 401 vs 403 vs 404: use 404 for resources the user isn’t allowed to know exist. Leaking existence through 403 is a real enumeration vector on admin tooling.
Client-rendered apps are where this breaks. A React or Vue router that renders a NotFound component still returned HTTP 200 for the document, so every mistyped URL and every stale backlink becomes an indexable 200 page with thin content. If you’re on Next.js, call notFound() in the server component or route handler so the framework sets a real 404. If you’re on a static host with client routing, the host’s 404 handling (Netlify’s 404.html, Cloudflare Pages’ equivalent) needs to be wired up rather than catch-all rewritten to index.html. This is one of several reasons the one-page versus multi-page decision has consequences well past the homepage.
<?php
// Retired SKUs: 410 removes them from the index faster than a soft 404 ever will
$gone = ['/products/legacy-widget', '/products/2019-bundle'];
$path = rtrim(parseurl($SERVER['REQUESTURI'], PHPURL_PATH), '/');
if (in_array($path, $gone, true)) {
httpresponsecode(410);
header('Cache-Control: public, max-age=86400'); // 410s are safe to cache, 500s are not
include __DIR__ . '/templates/gone.php';
exit;
}
Empty state design: six states, not one
Good empty state design starts with admitting that “no data” has multiple causes, and each one needs different words and different buttons. The six we design for, every time:
- First run: the account is new. Teach. One primary action, one line explaining what will appear here, optionally a link to docs. No apology.
- No results: filters or search returned nothing. Echo the query back, show which filters are active, and give a one-click “clear filters”. This is the state teams most often get wrong by hiding the filter state.
- User cleared it: inbox zero, all tasks done. This one can be positive. It’s also the only empty state where an illustration genuinely helps.
- Permission denied: the data exists, this user can’t see it. Say who to ask, not “contact your administrator” when you know the admin’s name.
- Load error: the request failed. Show a retry button and the status, never “no items”.
- Offline: no network. Different from a server error, and worth distinguishing because the user’s action is different.
Branch at the fetch boundary so the component receives a state, not a guess:
async function loadList(url, { timeout = 8000 } = {}) {
const ac = new AbortController();
const timer = setTimeout(() => ac.abort(), timeout);
try {
const res = await fetch(url, { signal: ac.signal });
if (!res.ok) return { state: 'error', status: res.status };
const data = await res.json();
// An empty array is a legitimate answer, not a failure
return data.length ? { state: 'ready', data } : { state: 'empty' };
} catch (err) {
// AbortError means we gave up waiting; TypeError usually means no network at all
return { state: err.name === 'AbortError' ? 'timeout' : 'offline' };
} finally {
clearTimeout(timer);
}
}
Markup-wise, keep it boring and announce it to assistive tech. The role="status" matters: screen reader users otherwise get silence when a table they just filtered empties out.
<div class="empty-state text-center py-5" role="status">
<h3 class="h5 mb-2">No invoices yet</h3>
<p class="text-muted mb-4">Invoices show up here once you send one. Drafts stay private until then.</p>
<a href="/invoices/new" class="btn btn-primary">Create your first invoice</a>
<a href="/docs/invoicing" class="d-block small mt-3">How invoicing works</a>
</div>
Two rules we hold to. Don’t render fake placeholder rows as a teaching device; people try to click them. Reserve the vertical space the populated view would occupy, otherwise the layout jumps when data arrives and your CLS score suffers for a state that lasts 400ms. If you’re building list-heavy admin screens, our notes on dashboard layout patterns cover where these slot into a shell.
Forms: the failure that costs real money
Form failures are the expensive ones because the user has already invested effort. Non-negotiables:
- Never clear the fields. If the submit fails, the data stays. For anything over three fields, persist to
sessionStorageon blur so a 500 followed by a refresh doesn’t wipe 10 minutes of typing. - Disable the button only while in flight, and give it a label change (“Sending…”). Then re-enable it with an error message that says what to do, not just that something failed.
- Send an idempotency key with anything that charges money or creates a record. This is what stopped that client’s four duplicate payment attempts from becoming four charges.
- Distinguish validation (4xx) from infrastructure (5xx). Validation errors go next to the field. Infrastructure errors go at the form level with a retry.
If you’re on a static site and don’t want to own the failure surface of a mail server, an endpoint like WebForms hands back proper status codes you can branch on, so your error handling is a real condition rather than a catch that shows a generic message. Longer flows have their own abandonment mechanics, which we covered in multi-step forms.
Testing the paths nobody tests
Failure states rot because nothing in CI touches them. Make them visible:
# Soft 404 audit: anything here returning 200 is an indexing problem
for p in /nope /blog/nope /products/nope /?s=zzzz; do
curl -s -o /dev/null -w "%{httpcode} %{urleffective}\n" "https://example.com$p"
done
Beyond that: stop PHP-FPM in staging and load the homepage. Use DevTools Network to go offline, then throttle to Slow 3G and watch what renders in the first second. Intercept routes in Playwright and force a 500 and a timeout for each key screen. Add a dev-only query flag (?_force=500) that’s compiled out of production builds so designers can open each failure state without engineering help.
Also check what happens when background work fails, because that’s a failure state with no page at all. A cron job that silently stops is the same category of bug as a swallowed promise, which is exactly the trap in WordPress cron.
Frequently Asked Questions
Should a 404 page include a search box?
Yes on content-heavy sites and stores where search already works well, no on small marketing sites of 10 to 20 pages where a link list is faster. If you include one, pre-fill it with a cleaned version of the requested slug so the user doesn’t retype anything. A search box that returns nothing useful is worse than no search box.
What status code should a maintenance page return?
503 Service Unavailable, with a Retry-After header giving seconds or an HTTP date. That tells crawlers to come back rather than deindex, and tells monitoring tools this is intentional. Returning 200 for maintenance is how sites lose rankings during a two-hour migration.
Do funny 404 pages hurt conversion?
Humour is fine on a 404 because nothing is actually broken, and it reads as confidence. It’s a bad idea on a 500 or a failed payment, where the user is losing something and a joke reads as not caring. Keep the tone proportional to the cost of the failure.
Should error pages load analytics and third-party scripts?
A 404 can, since the app is healthy and you want the data. A 500 page should not: every third-party dependency is another thing that can be down at the same time, and you want that file to render from disk in one request. Capture 5xx rates server-side or at the CDN instead, where the measurement doesn’t depend on the broken thing.
How do I avoid soft 404s in a client-rendered app?
Make sure the HTTP response for an unknown route is genuinely 404 before any JavaScript runs. With Next.js use notFound() server-side; with a static host, configure its 404 document instead of rewriting every path to index.html. Then verify with curl, because the browser will happily show your NotFound component over a 200 and you’ll never notice.
Where to start
Pick one thing this week: load your site with the application process stopped. Whatever you see is your real 500 page, and for most sites it’s the web server’s default. Fix that first. It’s the failure where users have the least agency and you have the most to lose. Then separate your “empty” branch from your “error” branch in the one list view your users touch most, and audit your 404 status codes with a 10 line curl loop.
Failure states are cheap to build and expensive to skip. Design them once, put them in your component library, and every project after this one inherits them.


