All WordPress HTML Templates Forms & Webhooks AI & Tools
AI & Tools

From Prototype to Production: Shipping an AI Feature Users Trust

A practitioner’s guide to ship an AI feature users trust: eval harnesses, pinned model snapshots, latency budgets, shadow rollouts and honest abstention.

A real working moment illustrating the theme of an article about ship ai feature. Wide 16:9 banner, one strong focal point, magazine editorial quality, authentic and unstaged.

A client demo last spring: a support-ticket summariser that worked beautifully on the twelve tickets we’d pasted into it. Shipped behind a flag to 5% of their inbox two weeks later. Within a day it had confidently summarised a legal threat as a billing question and told an agent the customer “seemed happy with the resolution.”

Nothing was broken. The API returned 200. The model did what models do when the input looks nothing like your twelve examples. That gap, between a prototype that demos well and a feature people will actually rely on, is where most ai product development stalls. It is almost never a modelling problem.

If you want to ship an AI feature users trust, treat the model as the least reliable component in your stack and design everything around that assumption. Here’s what that looks like in practice.

Key Takeaways

  • A prototype proves the model can do the task. An eval set of 80 to 150 real, labelled examples proves it does the task often enough to ship, and tells you when a prompt edit breaks something you fixed last month.
  • Pin model snapshots (for example gpt-4o-2024-08-06, not gpt-4o) and version your prompts in git. An upstream model update is a silent deploy you didn’t review.
  • Never call a model inline in a request that renders a page. Budget 4 seconds, stream or queue everything longer, and define the degraded path before the happy path.
  • Trust comes from the interface, not the accuracy number: show sources, make edits cheap, label confidence honestly, and give users a one-click way to say “that’s wrong.”
  • Roll out in shadow mode first, compare against human output offline, then go 1%, 10%, 50% with a kill switch that a non-engineer can flip.

The demo-to-production gap is an evaluation gap

Prototypes are built on inputs you chose. Production runs on inputs your users chose, which include empty strings, 40,000-word email threads, three languages in one message, pasted HTML, and someone’s entire CV in the “describe your issue” box.

The honest framing: your prototype measured nothing. It produced a vibe. Vibes are a legitimate signal that a feature is worth building, and a terrible basis for deciding whether it’s ready. The transition from ai mvp to shipped feature is mostly the work of replacing that vibe with a number you can regress against.

So before you write another prompt, collect real inputs. Pull 100 rows out of the production table the feature will touch, sample deliberately for the ugly ones, and label the correct output by hand. It takes an afternoon. On a ticket-classification project we found 14% of real tickets had no single correct category, which changed the product requirement: the feature needed an “unsure, route to human” state that nobody had specced.

Build the eval harness before you tune the prompt

An eval harness is a script, a JSON file of cases, and a pass threshold. It does not need a platform. We’ve shipped plenty of features where the whole thing was 60 lines of Node that ran in CI.

// evals/run.js  (node --experimental-vm-modules not required; plain ESM)
import { readFileSync } from 'node:fs';
import { classifyTicket } from '../src/ai/classify.js';

const cases = JSON.parse(readFileSync('./evals/tickets.json', 'utf8'));
const failures = [];
let pass = 0;

for (const c of cases) {
  // temperature 0 so a regression means the prompt changed, not the dice
  const out = await classifyTicket(c.input, { temperature: 0 });
  const ok = out.category === c.expected.category
    && (out.confidence >= 0.6) === c.expected.confident;
  ok ? pass++ : failures.push({ id: c.id, got: out.category, want: c.expected.category });
}

const rate = pass / cases.length;
console.log(${pass}/${cases.length} (${(rate * 100).toFixed(1)}%));

if (failures.length) console.table(failures);
// hard gate: below this, the build does not merge
if (rate < 0.92) process.exit(1);

Two rules that earn their keep. First, every production bug becomes a new eval case before it gets fixed. That is just regression testing with the labels written by a human instead of an assertion. Second, grade the thing you actually care about. If the output is free text, don’t score it on exact match; score it on the three facts that must be present, with a cheap string check or a second model call constrained to output {"hasrefundamount": true}. Model-as-judge works fine for fuzzy criteria as long as you spot-check the judge on 20 cases and keep its prompt in the same repo.

Where this advice doesn’t apply: genuinely open-ended creative output, where there’s no correct answer. For those, measure rejection rate in the UI instead and accept that your offline eval is weak.

What it takes to ship an AI feature users trust

Trust is not accuracy. We’ve shipped a feature at roughly 88% agreement with human labels that users loved, and a different one at 95% that they switched off. The difference was recoverability.

At 88%, every output arrived as an editable draft with the source passage highlighted, so a wrong answer cost two seconds to fix. At 95%, the output was auto-applied to a customer record, so the 5% cost a support ticket and a trust hit that never came back. Users do the maths on failure cost, not on your benchmark.

Concretely, the things that move perceived trustworthiness:

  • Attribution. Quote the source and link to it. If you can’t cite it, you probably shouldn’t be asserting it.
  • Editability. Output lands in a textarea, not in the database. The accept action is explicit and reversible for at least one session.
  • Honest abstention. “I couldn’t find this in your documents” beats a plausible paragraph. Make abstention a first-class output in your schema and an eval criterion, or the model will never choose it.
  • Visible scope. Tell users what the feature reads. “Based on your last 20 tickets” prevents the assumption that it saw everything.

Everything in our guardrails piece on hallucination, tone and escalation applies double here, because a feature that escalates well fails gracefully, and graceful failure is what users remember.

The boring architecture decisions that decide reliability

Model calls are slow, occasionally wrong, and sometimes just gone. Design for all three.

Budget the latency, then degrade

Pick a number. We use 4 seconds for anything a user is waiting on and 25 seconds for anything queued. Past the budget, abort and return the non-AI path: the keyword search, the template, the blank field. The page must render without the model.

export async function withBudget(fn, { ms = 4000, fallback }) {
  const ctrl = new AbortController();
  const timer = setTimeout(() => ctrl.abort(), ms);
  try {
    return await fn(ctrl.signal);
  } catch (err) {
    // timeout, 429 and provider 5xx all get the same answer: degrade, log, move on
    return fallback(err);
  } finally {
    clearTimeout(timer);
  }
}

Queue anything that isn’t on screen

Bulk summarisation, embedding backfills, enrichment triggered by a webhook: none of that belongs in a request handler. Provider rate limits mean retries, retries mean duplicate work, and duplicate work on a paid API is a line item. Use a real queue with idempotency keys and exponential backoff. The reasoning is the same as for webhook processing you should not do inline, with the extra wrinkle that each retry costs money.

Cache the deterministic parts

Hash the normalised input plus the prompt version plus the model snapshot, and cache on that key. On a documentation assistant with repetitive questions this cut provider spend by a bit over a third, and the cache key means a prompt change invalidates cleanly instead of serving stale answers. If you’re doing retrieval, read this on whether you actually need a vector database first, because for corpora under a few thousand chunks Postgres with pgvector is usually enough and one less thing to operate.

Model churn is a deploy you didn’t review

Pin snapshots. Always. Using a floating alias means your provider ships changes to your production behaviour on their schedule, and you find out from a user. Pinned versions get deprecated too, so put the deprecation date in your calendar and treat migration as a scheduled task with an eval run attached.

Prompts belong in version control as files, not in a database row someone edits in an admin panel. Log the prompt version with every call, alongside the model snapshot, token counts, latency and a hash of the input. When someone asks why output quality dropped on the 14th, you want to answer it with a query:

{
  "trace_id": "7f3c...",
  "feature": "ticket_summary",
  "prompt_version": "summary.v7",
  "model": "gpt-4o-2024-08-06",
  "latency_ms": 2140,
  "tokens": { "in": 1823, "out": 212 },
  "abstained": false,
  "user_action": "edited",
  "edit_distance": 94
}

user_action is the field that matters. Accepted, edited, discarded, reported: that’s your real quality metric, and it’s free. Track edit distance on accepted-after-edit and you’ll see prompt regressions days before anyone complains.

Rollout: shadow, slice, kill switch

Four stages, in order, no skipping.

  1. Shadow. Run the feature on live traffic, write the output to a log, show nobody. Compare 200 outputs against what humans actually did. This catches the input distribution problems your eval set missed.
  2. Internal. Your team and the client’s team, with a visible “report this output” control. Expect the first week to produce more product requirements than bugs.
  3. Percentage. 1%, then 10%, then 50%, with at least 48 hours and a look at user_action rates between steps. If discard rate climbs above your pre-agreed threshold, you stop.
  4. Default on with an off switch in user settings that persists.

The kill switch must be a config flag flippable without a deploy, by someone who isn’t an engineer. We’ve had to use ours twice in two years: once for a provider outage, once because a prompt change doubled abstention rate at 2am on a Saturday. Both times the feature degraded to the old non-AI path and nobody outside the team noticed.

Design the degraded state properly rather than letting it be an exception trace. The same thinking applies as in designing for 404s, 500s and empty states: the fallback is a real screen users will see, so give it copy and a next action.

One practical note for prototypes that live on a static page before the backend exists. You still want structured feedback from reviewers, and wiring a form endpoint on day one is wasted effort. A free endpoint like WebForms pushes the thumbs-down and the pasted output straight into Slack, which is enough to run stage two of this list before you’ve written a single database migration.

Frequently Asked Questions

How many eval cases do I actually need before shipping?

80 to 150 hand-labelled real inputs is the range where we start trusting the number, and it’s enough to detect a meaningful regression. Below about 50, a single flaky case swings your pass rate by two points and you’ll chase noise. Weight the set towards edge cases rather than mirroring production distribution, because the easy inputs are not where you break.

Should the AI feature be free or part of a paid tier?

Meter it from day one even if you don’t charge, because an unmetered model call is an unbounded bill. Track cost per user per month in your logs from the first week of the beta. If the median active user costs more than about 10% of their subscription price, either the feature needs a usage cap or the pricing needs revisiting before launch, not after.

Can we skip the eval harness if a human reviews every output?

Partly, yes. Human-in-the-loop lowers the accuracy bar considerably, which is exactly why draft-and-edit interfaces are the safest first shipped feature. You still need evals before any prompt or model change, otherwise you’re asking reviewers to notice a 6% quality drop by feel, and they won’t.

Which model should we start with in 2026?

Start with the strongest available model to establish whether the task is solvable at all, then try to move down to a cheaper or smaller one with your eval set as the gate. Doing it the other way round means you can’t tell whether a failure is the task, the prompt or the model. Keep the provider call behind a thin internal interface so swapping is a one-file change.

How do we handle user data and prompts sent to a third-party provider?

Get the provider’s data retention and training terms in writing, turn off training on your data where that’s an option, and document which fields leave your infrastructure. Strip or tokenise identifiers you don’t need in the prompt, since most tasks work fine on redacted text. For regulated clients we’ve run smaller open-weight models on dedicated infrastructure specifically to keep the data boundary clean, and accepted the quality hit.

Where to start tomorrow

If your feature is sitting at the prototype stage right now, don’t touch the prompt. Spend the next afternoon pulling 100 real inputs out of production and labelling them by hand. You’ll learn more about whether this feature should ship than another week of tuning will tell you, and you’ll leave with the one artefact that makes every later decision measurable.

Then pick your failure mode deliberately: draft-and-edit if you can tolerate 10% wrong, abstain-and-escalate if you can’t. Everything else, the queue, the pinned snapshot, the kill switch, follows from that choice.

ai eval harness in CI ai feature rollout shadow mode ai mvp to production checklist human in the loop ai interface design llm latency budget and fallback pinning model snapshot versions ship ai feature ship ai feature to production