From AI prototype to production: a checklist
Evaluations, guardrails, prompt injection, cost limits, observability and fallbacks — what an LLM feature needs before real users rely on it.

Getting a language model to do something impressive in a demo is easy. Keeping it useful, safe and affordable once real users arrive is the hard part. This is the checklist we work through before an AI feature goes live.
Most of it isn't about the model at all. It's about the stages a request passes through on its way in and out:
1. Evaluations
Build a test set of real inputs with expected outcomes — including awkward edge cases. Score what you can automatically (correct format, correct classification, cites a source) and review a sample by hand.
Run it on every change to prompts, models or retrieval. Without evaluations, you can't tell whether a change made things better or quietly made them worse.
2. Guardrails on inputs and outputs
- Validate structured output against a schema, and retry or fail cleanly when it doesn't match.
- Check outputs for things that should never appear: personal data, other customers' information, off-topic or harmful content.
- Decide how the feature fails. A clear "I can't help with that" beats a confident wrong answer.
3. Prompt injection
Treat any text the model reads — user messages, retrieved documents, web pages, emails — as untrusted. It can contain instructions that try to override yours.
If the model can take actions through tools:
- Give each tool the smallest permissions that work.
- Require confirmation for anything destructive or expensive.
- Never rely on the system prompt alone to keep data or actions safe.
4. Cost controls
LLM costs scale with usage, and a single runaway loop can be expensive.
- Set limits on output tokens, context size and requests per user.
- Use prompt caching where your provider supports it.
- Route simple requests to smaller, cheaper models and save the largest models for hard ones.
- Track cost per request and per customer from day one, and set budget alerts.
5. Observability
Log prompts, responses, latency, token counts and the exact model version — with sensible handling of personal data. For multi-step flows such as RAG or agents, trace each step, so when an answer is wrong you can see whether retrieval, the prompt or the model was at fault.
6. Latency and reliability
- Stream responses so users see progress immediately.
- Set timeouts, and retry transient errors with backoff.
- Plan a fallback — another model or region, or a clear error message — for when a provider is slow or down.
7. Version prompts and models like code
Keep prompts in version control, and pin model versions rather than always using "latest". A model upgrade can change behaviour in subtle ways, so treat it like any other release: run the evaluations first.
8. Data and privacy
Know where user data goes, how long it's kept, and which region processes it. Running models through Amazon Bedrock keeps requests inside AWS, and model providers don't see your prompts, which can simplify security reviews. If you use cross-Region inference, check which Regions your inference profile can route to, and check each model's data-retention and abuse-detection terms.
Update (October 2026): Bedrock now has per-Region data-retention settings, and some newer models keep traffic for up to 30 days for abuse detection (within AWS, not shared with the model provider).
None of this is glamorous, but it's the difference between an AI feature users try once and one they come to rely on.