ServicesWhat changes between an AI prototype and production?

Buyer guide

What changes between an AI prototype and production?

A practical checklist for data, permissions, evaluations, failures, deployment and ownership before expanding an AI application.

On this page
  1. The inputs stop being selected by the demonstrator
  2. Permissions become part of correctness
  3. Failures need useful states
  4. Evaluation becomes repeatable work
  5. The operating cost becomes a system question
  6. Someone must own the release
  7. A release decision can be conditional

A prototype can answer whether an idea deserves further work. Production asks a different question: can the intended users depend on the defined system under the agreed conditions?

The distinction is not a judgement about how polished the interface looks. It concerns the responsibilities around the result.

The inputs stop being selected by the demonstrator

A prototype often uses a few known examples. Operational users bring incomplete requests, unfamiliar files, contradictory records and unexpected volume.

Document the supported input boundary and test important exceptions. A system can have a narrow production scope; it should not silently pretend that every possible input is supported.

Permissions become part of correctness

A correct answer shown to the wrong person is still an unacceptable result. Check retrieval, stored data, exports and tool actions against the actual user and organisation.

Do not rely on an instruction telling a model to respect permissions when the application can prevent unauthorised access before the model receives the data.

Failures need useful states

Consider a provider timeout after work has already started. Does the user know what happened? Can the task be retried without duplicating a record or sending the same action twice? Does someone receive the exception?

Define durable states for pending, running, review-required, approved, completed and failed work as appropriate. Preserve enough information to investigate without indiscriminately logging sensitive content.

Evaluation becomes repeatable work

Create a test set with agreed expectations and version it. Record the system configuration used for each evaluation. Include failures that matter to the workflow, not only questions the model usually answers well.

When the model, prompt, corpus or integration changes, run the relevant checks again. A previous release decision does not automatically approve new behaviour.

The operating cost becomes a system question

Measure the actual workflow, including retries, document processing and human review. Define usage limits and an escalation path when the system reaches them.

A production budget should say who pays variable provider charges and who investigates unusual usage.

Someone must own the release

Name the owner of deployment, incident response, access changes and future improvements. Define how the team disables or rolls back an affected capability.

Provide deployment instructions, dependencies, tests and a runbook to the receiving team. A successful handover includes their ability to use those materials, not only an email containing links.

A release decision can be conditional

The sensible outcome may be a bounded user group, a smaller document family, draft-only behaviour or mandatory review for certain actions. Those boundaries should be explicit and observable.

A production-readiness review should identify what has been tested, what remains unknown and what conditions must stay in place. It should not convert “we ran some tests” into a blanket reliability claim.

Discuss production engineering · Evaluation service

AI Services