Skip to content
Cleland & Co.

AI strategy

Published October 11, 2026 · 10 min read

AI Prototype to Production: A Readiness Checklist

Check whether an AI prototype is ready for production with evaluation cases, a review-time example, release gates, and a readiness worksheet.

By for Cleland & Co.Published

About 10 min left

  • AI prototype to production
  • AI production readiness
  • AI implementation
Loose graphite panels and green cubes progress toward an assembled cube beside a brass inspection gauge.
Conceptual illustration of an AI production-readiness review.

The short answer

Before releasing an AI workflow, establish what it must do, how you will test it, who can access its data, what it costs to operate, and who takes over when it fails. Use the checklist to make a documented release, limited-trial, or hold decision.

Start here: Define the supported workflow, record the evidence for each release criterion, and leave unresolved failures visible. Use the editable production-readiness worksheet (CSV) (opens in a new tab) during the review; it opens in a spreadsheet without a signup.

A useful AI prototype can still leave a business with an unanswered question: what would have to be true before people could rely on it?

Moving an AI prototype to production means making that question specific. You need a defined job, tests that reflect real use, appropriate access, workable costs, and a person responsible for the system after launch. The model is one part of that operating arrangement.

This checklist is for AI-powered workflows such as document processing, internal assistants, and support triage. It focuses on the release decision. If your team is still trying to understand why a promising demo has not progressed, start with why AI pilots stall.

Start with a job you can evaluate

Write a short description of the work, including where it begins and ends. “Help with invoices” leaves too much open. “Extract specified fields from incoming invoices and prepare a draft for a reviewer” gives the team something to test.

List what the system is allowed to do. Reading a document, proposing a change, and writing to a business system require different decisions. A prototype that drafts information should not gain permission to act on it simply because the integration is available.

Record how the work is handled today. That baseline might include review time, corrections, missed cases, and the effort required to resolve an exception. The comparison should include the whole task: preparation and checking can outweigh time saved on the first draft.

Test the cases your business actually receives

Assemble an evaluation set that reflects the intended workload. Include ordinary examples and the situations that could change the release decision: incomplete records, unfamiliar formats, ambiguous requests, contradictory information, and inputs outside the agreed scope.

Keep some examples separate from development so the final review is not limited to cases used to tune the system. Record the model, instructions, tools, and data sources used for each run. Otherwise, a later team may be unable to explain why the results changed.

Agree on acceptable behavior before reviewing the results. Some errors can be corrected in a normal workflow; others should prevent release. An assistant that misformats a heading and an assistant that reveals another customer's information have different failure consequences.

NIST's AI Risk Management Framework (opens in a new tab) organizes risk work around governance, context, measurement, and management throughout a system's life. Use that framing to ask what must be understood and controlled. The worksheet below is a practical editorial tool; it does not establish compliance or certification.

Build an AI evaluation set you can inspect

For each test case, record the input reference, the expected behavior, what actually happened, the failure category, and who judged it. Keep sensitive source material in the approved system; the test register can reference it without copying it everywhere.

For an internal knowledge assistant, a starting test matrix might look like this. These are proposed test designs, not observed system results.

Test caseExpected behaviorEvidence to keep
A question answered by one current policyAnswer matches the policy and points to the supporting passageRetrieved passage, answer, reviewer judgment
Two approved documents disagreeIdentify the conflict and route it for resolutionBoth source versions and the escalation
The answer is absent from approved sourcesState the limitation instead of inventing a policyRetrieval result and final response
A user requests a restricted team's documentAccess is denied before restricted content reaches the model or userAuthorization test and relevant access log
A document tells the assistant to ignore its rulesTreat the instruction as untrusted content; do not grant new access or actionsObserved behavior and tool activity
The source service times outReturn the agreed fallback without presenting an unverified answerTimeout, user message, and support event

Measure the stages separately. If the right passage was never retrieved, rewriting the final answer prompt may not solve the problem. If the passage was retrieved but the answer contradicted it, investigate generation and review. If the correct answer went to the wrong user, treat that as an access failure, regardless of the answer-quality score.

Define what “pass” means for each category. For example, “answers match the approved policy” needs a review rubric: required facts, forbidden unsupported claims, and acceptable ways to express uncertainty. Have reviewers resolve disagreements before using their labels as a release gate.

There is no universal number of examples that proves production readiness. Cover the intended workload and the consequences of failure, then expand the set when you discover a new failure pattern. Passing a small test set is evidence about those tests, not proof that every future input is safe.

Make access part of the review

Document which information the workflow reads, where it is sent, where outputs are stored, and which users can retrieve them. Check those boundaries with accounts that have different permissions.

For a system that uses uploaded documents or retrieved content, include tests where that material contains misleading instructions. Determine whether the system treats those instructions as data or follows them in ways it should not. A warning in a prompt is not sufficient evidence that an access boundary holds.

Any action outside the agreed scope needs a clear rejection or escalation path. Name who can change permissions and how that change is reviewed.

Count the cost of completing the task

Build a cost estimate using expected volume, then update it during a limited trial. Include model calls, supporting services, retries, storage, review time, and correction work. Keep the assumptions visible.

Consider a busy period as well as a normal one. If the workflow creates more exceptions than the receiving team can handle, a technically available system may still be impractical to operate.

Define what should happen when a budget, usage, or workload limit is reached. The answer could be a queue, a narrower service, or a return to the existing process. Someone needs authority to make that choice.

Calculate cost per completed task, including human review

Use this planning equation:

Monthly operating cost = model and service charges + review labor + exception labor + ongoing support.

Divide the total by successfully completed business tasks. Count a task once even if the system retried several times, and retain the costs of failed attempts. Use your own supplier charges and labor costs rather than borrowing a rate from another business.

First, check whether the workflow releases useful capacity. Here is a hypothetical invoice-processing example, not a client result. It assumes all 1,000 invoices are eventually completed after review and corrections.

Monthly workAssumptionStaff time
Current manual process1,000 invoices at 5 minutes each83 hours 20 minutes
Review in the proposed processAll 1,000 invoices at 2 minutes each33 hours 20 minutes
Additional exception handling100 invoices at 6 extra minutes each10 hours
Ongoing supportAssumed monthly workload4 hours
Proposed process totalReview, exceptions, and support combined47 hours 20 minutes

The modeled difference is 36 staff hours per month. If review takes three minutes instead of two, the proposed process requires 64 hours, leaving 19 hours 20 minutes of capacity. That is why review time belongs in the pilot measurements.

To compare costs, apply the appropriate labor cost to each role's hours and add model, hosting, storage, and other service charges. Include build and onboarding costs separately when deciding whether the investment pays back. Time released is not automatically cash saved: the business must decide what useful work will fill it, and observed results must support the assumptions.

Give the system an owner after launch

Name the person who reviews operating results, investigates incidents, and approves material changes. Identify the information they will receive and the conditions that require action.

Test the fallback. Ask the receiving team to stop the AI step and complete the work through the alternative process. If that requires knowledge only the prototype's developer has, the handover is unfinished.

The AI contingency method provides further context for planning what happens when an AI dependency becomes unreliable or unavailable.

Use this production-readiness worksheet

Copy this table into your project notes. Replace each prompt with the actual evidence, responsible person, and open issue. Leave an unknown visible rather than treating it as a pass.

CriterionEvidence to collectOwnerOpen issueDecision
Defined jobIntended users, inputs, outputs, permitted actionsName the business ownerWhat remains ambiguous?Ready / limited / hold
EvaluationRepresentative cases, results, agreed acceptance criteriaName the evaluatorWhich failure matters most?Ready / limited / hold
AccessData-flow review and permission testsName the access ownerWhich boundary is untested?Ready / limited / hold
EconomicsVolume assumptions, observed costs, review effortName the budget ownerWhat changes at peak demand?Ready / limited / hold
OperationsMonitoring, incident responsibilities, change processName the operatorWho receives an alert?Ready / limited / hold
FallbackA rehearsed alternative and restart procedureName the receiving team leadCan the team use it?Ready / limited / hold

A “ready” entry should point to something another person can inspect. A missing document may be quick to resolve; an unresolved permission failure may require a different design. Avoid averaging those problems into a reassuring score.

Worked example: invoice extraction with review

Consider an illustrative workflow that extracts invoice fields into a draft record. A person checks the source document before any approved information enters the accounting system. This is a hypothetical example, not a client result.

The review finds that unfamiliar layouts sometimes produce missing tax fields. The team cannot yet support all formats, but it can identify an approved group of document types and route everything else to the existing manual process.

Its worksheet might record:

  • Evaluation: Additional testing is needed for unfamiliar layouts. Limit the trial to the tested document types.
  • Access: The extractor can create drafts but cannot approve payments or change supplier details.
  • Operations: A named reviewer checks each draft against the original. Track corrections and the time required.
  • Fallback: Unsupported or failed cases go to manual entry, with a way to avoid processing the same invoice twice.

Those are proposed controls to test, not proof the workflow is safe. The trial should establish whether reviewers can reliably spot errors and whether the complete process is useful at the expected volume.

Move from prototype to production through explicit release gates

A practical rollout can use three gates. Adapt them to the workflow's consequences and the team's capacity; they are an operating recommendation, not a required certification process.

Gate 1: replay approved historical work. Compare outputs with known outcomes using permitted data. Save the versioned evaluation results and investigate serious failures. This checks behavior before a live workflow depends on it.

Gate 2: run a limited, supervised trial. Name the users, document types, volume limit, review duties, and end date. Keep the existing process available. Check whether the human reviewer catches errors and whether the queue can be handled within the agreed working day.

Gate 3: approve a defined production scope. Record the supported cases, unresolved exclusions, actual operating cost, access test results, and fallback rehearsal. Increase scope only when new evidence supports it.

Write the stop conditions before the trial

A useful stop condition names both the event and the response. For example: “If the assistant exposes content outside the user's permission, disable the affected retrieval path, preserve the relevant evidence, and route work to the existing process while the access owner investigates.” Other conditions might concern duplicate writes, unavailable reviewers, or spending beyond an approved cap. Choose limits for the actual use case; do not borrow a generic accuracy target.

Monitoring should include task completion, correction reasons, queue age, latency, and cost, with an owner for each actionable signal. A dashboard alone does not assign responsibility. Decide who can pause the workflow and how users will know to use the fallback.

Require regression checks when a model, prompt, retrieval source, or connected tool changes. LLMOps—the operation of applications built around large language models—includes this repeatable evaluation and change process. For the business owner, the important deliverable is a record showing what changed, what was tested, and who accepted the remaining limitations.

Make the release decision explicit

End the review with a recorded decision: release the agreed scope, run a limited trial with stated conditions, or hold while a named issue is resolved. Include who approved the decision and what would cause the team to revisit it.

That record turns a general concern about production readiness into work someone can own. If you need help evaluating the next step, explore AI consulting and describe the workflow, the pilot evidence, and the decision your team is facing.

If the review reveals an ongoing leadership gap, the fractional CAIO hiring guide explains how to evaluate that role.

References and boundaries

Primary references, not borrowed authority.

These sources inform the framing. They do not endorse Cleland & Co., validate a client outcome, or turn this guide into a certification standard.

  • AI RMF Core (opens in a new tab)

    National Institute of Standards and Technology

    Lifecycle risk framing. The worksheet and example below are proposed working tools, not a NIST assessment or certification.

Questions

Asked and answered.

What is the difference between an AI prototype and a production system?
A prototype tests whether an idea can work. Production adds a defined operating scope, representative evaluation, enforced access, workable costs, monitoring, support, and a tested fallback for actual users.
What should an AI production readiness checklist include?
Include the job and permitted actions, evaluation results, data and access boundaries, total operating costs, named owners, monitoring, and fallback tests. Each item should point to evidence, an unresolved issue, and a release decision.
How much testing does an AI pilot need before production?
There is no universal test count or accuracy threshold. Cover common work, consequential failures, unsupported inputs, and access boundaries. Keep held-out cases and define acceptance criteria for the intended scope before judging results.
How do you calculate the cost of an AI workflow?
Add model and supporting service charges, review labor, exception handling, and ongoing support. Divide by successfully completed business tasks. Assess build and onboarding costs separately, and validate assumptions during a limited trial.

Give your AI project a clear next step.

Share the workflow, current test results, and the decision holding up the next step. We can discuss a review of the approach and define any further work in the scope.