Skip to content
ClelandCo

The method — AI Contingency Framework

Five phases

What happens when the model meets reality.

Every AI plan works until it meets reality.

A framework for the questions nobody asks until the answer is expensive: what happens when the model is wrong, the vendor changes terms, or an auditor wants to know how any of this is governed. Rehearsed before the system is load-bearing, not after.

Briefing

The AI market sells adoption. Readiness assessments, roadmaps, pilot programs, ROI models — an entire industry organized around getting a system into production. Almost none of it addresses the second half of the problem, which is that the system will eventually behave in a way nobody planned for, on a day nobody chose.

This is not a hypothetical risk register. It is the ordinary operating condition of a probabilistic system placed inside a deterministic business process. The model will be confidently wrong. Someone will paste customer data into a consumer chatbot. A vendor will change its pricing, its terms, or its model behind the same API name. A regulator or an acquirer will ask how the thing is governed and the answer will be assembled in a panic.

In Special Operations planning, the response to this is not optimism. It is rehearsal: you run the plan against a thinking adversary before it matters, you find where it breaks, and you build the branch you will execute when it does. The plan that survives contact is rarely the clever one. It is the one that was rehearsed.

The same discipline applies to AI, and for the same reason: the failure is almost never the technology. It is that nobody planned for the technology to fail.

Six questions

Cheap to ask now. Expensive to ask later.

01

What happens when the model gives bad advice?

Not whether it will — it will. Whether there is a defined threshold at which output is not trusted, a human in the loop at that threshold, and a record of what the system said when someone acted on it.

02

What happens when a vendor changes pricing or access?

Unit economics that work at one price per million tokens can invert at another. Model deprecation, rate-limit changes, and silent behavior changes behind a stable API name are all routine, and all of them arrive without your consent.

03

What happens when sensitive data leaks?

Through a prompt, a log, a fine-tune, a vector store, or a screenshot in a support ticket. The question is what is recoverable, what is notifiable, and who finds out first — you or the customer.

04

What happens when employees misuse it?

Shadow AI is not a policy failure, it is a demand signal. People route around tools that are slower than the alternative. The plan has to assume use you did not sanction and make the sanctioned path the faster one.

05

What happens when regulations change?

Documentation you cannot produce retroactively is the expensive kind. Provenance, evaluation results, and decision records are cheap to keep as you go and near-impossible to reconstruct a year later.

06

What happens when the model fails mid-workflow?

Degradation, not just outage. A system that returns something plausible while quietly performing worse is more dangerous than one that returns an error, because nothing alerts and everyone keeps trusting it.

The framework

Five phases. Each one produces something.

A phase that produces no artifact did not happen. Each one below ends in a document, a test result, or a decision on the record.

Phase 1

Mission analysis

Before any tooling question, the business problem in one sentence, the person it serves, and the definition of success specific enough to be falsified. Most stalled AI programs fail here and discover it eighteen months later. If success cannot be stated in a way that could turn out to be false, there is nothing to build toward and nothing to measure against.

Produces

  • The problem statement, in one sentence, agreed in writing.
  • The end user, named, and what changes for them if this works.
  • Success and failure criteria that a disinterested party could score.
  • The constraints that are real — budget, data, regulatory, political.

Phase 2

Intelligence preparation

An honest survey of the ground before committing to it: the data environment as it actually is rather than as the architecture diagram claims, the threats specific to this deployment, and the capability you genuinely have in-house. This phase produces the uncomfortable findings, which is the point of running it before the roadmap rather than during the incident.

Produces

  • Terrain — data quality and lineage, existing systems, and the organizational appetite for change.
  • Threats — prompt injection surface, data exposure paths, hallucination consequences, compliance exposure, vendor concentration.
  • Friendly forces — the engineering and data capability you have, and the executive sponsor who will still be there in six months.
  • The gap list, ranked by what would hurt most.

Phase 3

Courses of action

Multiple viable options developed in parallel and compared on the same criteria, then a recommendation with the reasoning attached. Buy is fast and cedes control. Build is controlled and expensive. Hybrid — a commercial model against a proprietary data layer — is usually right and usually harder than it looks. A single option presented as inevitable is not a recommendation, it is a preference.

Produces

  • Two or three genuinely viable courses of action, costed and scoped.
  • The same evaluation criteria applied to each, including the cost of reversing it.
  • A recommendation, with the reasoning written down so it can be argued with.
  • The decision record: what was chosen, by whom, on what date, on what evidence.

Phase 4

Rehearsal

The phase that gets skipped, and the reason this framework exists. Before the system carries real work, it is deliberately attacked, degraded, and misused under controlled conditions. Everything found here is found cheaply. Everything not found here is found by a customer.

Produces

  • A red-team report with reproductions, not a list of concerns.
  • Evaluation results against a held-out set, with the failure cases kept.
  • A tested fallback path for each dependency that can go away.
  • The runbook someone who is not you can execute at 2 a.m.

Phase 5

Sustainment

Deployment is the beginning of the measurement problem, not the end of it. Drift is gradual and unannounced; adoption decays quietly; the vendor changes the model without changing the version string. The monitoring has to be able to distinguish a real change from noise, which means it needs a baseline established before launch and error bars afterward.

Produces

  • Drift and performance monitoring against the pre-launch baseline.
  • Adoption measured by use, not by seats provisioned.
  • A standing review of security events, vendor changes, and regulatory movement.
  • An after-action review after every incident, written down and acted on.

Phase 4, in detail

The rehearsal nobody runs.

Six exercises, run against the real system under controlled conditions. Each one is a thing deliberately done to it, and a specific question it answers.

  • Exercise 01

    Prompt injection

    Hostile instructions placed where the system will read them as input — a document, a web page, a support ticket, a filename, a calendar invite.Whether the model can be talked out of its instructions, and whether anything downstream of it acts on the result without a check.
  • Exercise 02

    Data poisoning

    Corrupted or adversarial records introduced into a retrieval corpus, a fine-tuning set, or a feedback loop.Whether bad inputs are detectable before they are trusted, and how far a single poisoned record propagates before anything notices.
  • Exercise 03

    Failure injection

    The dependency is removed. The API returns errors, or latency, or a truncated response, or valid-looking nonsense.Whether the system degrades safely or fails open, and whether the fallback path has ever actually been run.
  • Exercise 04

    Bias and disparate impact

    Held-out evaluation across the populations the system will actually encounter, including the ones underrepresented in the training data.Where performance differs by group, by how much, and whether that difference has legal or ethical consequences in this specific use.
  • Exercise 05

    Misuse simulation

    The system is used the way busy people will actually use it, including the shortcuts, the copy-pasted customer data, and the uses nobody sanctioned.What the sanctioned path costs in friction, and whether the unsanctioned one is faster — because if it is, it wins.
  • Exercise 06

    Disaster recovery

    The full scenario: the model is wrong, someone acted on it, and a customer noticed before you did.Time to detect, time to contain, who is called, what is disclosed, and whether the record needed to reconstruct the decision exists.

Where this runs

A method, not a product.

The framework is not sold separately. It is the method behind the two senior engagements, and how much of it runs depends on which one you buy.

AI Strategic Advisor

Phases 1 through 3, plus the rehearsal plan. Your team runs the exercises; I design them, review the findings, and pressure-test the courses of action.

Fractional CAIO

All five phases, continuously, with me accountable for the results. Rehearsal becomes a standing gate before anything ships, and sustainment is reported monthly to the exec team or board.

What this is not.

This is not a penetration test. Network and application security are a different discipline with different certifications. The rehearsal phase covers the model, the data around it, and the workflow it sits in. Where the finding is a conventional security issue, it gets handed to people who do that for a living.

This is not a compliance certification. It produces the documentation, evaluation records, and decision trail that an audit asks for, and it will tell you where you would fail one. It is not an audit and it is not legal advice.

This is not a guarantee the system will not fail. It will. The framework exists so the failure is one you have already seen under controlled conditions, with a tested response, rather than one you are meeting for the first time in front of a customer.

Questions

Asked and answered.

Is this a separate product?
No. It is the method behind the advisory and fractional engagements, published so you can evaluate how the work is done before deciding whether to buy the person doing it. The advisory tier applies the planning phases and designs the rehearsals for your team to run; the fractional seat runs all five phases continuously and is accountable for the results.
We already have a risk register. How is this different?
A risk register lists what could go wrong. Rehearsal finds out what actually does, by doing it. The difference matters because the failures that hurt are rarely the ones anyone thought to write down — they surface when a real system is deliberately attacked, degraded, and misused under conditions you control.
Can you run this before we have anything in production?
That is the best time. Phases 1 through 3 are cheapest before a course of action is locked in, and the rehearsal phase is designed to run against a prototype. Contingency planning after launch is incident response with better paperwork.
What does the military framing actually buy me?
One specific thing: the assumption that the plan will meet an adversary and should be tested against one before it matters. That is not a metaphor borrowed for marketing — it is the planning discipline I spent eleven years using in Special Operations, where the cost of being confidently wrong was not a bad quarter. The failure modes are different. The planning problem is the same.
How long does it take?
The planning phases run in weeks and are scoped to the decision in front of you. Rehearsal is sized to the system: a single retrieval workflow is days, a multi-agent system touching customer data is longer. Sustainment is continuous by definition, which is why it belongs to a standing seat rather than a project.
Do you fix what you find, or just report it?
Both, and which one depends on the tier. The advisory engagement hands your team a report with reproductions and a remediation plan. The fractional seat fixes the pieces where a handoff would kill the work, because a finding nobody has capacity to act on is not a finding, it is a liability with a date on it.

Find it in rehearsal, not in production.

If AI is becoming load-bearing in your business and nobody has asked what happens when it fails, that is the conversation to have. It is usually a short one, and it usually changes the roadmap.