Skip to content
ClelandCo

The method — AI Contingency Framework

Five phases

What happens when the model meets reality.

Every AI plan works until it meets reality.

A framework for questions that are easy to defer and expensive to meet unprepared: what happens when the model is wrong, a vendor changes terms, or an auditor asks for the governance record. Rehearsed before the system is load-bearing, not after.

Briefing

AI adoption work often concentrates on readiness, roadmaps, pilots, and the path to launch. The second half of the operating problem is what happens after that plan meets changing inputs, users, dependencies, and model behavior.

A probabilistic component inside a business process can return plausible but incorrect output, receive data outside its intended boundary, or change when a vendor updates the model or terms. Customers, auditors, and acquirers may also ask for the decision and evaluation record. The framework turns those possibilities into named exercises and artifacts before the response is urgent.

Jeremy’s Special Operations planning background informs the emphasis on rehearsal and fallback branches. This framework adapts that planning habit to AI systems; the analogy does not substitute for system-specific technical, legal, security, or operational evidence.

The same planning discipline is useful for AI: technical failure is only one part of the incident. Detection, authority, fallback, communication, and recovery also have to be designed and rehearsed.

Six questions

Cheap to ask now. Expensive to ask later.

01

What happens when the model gives bad advice?

Define the threshold at which output is not trusted, the human review or fallback at that threshold, and the record needed to reconstruct what the system said before someone acted on it.

02

What happens when a vendor changes pricing or access?

Unit economics that work at one price per million tokens can invert at another. Model deprecation, rate-limit changes, and behavior changes behind a stable API name are provider risks worth modeling because they may arrive outside the client’s control.

03

What happens when sensitive data leaks?

Through a prompt, a log, a fine-tune, a vector store, or a screenshot in a support ticket. The question is what is recoverable, what is notifiable, and who finds out first — you or the customer.

04

What happens when employees misuse it?

Unapproved AI use can indicate unmet task demand as well as a policy/control gap. The plan should inventory the task and data, provide a usable sanctioned path where appropriate, and retain privacy, security, employment, and records review.

05

What happens when regulations change?

Documentation you cannot produce retroactively is the expensive kind. Provenance, evaluation results, and decision records are generally easier to preserve contemporaneously than to reconstruct later, but the required record depends on the system and applicable obligations.

06

What happens when the model fails mid-workflow?

Degradation, not just outage. Plausible degradation can be harder to detect than an explicit error when no threshold, alert, or review path is in place.

The framework

Five phases. Each one produces something.

Each phase ends in a written artifact: a document, a test result, or a decision on the record. A phase that produces none of those did not happen.

Phase 1

Mission analysis

Before any tooling question, state the business problem in one sentence, name the person it serves, and define success specifically enough to be falsified. If the result cannot turn out to be false, there is no agreed target to build toward or measure against.

Produces

  • The problem statement, in one sentence, agreed in writing.
  • The end user, named, and what changes for them if this works.
  • Success and failure criteria that a disinterested party could score.
  • The constraints that are real — budget, data, regulatory, political.

Phase 2

Intelligence preparation

An evidence-based survey of the ground before committing to it: the data environment as it actually is rather than as the architecture diagram claims, the threats specific to this deployment, and the capability you genuinely have in-house. This phase produces the uncomfortable findings, which is the point of running it before the roadmap rather than during the incident.

Produces

  • Environment — data quality and lineage, existing systems, and the organizational appetite for change.
  • Threats — prompt injection surface, data exposure paths, hallucination consequences, compliance exposure, vendor concentration.
  • Capacity — the engineering and data capability you have, and the executive sponsor who will own the work.
  • The gap list, ranked by what would hurt most.

Phase 3

Courses of action

Develop multiple viable options and compare them on the same criteria, then attach the reasoning to the recommendation. Buying, building, and combining commercial services with owned data each trade speed, control, cost, reversibility, and operating burden differently. A single option presented as inevitable is not a comparison.

Produces

  • Two or three genuinely viable courses of action, costed and scoped.
  • The same evaluation criteria applied to each, including the cost of reversing it.
  • A recommendation, with the reasoning written down so it can be argued with.
  • The decision record: what was chosen, by whom, on what date, on what evidence.

Phase 4

Rehearsal

Before the system carries material work, it is deliberately attacked, degraded, and misused under controlled conditions. Findings made here can be reproduced and assigned before an uncontrolled incident; rehearsal cannot prove that every failure mode has been found.

Produces

  • A red-team report with reproductions, not a list of concerns.
  • Evaluation results against a held-out set, with the failure cases kept.
  • A tested fallback path for each in-scope dependency named in the rehearsal plan.
  • The runbook someone who is not you can execute at 2 a.m.

Phase 5

Sustainment

Deployment is the beginning of the measurement problem, not the end of it. Drift is gradual and unannounced; adoption decays quietly; the vendor changes the model without changing the version string. The monitoring has to distinguish a real change from noise, which means a baseline before launch and a rule for what counts as a real change afterward.

Produces

  • Drift and performance monitoring against the pre-launch baseline.
  • Adoption measured by use, not by seats provisioned.
  • A standing review of security events, vendor changes, and regulatory movement.
  • An after-action review after every incident, written down and acted on.

Phase 4, in detail

Rehearse failure before release.

Six exercise families, scoped to the threat model and run against the system or a representative prototype. They are not exhaustive. Each one is a thing deliberately done to it, and a specific question it answers.

  • Exercise 01

    Prompt injection

    Hostile instructions placed where the system will read them as input — a document, a web page, a support ticket, a filename, a calendar invite.Whether anything downstream of the model acts on manipulated output without a check.
  • Exercise 02

    Data poisoning

    Corrupted or adversarial records introduced into a retrieval corpus, a fine-tuning set, or a feedback loop.Whether bad inputs are detectable before they are trusted, and how far a single poisoned record propagates before anything notices.
  • Exercise 03

    Failure injection

    The dependency is removed. The API returns errors, or latency, or a truncated response, or valid-looking nonsense.Whether the system degrades safely or fails open, and whether the fallback path has ever actually been run.
  • Exercise 04

    Bias and disparate impact

    Held-out evaluation across the populations the system will actually encounter, including the ones underrepresented in the training data.Where performance differs by group, by how much, and whether that difference has legal or ethical consequences in this specific use.
  • Exercise 05

    Misuse simulation

    The system is used the way busy people will actually use it, including the shortcuts, the copy-pasted customer data, and uses outside the approved path.What the sanctioned path costs in friction, and whether the unsanctioned one is faster — because if it is, it wins.
  • Exercise 06

    Disaster-recovery walkthrough

    A tabletop of the full scenario: the model is wrong, someone acted on it, and a customer noticed before you did.Time to detect, time to contain, who is called, what is disclosed, and whether the record needed to reconstruct the decision exists.

Where this runs

A method, not a product.

The framework is not sold separately. It is the method behind the two senior engagements, and how much of it runs depends on which one you buy.

AI Strategic Advisor

Phases 1 through 3, plus the rehearsal plan. Your team runs the exercises; I design them, review the findings, and pressure-test the courses of action.

Fractional CAIO

All five phases within the agreed mandate, with me accountable for designing, reviewing, and reporting the evidence. Reserved legal, financial, and headcount matters stay with authorized officers. Rehearsal becomes a standing gate before in-scope systems ship, and sustainment is reported on the agreed leadership cadence.

What this is not.

This is a published method, not a client result. Publishing the framework does not establish that it was independently validated or that any client engagement produced a particular outcome.

This is not a penetration test. Network and application security are a different discipline with different certifications. The rehearsal phase covers the model, the data around it, and the workflow it sits in. Where the finding is a conventional security issue, it gets handed to people who do that for a living.

This is not a compliance certification. It can organize documentation, evaluation records, and a decision trail against stated requirements and identify evidence gaps for qualified reviewers. It is not an audit, readiness opinion, certification, or legal advice.

This is not a guarantee that the system will not fail. Failures remain possible. The framework increases the set of conditions the team has already exercised and gives each tested condition an owner, evidence record, and response path.

Questions

Asked and answered.

Is this a separate product?
No. It is the method behind the advisory and fractional engagements, published so you can evaluate how the work is done before deciding. The advisory engagement applies the planning phases and designs rehearsals for your team to run; the fractional mandate designs, reviews, and reports the agreed phases. Reserved legal, financial, and headcount matters stay with authorized officers.
We already have a risk register. How is this different?
A risk register lists what could go wrong. Rehearsal tests selected risks by deliberately attacking, degrading, and misusing the system under controlled conditions. It produces reproductions and response evidence; it does not establish that untested risks are absent.
Can you run this before we have anything in production?
Yes. Phases 1 through 3 can run before a course of action is locked in, and rehearsal can run against a representative prototype. After launch, the same method supports sustainment and incident preparation, but changes may carry more dependencies and migration work.
What does the military framing actually buy me?
It explains the framework’s emphasis on rehearsal, branches, and explicit authority under uncertainty. That background is provenance for the method, not proof that a military plan and an AI system share the same risks or that the framework is independently validated.
How long does it take?
Timing depends on the decision, system, integrations, data sensitivity, exercise set, participants, evidence requirements, and retest work. A proposal should name those assumptions rather than promise a generic number of days or weeks. Sustainment can be assigned inside a standing mandate or to an internal role.
Do you fix what you find, or just report it?
Both; the mandate determines who executes. In advisory scope, the client team runs the exercises and records reproductions; the advisor designs the plan, reviews supplied evidence, and drafts recommendations. A fractional mandate may implement only named pieces assigned in scope. Findings without assigned capacity remain open risks.

Name the failure conditions before they are load-bearing.

If AI is becoming load-bearing in your business, name the failure conditions, evidence, owner, and fallback before the system carries material work. That is the question this framework is designed to structure.