Capstone: Write a Decision Memo

LESSON

Decision Making, Uncertainty, and Judgment

008 40 min intermediate CAPSTONE

Capstone: Write a Decision Memo

By the end of this lesson, you will be able to...

  • frame a consequential choice as a decision another person can inspect;

  • compare options with evidence, probabilities, values, reversibility, and explicit trade-offs;

  • specify failure paths, owners, triggers, and a review plan before the outcome is known;

  • evaluate the memo with a rubric that separates decision quality from outcome luck.

Idea in one sentence: A good decision memo makes the choice, uncertainty, values, controls, and next review visible before action, so disagreement can improve the plan instead of becoming retrospective blame.

Core Insight

Consider the checkout team before its next high-traffic campaign. The team has a real choice: expand a new caching path to more customers, hold the current canary, or delay the campaign while it collects evidence. The decision has no option that is free. Waiting may lose revenue and learning time. Expanding may expose customers to payment or availability failures. A confident sentence such as “the new path is safe” hides the actual work.

The capstone is to write a memo that another person can challenge and execute. It is not a prediction contest and it is not a record written after the result. It is a compact model of the decision before the result is known. The memo should make four things inspectable:

  1. what is being chosen and what remains uncertain;
  2. why the options differ in consequences and values;
  3. how the team will detect a wrong turn and change course;
  4. when the decision will be reviewed and what would update it.

The previous lessons supplied the parts. Lesson 001 separated decision quality from outcome quality. Lesson 002 generated options and a baseline. Lesson 003 defined events and calibrated probabilities. Lesson 004 made values, constraints, and stakeholder costs explicit. Lesson 005 matched commitment to reversibility, information gain, and cost of delay. Lesson 006 turned imagined failures into indicators, defenses, owners, and stop conditions. Lesson 007 protected the before-state so later feedback would not rewrite history.

The memo is the track's synthesis artifact. It should be short enough to use under pressure, but detailed enough that a reviewer can find the assumptions and disagree with a particular step.

Start With the Decision, Not the Solution

Open with a decision question that has an actor, an action, a time horizon, and a review point.

Weak:

“Should we improve checkout?”

Stronger:

“Should the checkout team expand the cache-backed payment path from 10% to 25% of campaign traffic this Friday, with a review after the first hour?”

The stronger question narrows the object. It does not claim that 25% is correct. It tells the reader what choice the memo must help them make. If the decision is not time-sensitive, state what makes delay valuable or costly. If the decision is reversible, say how; if it is not, identify the commitment that cannot be recovered.

Then name the owner and the decision authority. A memo without an owner is an essay. A team may provide evidence, but somebody must be able to choose, pause, or roll back.

Use a Decision-Memo Structure

Use the following structure. Each section should answer a different question; do not repeat the same confidence statement under several headings.

1. Decision and success condition

State the question, owner, deadline, and review horizon. Define success and unacceptable failure in observable terms. “Reliable” is not a success condition until it becomes a metric, range, or user-visible boundary.

For the checkout decision:

The thresholds are not universal truths. Their value is that a different reviewer can question them before the action.

2. Context, evidence, and unknowns

Separate what is known from what is inferred. Name the reference period, the relevant comparison, and the missing evidence that could change the choice.

For example, the team knows that the canary handled normal traffic for two hours, that payment idempotency checks passed in replay, and that warm-cache recovery was rehearsed once. It does not know how the cache behaves when two zones lose capacity during peak traffic. That unknown should not disappear inside the word “tested.”

Avoid turning the context into a data dump. Include evidence only when it changes an option, probability, value, or control. Put additional material in an appendix if the decision can be inspected without it.

3. Options, baseline, and learning path

List at least three meaningful alternatives, including the current baseline. A useful set often contains:

Option What it does now What it teaches How it can be undone
Hold at 10% protects current exposure adds normal-traffic evidence easy; no new commitment
Expand to 25% with rollback buys campaign capacity tests a wider traffic range rollback to 10% if triggers fire
Delay and run a stress rehearsal spends time before exposure tests the cold-cache unknown preserves the choice, costs campaign time

Do not label options “safe,” “risky,” or “obvious” before showing their mechanisms. Include the option of doing nothing when it is genuinely available. If one option is only a variant of another, explain what information or reversibility makes it distinct.

4. Beliefs, probabilities, and update triggers

Define an event before assigning a probability. Give a range when precision is not justified, and state what would move the estimate.

The team might write:

“Given the canary, replay, and the missing two-zone stress test, we estimate a 20–30% chance of a customer-visible checkout failure during the first hour at 25% traffic. A successful zone-loss rehearsal would lower the range; a warm-cache miss or payment queue growth would raise it.”

This statement does three jobs. It records the belief before the outcome, makes the uncertainty visible, and gives the review a future update rule. Do not use a probability to disguise a value choice. A 5% chance of duplicate payment may be unacceptable even if it is less likely than a 30% chance of temporary feature loss.

If probabilities are too weak to support a numerical range, use an ordinal comparison and explain the evidence gap. “The zone-loss failure is less likely than a degraded-mode activation, but its impact is much higher” is more useful than false precision.

5. Values, constraints, and trade-offs

State what the decision protects and what it is willing to spend. Separate hard constraints from preferences. This is where a memo shows why two reasonable people can choose differently without either having made a calculation error.

For the checkout team, payment integrity is a veto constraint. Availability and campaign revenue matter, but they cannot justify duplicate charges. Customer trust and operator fatigue are also costs, even if they are harder to count. The memo should say which cost is accepted, by whom, and for how long.

Write one explicit trade-off sentence:

“We accept lower campaign capacity and a temporary degraded experience to keep payment integrity and rollback authority intact.”

If a stakeholder would reject that sentence, the disagreement is useful. Change the values or constraint before pretending the options are comparable.

6. Reversibility, optionality, and commitment

Describe the undo path, its cost, and the information gained by each step. A reversible action can be a rational way to learn, but “we can roll back” is incomplete if rollback takes an hour or leaves inconsistent state.

For the 25% expansion, the memo names a rollback command, an owner, a maximum recovery time, and the data that must remain valid after rollback. It also records the cost of waiting: the campaign may lose an hour of capacity and the team may miss a useful learning window.

Use commitment levels deliberately:

The label is less important than the reasoning. Explain why this commitment fits the uncertainty and the cost of delay.

7. Pre-mortem, signals, defenses, and owners

Imagine that the decision failed, then trace a plausible path rather than listing dramatic disasters. For each important path, record the hidden assumption, leading indicator, defense, owner, and stop or restore condition.

Failure path Assumption under pressure Leading signal Defense and authority
Warm cache is unavailable after a zone loss reserved capacity is ready when needed fill rate misses target; queue depth rises rehearse zone loss; platform owner blocks expansion
Retry creates duplicate payment all paths preserve idempotency duplicate-key mismatch replay payment requests; payments owner vetoes
Alert arrives but nobody acts an owner will interpret the signal acknowledgement exceeds 2 minutes incident commander stops expansion automatically

Each row must be able to change the plan. “Monitor performance” is not a defense until the memo says which metric, threshold, authority, and action it implies.

8. Review plan and feedback loop

End before the outcome, not with a prediction of victory. State when the memo will be reviewed, what evidence will be compared with the forecast, and which updates are possible.

The review may revise:

Write the review rule now. For example: “After the first hour, compare the recorded range with error rate, queue depth, cache-fill behavior, and payment-integrity checks. Classify any gap as luck, model, execution, or framing before changing the next plan.” This keeps the memo connected to the decision journal from lesson 007.

A Complete Worked Memo

The following condensed memo shows how the pieces fit without claiming certainty.

Decision: Should we expand the cache-backed checkout path from 10% to 25% during Friday's campaign?

Owner and horizon: The campaign incident commander decides by 10:00 and reviews after 60 minutes at 25%. The payments lead can veto any step that threatens payment integrity.

Evidence and unknowns: The path passed two hours of normal canary traffic, payment replay, and one warm-cache rehearsal. The two-zone peak-load case remains untested. Current checkout errors are 0.2%; the reference period is the last three comparable campaigns.

Options: Hold at 10% and collect more data; expand to 25% with an automatic rollback; delay for a two-zone stress rehearsal. The baseline protects exposure but loses capacity and learning time.

Belief: We estimate a 20–30% chance of customer-visible checkout failure in the first hour at 25%. The range decreases after a successful stress rehearsal and increases if cache-fill rate or payment queue depth deteriorates.

Values and constraints: Payment integrity is a veto. We accept reduced campaign capacity and temporary feature loss before accepting duplicate payment or an unowned incident.

Commitment: Choose the reversible 25% pilot. Roll back if errors exceed 1% for five minutes, queue depth grows for three intervals, or any duplicate-key check fails. The incident commander owns the action; rollback must complete within 60 seconds.

Pre-mortem controls: Rehearse zone loss before expansion where possible. The platform owner watches warm-cache fill; the payments lead watches idempotency; the incident commander owns alert acknowledgement and can stop expansion.

Review: After 60 minutes, compare the observations with the probability range and record whether any gap came from luck, model, execution, or framing. Update one estimate, test, trigger, owner, or constraint before the next campaign.

Notice what the memo does not say. It does not promise that the pilot will succeed. It does not hide the missing stress test. It does not make “customer trust” a decorative value while allowing duplicate charges. It gives a reviewer several precise places to disagree.

Review the Memo Before You Approve It

Use this rubric. Score each dimension as clear, partial, or missing, and write one sentence of evidence.

  1. Decision clarity: Is there one question, owner, deadline, and review horizon?
  2. Option quality: Are the baseline, a reversible step, and a delay or learning path meaningfully different?
  3. Evidence discipline: Are known facts, unknowns, and reference periods separated?
  4. Belief calibration: Is the event defined, the range defensible, and the update trigger visible?
  5. Value transparency: Are constraints, stakeholder costs, and one explicit trade-off stated?
  6. Commitment fit: Does the action match reversibility, information gain, and cost of delay?
  7. Failure control: Does each important path have a signal, defense, owner, and trigger?
  8. Feedback design: Can a later reviewer distinguish outcome luck from model, execution, or framing gaps?

A memo with one missing dimension is not automatically useless. The point of the rubric is to expose the missing part before commitment. If the missing part is a veto constraint or a rollback owner, pause the decision. If it is a low-impact evidence detail, record the limitation and decide whether the cost of delay justifies filling it.

Common Failure Modes

The memo is a disguised recommendation

The author chooses a favorite option first and fills the sections with supporting evidence. Force the baseline and a real alternative into the table, then state what evidence would make the recommendation change.

The memo confuses probability with permission

A low probability is not permission to violate a hard constraint. Put values and vetoes beside probabilities, not after them.

The memo lists risks without controls

“There may be latency” does not tell anyone what to watch or who can act. Trace one mechanism from assumption to signal to defense.

The memo hides ownership in “the team”

If no named person can stop, roll back, or escalate, the trigger is only decoration. Assign authority before action.

The memo becomes too long to use

Move supporting evidence to an appendix. Keep the decision question, options, beliefs, values, commitment, controls, and review rule on the main page. The trade-off is between completeness and attention; a memo earns length only when each sentence helps a choice or a review.

The memo is judged by the outcome

After the decision, use the two-pass review from lesson 007. Inspect the before-state first. A good result can follow weak process; a bad result can follow a defensible choice and bad luck.

Practice: Write Your Own Memo

Choose a real or realistic choice with moderate stakes: a migration step, a study plan, a product experiment, a household purchase, or an organizational change. Do not choose a situation where writing the exercise would expose confidential information.

Draft one page using the eight sections. Then perform three passes:

  1. Framing pass: underline the question, owner, horizon, baseline, and hard constraint. If one is missing, repair it.
  2. Uncertainty pass: circle every probability, range, unknown, and update trigger. Remove numbers that have no event definition or evidence.
  3. Control pass: trace the most consequential failure path. Add its signal, defense, owner, and stop condition. Add a review date before you finish.

Ask a reviewer to challenge one option, one value, and one assumption. Do not defend the memo immediately. Record the challenge as evidence that may improve the decision. If the reviewer cannot tell what would change your mind, the memo is not finished.

What This Track Leaves You With

The track began with a common confusion: people often judge a choice by the outcome it happened to produce. It ends with a reusable artifact that keeps the distinction visible. A defensible decision is not guaranteed to win. It is a choice whose question, alternatives, beliefs, values, commitment, controls, and review path can be inspected before the world supplies its noisy answer.

The memo is also a handoff point. A later specialization in causal decision-making can ask which interventions change the outcome. A track on sequential decision-making can formalize repeated updates and policies. Those tracks build on this foundation rather than replacing it: the decision still needs a frame, an owner, a value boundary, and a feedback loop.

Resources

Key Takeaways

PREVIOUS Feedback, Regret, and Decision Journals