Skip to main content

Chapter 12: Narratives and Outcomes

The project that became a success

Northbridge Services launches a new scheduling system for its field teams. The investment case promises three things: shorter customer waiting times, less administrative work, and fewer missed visits. Six months later, missed visits have fallen, but average waiting time has risen and dispatchers spend longer correcting automated schedules.

At the benefits review, the program director presents the launch as a success. The main slide shows the drop in missed visits. Waiting time appears in an appendix, divided into categories that make comparison with the original baseline difficult. Administrative effort is described as “temporary adoption activity.” The director reminds the committee that the board’s strategic goal was digital modernization.

None of these statements is necessarily false. Missed visits matter. Early implementation often requires extra work. Modernization was a strategic goal. Yet the account has changed. A project approved on three promised benefits is now judged mainly on one achieved benefit and the fact that new technology was deployed.

The story may also move the other way. A technically successful project can be called a failure because it did not solve a problem it was never funded to solve. A pilot can be criticized for lacking national scale. A team can meet an agreed target only to learn that leaders now consider a different measure decisive.

Organizations need narratives. Data do not arrange themselves into meaning. Leaders must explain what happened, why it matters, and what should happen next. The risk begins when the story controls which evidence counts, disguises a changed standard, or makes a contested choice appear inevitable.

This chapter examines five patterns: selecting measures that favor one interpretation, changing the success story after results arrive, describing a decision as more agreed than it was, reclassifying an outcome after the fact, and presenting too few alternatives.

A result has more than one story

Consider a project delivered two months late but under budget. It may be:

  • a scheduling failure;
  • a financial success;
  • a wise response to new safety evidence;
  • an example of excessive original optimism;
  • a deliberate trade of time for scope;
  • a warning that the budget excluded hidden internal labor.

Several descriptions can be true at once. The task is not to discover the single neutral story. It is to connect claims to agreed purposes, complete enough evidence, and transparent judgment.

A disciplined review separates:

  1. Promise: what result, constraint, and measure were agreed?
  2. Observation: what happened, including uncertainty and missing data?
  3. Attribution: what factors probably contributed?
  4. Evaluation: good or bad against which standard and for whom?
  5. Action: what should be continued, corrected, stopped, or tested?

When these steps blur, evaluation can masquerade as observation. “Adoption was strong” may mean 80 percent of accounts were activated, 30 percent of employees used the core function weekly, or leadership believes resistance has ended. Each is a different claim.

1. Choosing measures that favor one interpretation

Measurement requires selection. A dashboard cannot show every consequence. The choice of numerator, denominator, period, comparison group, threshold, and missing-data rule will emphasize some realities and mute others.

Suppose Northbridge reports missed visits as a percentage of all scheduled visits. The rate falls from 7 percent to 4 percent. If the new system also schedules fewer difficult rural visits, the improvement may partly reflect a changed case mix. This does not make the measure fraudulent. It means the interpretation needs volume and case context.

Ask seven questions about a consequential measure:

  1. What decision or purpose is it meant to inform?
  2. What exactly is counted, and what is excluded?
  3. What comparison or baseline is used?
  4. Did the population, process, or definition change?
  5. What important trade-off does the measure not show?
  6. Can the people judged by it inspect and correct the underlying data?
  7. What behavior could improve the number without improving the real outcome?

The last question matters because measures become targets. A service center can reduce average waiting time by ending complex calls early. A hospital can improve a recorded timeliness indicator by changing when the clock begins. A sales team can raise conversion by excluding difficult prospects. Most people respond rationally to what an accountability system rewards.

Use a small family of measures when a single indicator creates an obvious blind spot: outcome, quality, volume or access, time, cost, and unintended consequence. Do not add all six automatically. Too many indicators make selective storytelling easier because presenters can choose whichever moved favorably.

Predefine the main measures and acceptable changes before the result where possible. Keep the original definition alongside a revised one during transition. If a better measure becomes available, adopt it; do not preserve a weak metric merely for consistency. Explain the change, restate prior periods if valid, and show whether the conclusion depends on the new definition.

Qualitative evidence matters too. Customer accounts, staff observations, incident narratives, and expert judgment can reveal effects numbers miss. State how cases were selected and avoid using a vivid anecdote as proof of prevalence. A story can identify a mechanism without estimating how often it occurs.

2. Changing the success story after the result

Goals evolve. During a project, regulation may change, a competitor may enter, a safety issue may emerge, or leaders may deliberately trade one benefit for another. Adaptation is not dishonesty. The problem is changing the basis of evaluation without recording when, why, and by whom it changed.

At approval, Northbridge emphasizes customer waiting time, administrative effort, and missed visits. After launch, the program is defended as a modernization initiative. If modernization genuinely became the primary objective, the governance record should show that decision and the accepted consequences. Otherwise, the new story protects the project from the promises that secured investment.

Motivated reasoning helps explain why this can happen without a conscious plan to deceive. People are often more likely to search for and accept reasons supporting a preferred conclusion, while still experiencing themselves as reasonable.1 Commitment can intensify after resources, reputation, and identity are invested. Research on escalation of commitment examines why decision-makers may continue a course despite adverse information.2

Use a benefits history rather than a final benefits list. Record:

  • the original purpose and material assumptions;
  • baseline and target definitions;
  • approved changes to scope, benefits, timing, or risk;
  • the decision-maker and date;
  • what was knowingly traded away;
  • current results and uncertainty;
  • which claims are forecasts rather than realized benefits.

Separate delivery from benefit. Installing the system may be a delivery achievement. Reducing waiting time is an outcome. Creating the capacity to improve waiting time is an intermediate result. All can be worth reporting, but one does not prove the next.

Also resist outcome bias: judging the quality of a decision solely through the result. Experiments have shown that people can evaluate the same decision more favorably when they are told it produced a good outcome.3 A sound decision under uncertainty can end badly; a reckless decision can get lucky. Review the information, alternatives, process, and risk treatment available at the time, then learn from the outcome as new evidence.

3. Describing a decision as more agreed than it was

After a meeting, language tends to smooth conflict. “The group agreed.” “Stakeholders endorsed the direction.” “There was broad alignment.” Sometimes this is accurate shorthand. Sometimes participants agreed only to continue, preferred different options, raised conditions, or were informed after the choice.

Inflated agreement has practical effects. It discourages later questions—“we already agreed”—and transfers ownership of a decision to people who did not make it. It can also deprive the accountable leader of proper credit and responsibility.

Use a decision record that distinguishes:

  • decision: what was chosen;
  • authority: who chose under what mandate;
  • participation: who attended or submitted input;
  • support: who explicitly supported the choice, if relevant;
  • conditions: what qualifications or dependencies were attached;
  • dissent: material alternatives or objections that remain relevant;
  • implementation obligation: what participants must do despite disagreement;
  • reopening trigger: what new evidence or event returns the issue to review.

Not every meeting needs a roll-call vote. Consensus can be sensed through discussion, and recording every passing doubt would make minutes unreadable. Ask for explicit confirmation when the consequence is high or the group’s endorsement is part of the authority. “I am hearing support for Option B subject to legal confirmation by Friday. Does anyone believe that misstates their position?”

Silence is weak evidence of agreement, especially in a hierarchy. It can mean assent, uncertainty, deference, fatigue, a lost connection, or belief that the decision is already made. Invite views before declaring consensus. Written follow-up can help participants who need time or could not safely interrupt.

The accountable decision-maker may choose against the majority. Say so plainly: “After consultation, I have selected Option B because it best meets the safety threshold, despite the operations team’s preference for Option A.” Honest authority is more trustworthy than invented unanimity.

Material dissent should be bounded, not dramatized. Record the substance and decision relevance, not a transcript of personalities. Dissenters do not need permanent exemption from implementation. They do need assurance that the record will not later claim they endorsed an assumption they challenged.

4. Reclassifying an outcome after the fact

Classification determines the standard. A project can become a “pilot” after it underperforms, reducing expectations that were used to secure funding. An experiment can become an “operational launch” after a favorable result, bypassing the controls for wider deployment. A missed target can become “directional.” A cancelled initiative can become “completed discovery.”

Classification can genuinely change as learning occurs. A limited launch may reveal that further testing is necessary; leaders can convert the next phase into a pilot. The distinction is between a prospective decision and a retrospective relabeling. Organizational research on issue categorization illustrates the broader mechanism: how an issue is framed and categorized can shape the action that follows.4

Before work begins, define the category in operational terms:

  • What decision will this work inform?
  • What population, duration, and exposure are authorized?
  • What safeguards and approvals apply?
  • What constitutes completion?
  • What evidence permits expansion, revision, or stopping?
  • Which promises are exploratory and which are commitments?

If the category changes, state the effective date and consequences. Do not use the new label to erase obligations or harms that arose under the old one.

Post hoc explanation has a close cousin in research: presenting a hypothesis developed after seeing results as though it had been specified beforehand. The term HARKing—hypothesizing after results are known—was developed in the context of scientific reporting.5 Organizational decisions are not journal articles, but the underlying transparency lesson travels well. Discovery after the result is valuable. Label it as discovery, then test it where the stakes warrant.

Financial and operational categories need similar care. Moving costs between “run” and “change,” incidents between severity levels, or customers between segments may be valid under established rules. If classification changes only after the number is known, require an independent review and show the result under both treatments where possible.

Never assume a benign label removes real duties. Calling work a pilot, learning exercise, or prototype may not change obligations concerning safety, employment, privacy, research, contracts, or regulated activity. Qualified owners should determine the applicable requirements.

5. Presenting too few alternatives

Control of the option set is control of the decision. If a paper offers only “approve the full program” or “do nothing,” approval can look inevitable even when phased, smaller, reversible, or differently owned approaches exist.

Too many options can also obscure. Decision-makers need a manageable set of materially distinct choices, not ten minor variants. The safeguard is to show the credible range and explain why other plausible paths were excluded.

A useful option paper includes:

  1. the problem and decision deadline;
  2. the baseline or “continue current course” option, with real consequences;
  3. at least one meaningfully different approach where available;
  4. assumptions, benefits, costs, risks, reversibility, and people affected;
  5. dependencies and authority;
  6. the recommendation and why;
  7. uncertainty and what evidence could change the choice.

Avoid the fictional “do nothing” baseline. Current operations require money, labor, and risk acceptance. Describe them. A favored option should not receive a complete design while alternatives are represented by vague disadvantages.

Framing research demonstrates that equivalent outcomes described as gains or losses can produce different preferences.6 This does not mean every persuasive presentation is manipulation. Decisions require framing. Make the frame inspectable: “We have described this as avoiding service loss rather than producing savings because the current capacity expires in June.” Where a different reasonable frame changes the preference, show both.

Ask who generated the alternatives. A team that built one solution may search narrowly around it. Invite operations, users, finance, risk, or an independent reviewer to propose constraints and alternatives before the paper hardens. Do this proportionately; a routine purchase does not need a strategy exercise.

Preserve a no-go or pause option when authority genuinely permits it. If delay is impossible because of law, safety, contract, or expiry, state that constraint and confirm its source. Then the decision may concern how, not whether.

Targets produce stories before anyone writes the report

A target does more than measure performance. It directs attention, status, effort, and explanation. People learn which delays must be justified, which quality failures receive silence, and which result earns access to leaders. The eventual narrative is partly built into the incentive system.

Suppose a repair service promises that 90 percent of jobs will close within two days. The target may improve speed. It may also encourage teams to classify a complicated repair as several separate jobs, close a case before the customer confirms the result, or avoid accepting work likely to miss the threshold. These responses range from sensible process design to clear misconduct. The target alone does not tell us which occurred.

Before attaching consequence, identify the behavior the target should encourage and the most likely substitute behavior. Ask the people doing the work how they would improve the number if the underlying purpose did not matter. Their answers often reveal weak definitions faster than a technical audit.

Use three protections. First, pair the target with a constraint or quality check where a predictable trade-off matters: speed with first-time resolution, growth with returns, output with safety, or case closure with recurrence. Second, review the tails and exceptions, not only the average. Third, give teams a safe route to explain when meeting the target would harm the real purpose.

Targets should have governance. Who owns the definition? Who validates data? Who approves exclusions? When will the threshold be reviewed? If local managers can reclassify cases that determine their own reward, use independent sampling or approval proportionate to risk.

Do not assume every unexpected behavioral response is gaming. Workers may adapt because the target contradicts another duty, because tools are inadequate, or because the recorded process differs from real demand. Investigate the mechanism. Punishing the visible workaround while preserving an impossible target teaches people to hide better.

Incentives extend beyond money. Promotion, praise, executive attention, public rankings, contract renewal, and relief from scrutiny all tell people which story to produce. A leader who celebrates launches but never maintenance should expect reports full of launches. A board that asks only whether a project is green should expect definitions that keep it green.

Counter this by rewarding accurate bad news, early stopping, prevented harm, corrected forecasts, and knowledge transferred. Leaders can say, “This team missed the original date because testing found a serious access flaw; the discovery and transparent reset are evidence of control, not a reason to recolor the history.” That does not make every miss admirable. It makes learning and truth compatible with status.

Finally, keep reward and evaluation intervals long enough to capture consequence. A sales incentive paid before cancellations, a project award given before benefits review, or a speed ranking published before quality data arrives builds premature success into the organization’s memory. Where delay is impractical, treat recognition as provisional and return to the complete result.

Numbers do not speak; nor are they merely opinions

It is tempting to respond to selective narrative by saying “just show the data.” Data are produced through definitions, instruments, systems, and human choices. That does not make them arbitrary. Some measurements are more valid, reliable, complete, and decision-relevant than others.

Treat measurement as an argument with evidence. The presenter should be able to explain why a measure represents the claimed concept, how it was generated, what error or missingness exists, and which conclusion it can support. The audience should be able to ask whether another plausible measure changes the result.

Use visual discipline. Begin axes at a meaningful point or clearly signal a restricted range. Keep units and time periods consistent. Show absolute numbers alongside percentages where scale matters. Explain exclusions near the chart. Do not use decorative precision: a forecast reported to one decimal place may still have wide uncertainty.

For high-stakes analysis, preserve lineage from source to result. Who extracted the data? Which version, query, and transformation were used? Who can reproduce it? Restrict personal or confidential data to authorized users and collect only what the purpose requires. Transparency does not mean publishing raw sensitive records.

Narrative evidence needs lineage too. If a report says “staff welcomed the change,” identify the source: a representative survey, voluntary focus groups, manager impressions, or selected quotations. Each may be useful, but they warrant different confidence.

Precommitment without rigidity

One defense against a story that follows the result is to state important rules in advance. Before launch, record the primary outcome, threshold, review date, exclusions, decision owner, and likely responses. This is a practical form of precommitment.

Precommitment should not force leaders to ignore reality. Add a change rule: criteria may be revised when an assumption fails or new evidence emerges, provided the reason, authority, date, and effect on interpretation are recorded. Show old and new analyses where feasible.

Use a “decision time” view and a “current knowledge” view:

  • Decision time: Was the original choice reasonable with the information and authority then available?
  • Current knowledge: Given what happened, what should we believe and do now?

These questions prevent two common errors. The first condemns every bad outcome as a bad decision. The second defends continuation merely because the original choice was reasonable.

Teams should also define a stop or review threshold. What level of harm, cost, delay, or failed benefit triggers reconsideration? Escalation-of-commitment research suggests that sunk investments and responsibility for an earlier decision can shape continued commitment.2 A pre-agreed review gives leaders permission to change course without presenting adaptation as defeat.

A response ladder when the story and record diverge

Recover the promise. Find the approved purpose, measures, assumptions, scope, and constraints.

Reconstruct changes. Identify authorized revisions and the dates they took effect.

Show the complete result. Include material benefits, costs, access, quality, delay, risk, and unintended consequences at a useful level.

Separate facts from judgments. Label observations, estimates, causal interpretations, and recommendations.

Restore decision ownership. State who chose, who advised, who dissented on material grounds, and who must implement.

Expand the option set. Add credible alternatives or explain why they are unavailable.

Decide forward. Use present evidence to continue, adapt, pause, stop, repair, or investigate. Do not make historical reputation the only objective.

If you suspect deliberate falsification, unauthorized reclassification, fraud, safety concealment, or misleading regulatory or financial reporting, preserve authorized evidence and use the proper specialist or protected route. Do not alter source records or conduct an informal investigation beyond your authority.

For managers: run a narrative audit

Take a consequential proposal or review and ask a person not responsible for the preferred story to read it. Have them mark:

  • every evaluative word—successful, delayed, efficient, strong, accepted;
  • the measure or evidence supporting it;
  • the missing comparator or denominator;
  • changes from the original promise;
  • claims of agreement and the decision record behind them;
  • category changes;
  • plausible alternatives not shown;
  • statements of causation that may be only correlation or interpretation.

The reviewer should not rewrite the document into false balance. If the evidence strongly supports success, say so. Their job is to make the route from evidence to conclusion visible.

At the meeting, begin with the decision required, not a long history designed to exhaust scrutiny. Give members materials with enough time, identify changes since circulation, and record the chosen option and reasons. Afterward, publish a proportionate account for affected people. Confidentiality may limit details, but it should not require claiming consensus that did not exist.

Review communication later. Did the public or internal summary preserve the material conditions and trade-offs? A clear decision can be distorted during cascade as each layer simplifies it. Provide a short canonical account: what was decided, by whom, why, what remains uncertain, what happens next, and where questions go.

Returning to Northbridge

The benefits committee asks the program director to restore the original three measures. Missed visits improved from 7 to 4 percent, although part of the change reflects a shift in scheduled case mix. Median customer waiting time rose by two days. Dispatcher correction time increased during the first four months and then began to fall, but remains above baseline.

The committee records two separate judgments. The deployment met its technical scope and produced a meaningful reduction in missed visits. It has not yet delivered the complete business case. The board’s modernization objective supports continued investment but does not replace the customer and workload promises.

Three options are considered: continue unchanged, roll back, or keep the system while pausing expansion and redesigning the exception rules with dispatchers. The committee selects the third. It sets a twelve-week review with the original measures, case-mix information, and a dispatcher workload sample.

The final communication calls the project neither a triumph nor a disaster. It says what worked, what did not, what was learned, and what decision follows. That account is less dramatic. It is also more useful.

Practice: give one outcome two honest readings

Choose a completed, non-sensitive project or decision. Write two short interpretations that are both supported by the available evidence—for example, a delivery view and a user view, or a cost view and a quality view.

Then record:

  1. the original promise and source;
  2. the main measure and its definition;
  3. one material measure that was not emphasized;
  4. any goal, scope, or category change;
  5. who made the final decision;
  6. the level of support and any relevant dissent;
  7. at least one credible alternative;
  8. which interpretation you find more persuasive and why;
  9. what evidence would change your conclusion;
  10. the forward decision the evidence supports now.

The exercise is not a demand for indecision. It is practice in making a strong conclusion without hiding the choices that produced it.

Notes

Footnotes

  1. Ziva Kunda, “The Case for Motivated Reasoning,” Psychological Bulletin 108, no. 3 (1990): 480–498, https://doi.org/10.1037/0033-2909.108.3.480.

  2. Barry M. Staw, “Knee-Deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action,” Organizational Behavior and Human Performance 16, no. 1 (1976): 27–44, https://doi.org/10.1016/0030-5073(76)90005-2; David J. Sleesman et al., “Cleaning Up the Big Muddy: A Meta-Analytic Review of the Determinants of Escalation of Commitment,” Academy of Management Journal 55, no. 3 (2012): 541–562, https://doi.org/10.5465/amj.2010.0696. 2

  3. Jonathan Baron and John C. Hershey, “Outcome Bias in Decision Evaluation,” Journal of Personality and Social Psychology 54, no. 4 (1988): 569–579, https://doi.org/10.1037/0022-3514.54.4.569.

  4. Jane E. Dutton and Susan E. Jackson, “Categorizing Strategic Issues: Links to Organizational Action,” Academy of Management Review 12, no. 1 (1987): 76–90, https://doi.org/10.5465/amr.1987.4306483.

  5. Norbert L. Kerr, “HARKing: Hypothesizing After the Results Are Known,” Personality and Social Psychology Review 2, no. 3 (1998): 196–217, https://doi.org/10.1207/s15327957pspr0203_4.

  6. Amos Tversky and Daniel Kahneman, “The Framing of Decisions and the Psychology of Choice,” Science 211, no. 4481 (1981): 453–458, https://doi.org/10.1126/science.7455683.