Companion to Chapter 13. Chapter 13 names vibe charting — the verification-choice failure mode of accepting AI output because it feels right rather than because its evidence supports the claim — and prescribes the discipline that catches it: five diagnostics (three or more yes-answers is the signal), four failure modes (hallucinated insight, misleading default, over-styled output, confidently-wrong narrative), three cognitive forcing functions (predict-before-reveal, evidence requirement, restatement-before-action) that interrupt the underlying bias mechanisms, and the explicit ship / iterate / discard decision rule. Underneath is the bias ratchet — automation bias plus anchoring plus confirmation bias compounding across a session with no natural correction. This page carries what the book cannot: curated further reading beyond Appendix C on the cognitive-science literature the chapter builds on, a five-question self-check quiz on vibe charting and the five diagnostics, and a hands-on activity that scores a shipped artifact against the diagnostics.
Use the further reading to see the bias mechanisms in the primary literature. Use the quiz to move vibe-charting recognition into reflex. Use the activity to feel where a shipped artifact from your own team is on the vibe-charting spectrum — the answer is usually more surprising than the analyst expects.
Further reading
Books and papers beyond the book’s own bibliography (Appendix C, updated in errata) on cognitive-bias mechanisms and human-AI over-reliance.
- Parasuraman, R., & Manzey, D. H. (2010). «Complacency and Bias in Human Use of Automation.» Human Factors, 52(3), 381–410. The same paper Ch 1 cites — here it also underlies the automation-bias leg of Ch 13’s bias ratchet. Reads well as a survey of the empirical work across aviation, medicine, and process control. Available via doi.org/10.1177/0018720810376055.
- Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux. The System 1 / System 2 framing under the «cognitive forcing function» concept. Chapter 3 («The Lazy Controller») explains why explicit forcing functions are what interrupt System-1 acceptance of well-formatted output. Chapters 6, 7, 9 for the anchoring and availability biases the ratchet composes.
- Buçinca, Z., Malaya, M., & Gajos, K. (2021). «To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-making.» Proceedings of the ACM on Human-Computer Interaction, 5, CSCW1, Article 188. The empirical study the «cognitive forcing function» term comes from. Buçinca et al. tested three specific forcing functions on human-AI decision-making tasks and measured reduction in AI over-reliance. Ch 13’s three forcing functions are the visualization-domain adaptation. Free at doi.org/10.1145/3449287.
- Nickerson, R. S. (1998). «Confirmation Bias: A Ubiquitous Phenomenon in Many Guises.» Review of General Psychology, 2(2), 175–220. The canonical review of confirmation bias — the leg of the bias ratchet that keeps the analyst from updating their mental model when Copilot is wrong. Long and thorough; the summary in the abstract stands alone for readers who want the shape without the full survey.
- Karpathy, A. — on «vibe coding.» Andrej Karpathy (co-founder OpenAI, former Tesla AI) coined the term «vibe coding» in a February 2025 tweet describing coding sessions where the developer accepts LLM-generated code on the feel of the response rather than the review of the diff. Vibe charting is the direct visualization-domain analog. Karpathy’s posts at x.com/karpathy. The Verification Habit sister book’s Chapter 12 catalogs vibe coding in its original coding-domain form.
Last curated: 2026-07-18. Reviewed quarterly. Suggestions welcome via the errata page.
Self-check quiz
Five questions on Ch 13’s load-bearing concepts. Answer them out loud before opening the reveal.
-
Q1. Define vibe charting in one sentence. What is it not?
Show answer
Vibe charting is the verification-choice failure mode of accepting AI output because it feels right rather than because its evidence supports the claim. It is not the same as using AI to build charts. Using AI is a tool choice; vibe charting is a discipline choice — specifically, the (often unconscious) choice to evaluate output by intuition rather than criteria. The Yuki 1:47 a.m. scenario in the chapter opener is the archetypal case: four Copilot iterations, no principled review, accepted because it looked right, shipped because it looked right, and nobody caught the accessibility failure because nobody checked.
-
Q2. Name the five diagnostics for vibe charting. What is the threshold for «vibe charting is happening»?
Show answer
(1) Ambiguity burden — is the request clear enough that a third party could state what would make the output pass or fail? (2) Evidence chain — does each load-bearing claim trace back to specific evidence in the data? (3) Telemetry trio — acceptance rate, verification prompts per session, time-to-acceptance; each low signals over-acceptance. (4) Repair loop maturity — are revisions structured repairs (named failure, targeted fix, verification) or prompt roulette (reroll and hope)? (5) Over-reliance signals — behavioral patterns of agreement without review. Threshold: three or more yes-answers on the same workflow signals vibe charting is happening.
-
Q3. Name the four AI-chart failure modes Ch 13 catalogs, and one signature detail for each.
Show answer
Hallucinated insight — a chart or Smart Narrative claim the data does not support; often confidently phrased. Misleading default — right data, wrong rendering (truncated y-axis, dual axis, stacked-area-as-trend, categorical palette on ordered data). Over-styled output — visual sophistication masking analytic poverty (Yuki’s rounded-corner soft-blue dashboard); polish signals effort but is not the same as review. Confidently-wrong narrative — Smart Narrative that reads cleanly but names causal claims the visual does not visibly demonstrate; the cross-surface generalization of Ch 9’s wrong-narrative failure. Each failure mode has its own catch move (Table 13.1 in the book).
-
Q4. Name the three cognitive forcing functions and, for each, name the specific bias it interrupts.
Show answer
Predict-before-reveal — write your one-sentence prediction of what the AI will produce before looking at the output. Interrupts anchoring bias: your prediction is not anchored on the AI’s output because the AI’s output does not exist yet. Evidence requirement — before accepting any claim, identify one specific citable piece of evidence supporting it. Interrupts automation bias: the requirement forces the analyst to look at the underlying data rather than accepting the surface claim. Restatement before action — before shipping, restate each visual’s claim in your own words without looking at the AI’s framing. Interrupts confirmation bias: the analyst’s own words expose whether they were reading the AI’s claim or reading the visual.
-
Q5. Ch 13 makes the ship/iterate/discard decision explicit rather than a binary accept-or-reject. Why the three-way, and which option is most often skipped?
Show answer
Because AI-directed artifacts can be salvageable (iterate), unsalvageable-but-fixable-with-a-rebuild (discard), or shippable (ship) — not just yes/no. The most-often-skipped option is discard. Analysts default to iterate because a draft feels like progress — discarding feels like waste even when the iteration would cost more than a rebuild. Discard means going back to Crystallize with a sharper specification, not patching the current output. In the sibling book’s framing, an iteration that patches over a fundamentally under-specified Assemble output is prompt roulette in a longer form.
Yuki’s 1:47 a.m. session, up close
Yuki’s scenario in Chapter 13 is the canonical vibe-charting case: a late-night quarterly-review dashboard, four prompts («Make it look better», «Warmer», «Less corporate», «More like the Stripe dashboard but not exactly»), zero verification prompts, sub-minute time-to-ship. The shipped dashboard runs at a 3.1:1 text-contrast ratio — below the WCAG-recommended 4.5:1 minimum — and nobody notices because nobody checks. All five of the Chapter 13 diagnostics fire on this session; the 5-of-5 pattern is what makes the failure nameable.
- Diagnostic 1 — ambiguity burden. None of Yuki’s four prompts is testable. «Make it look better» has no pass/fail condition; neither does warmer, less corporate, or more like Stripe. Yuki carries the whole ambiguity burden and evaluates by feeling because there is nothing else to evaluate against.
- Diagnostic 2 — evidence chain. No oracle exists. WCAG’s
4.5:1contrast standard was available and would have caught the3.1:1miss in under a minute, but no criterion is named upstream so the check is never triggered. - Diagnostic 3 — telemetry trio. Acceptance rate near
1.0, zero verification prompts across the session, time-to-ship under a minute after the final render. Any one number in isolation is not diagnostic; all three together on a non-trivial dashboard are. - Diagnostic 4 — repair-loop maturity. Four rerolls, zero repairs. Each prompt restates a feeling instead of naming what was wrong against a rubric. The information from each round is discarded; the next round starts from scratch. This is the prompt-roulette pattern.
- Diagnostic 5 — over-reliance signals. High-stakes audience (quarterly review), sub-minute acceptance, Smart Narrative unchecked against the visuals, polish reading as truth. Three of the three over-reliance signals fire.
Five of five. That is the vibe-charting verdict. Chapter 13’s point is not that the reroll cycle is inherently wrong — sometimes iteration is the right move — but that the reroll cycle here happened without any of the diagnostics ever being consulted. The moment of acceptance was invisible; there was no gate to skip because no gate was set. The five diagnostics exist to make the moment of acceptance visible, so the analyst can choose whether to verify or skip with the cost of the skip in front of them.
Inspect the session. The full reroll timeline and diagnostic scorecard are available as a flat CSV so you can rebuild the scorecard yourself, or score your own session against the same template. Download yuki-session.csv · Dataset README and schema.
Try this activity — score a shipped artifact on the five diagnostics
This activity runs the Ch 13 diagnostics against something your team shipped recently. Total time: 30–45 minutes; no Power BI required — you can do this against a screenshot and your session notes.
-
Pick the artifact. A Copilot-generated report, an AI visual on a hand-built page, a Q&A answer that made it into a leadership deck, or a dashboard composition — anything AI-directed your team has shipped in the last quarter. Have the artifact + your session notes (if you kept any) available.
-
Score yourself yes/no on each of the five diagnostics.
- Ambiguity burden: is the original request clear enough that a fresh third party could name what would make the output pass or fail? (No = ambiguity burden high.)
- Evidence chain: for each load-bearing claim on the artifact, can you name the specific data point or measure it traces to? (No = evidence chain broken.)
- Telemetry trio: did you accept too quickly (under thirty seconds for a complex artifact)? Did you skip verification prompts? Did you make zero edits? (Any yes = telemetry-trio signal fires.)
- Repair loop maturity: did each revision address a named defect, or was it prompt roulette (reroll and hope)? (Roulette = yes on this diagnostic.)
- Over-reliance signals: did trust in the AI replace review at any decision point? (Yes = over-reliance signal fires.)
Count your yeses. Three or more is the vibe-charting signal.
-
Run the four-pattern check. For each visual or narrative on the artifact, ask which of the four failure-mode patterns (if any) fires. Some artifacts pass cleanly on this check even when the diagnostics fire; others surface one or two failure patterns. Note them by name.
-
Run the three forcing functions retrospectively. Could you have predicted what Copilot produced before it produced it? Can you cite one specific piece of evidence for each major claim? Can you restate each visual’s claim in your own words without looking at the AI’s title or narrative? Note where the answer is no.
-
Make the decision. If the artifact has already shipped, write the would-revise-if clause: what specific evidence, if it surfaced, would cause you to revise or retract? If the artifact has not shipped yet, write the ship / iterate / discard decision and the deciding evidence check.
-
Update one upstream discipline. If any pattern surfaced, name which per-surface move from Ch 9–12 would have caught it earlier. Commit to running that move on the next artifact of the same type — the compounding is what turns individual vigilance into team discipline.
-
Reflect (three prompts). Write one sentence in response to each:
- Which diagnostic scored highest, and does the result match your prior sense of the artifact’s defensibility?
- Which forcing function, run retrospectively, would have caught the most?
- If you had to choose one team-level change to prevent this failure mode next time, what would it be?
The reflection is the activity’s payoff. Skipping it turns the exercise into paperwork.
Optional extension: repeat the activity on a widely-cited chart or dashboard you have seen recently (a vendor demo, a colleague’s screen, a public news graphic). The diagnostics score things you did not build the same way as things you did — running them on someone else’s artifact is calibration for running them on your own.
Related site resources
- Yuki’s session dataset: the reroll timeline and diagnostic scorecard behind Figure 13.2 — four iterations, five diagnostics, downloadable as a flat CSV.
- Chapter 13 — Practice-on-your-own solutions: worked answers to the disagreement scenario and the discard-call problem.
- Chapter 9, 10, 11, 12: the per-surface verification moves the failure-mode taxonomy generalises. Each chapter’s move catches its surface’s failure earlier than the cross-surface diagnostics do.
- Prompt Library: the CSAR prompts include the predict-before-reveal and confirmation-turn patterns that operationalise the forcing functions.
- Chapter 14 companion: the next chapter — responsible AI in analytics closes the vibe-charting failure loop with organisational discipline (reliance drills, decision boundaries).
- Errata: if the community surfaces a documented instance of one of the four failure patterns in the book’s worked examples, it goes here first.
Cited in the book
Chapter 13’s bibliography lives in the book’s Appendix C (living version in errata). The chapter’s operative references:
- Parasuraman, R., & Manzey, D. H. (2010). Complacency-bias mechanism.
- Tversky, A., & Kahneman, D. (1974). «Judgment under Uncertainty: Heuristics and Biases.» Science, 185(4157), 1124–1131. Anchoring-bias origin.
- Nickerson, R. S. (1998). Confirmation bias review.
- Kahneman, D. (2011). Thinking, Fast and Slow. System-1 / System-2 framing.
- Buçinca et al. (2021). Cognitive forcing functions empirical study.
- The Verification Habit (sister book), Chapter 12. Vibe-coding origin and the bias-ratchet mechanism generalised.