Graphical perception, from the 1984 experiment to a chart you can defend.
Rank 6 of 6 — shading0 / 22 blocks
What you'll be able to do
Take a chart you are responsible for, state the decoding task it asks of its reader, predict which encodings will fail and roughly by how much, and defend the redesign with evidence — including evidence you collected yourself.
Not "understand graphical perception." The deliverable is a stack of decision records and one rebuilt chart with a measured defense.
Where you're starting from
Assumed: you build charts and progress meters in HTML/SVG regularly; log scales, robust estimators and bootstrap confidence intervals are familiar or a one-line refresher; you've absorbed the folk rule that bar charts beat pie charts. Assumed absent: any contact with the primary literature, and any experience running a perception experiment. If that's wrong in either direction, the fading in Units 3–5 will feel off — adjust by skipping the worked examples or by slowing down.
What this course does not do
No color science. Colormaps, gamut, contrast — a different course. Color appears here only as a low-ranked channel.
No uncertainty visualization. Error bars, hypothetical outcome plots, quantile dotplots: out of scope, and genuinely important.
No task taxonomies. Brehmer & Munzner's why/what/how is the natural sequel; this course stops at the elementary level on purpose.
No aesthetics, branding, or chart-junk arguments. Where the evidence runs out, the course says so rather than substituting taste.
No accessibility treatment. Named where it collides with the material, never taught.
No interaction. Everything here is about static decoding.
Time
About 20 hours across five units. That figure already has the 2.5× correction applied — the raw design is roughly 8 hours, and learners consistently take far longer than the raw estimate. Every block below carries its own estimate. Blocks are sitting-sized: you can open one, close it, and stop.
One more thing the attrition data is clear about, and it has nothing to do with the material: cohort formats complete at substantially higher rates than solo self-paced study, and social isolation is a commonly cited reason people abandon courses. If you actually want to finish this, the highest-leverage move available is to talk one other person into working through it alongside you.
How claims are marked
Every load-bearing claim carries a marker. S4 a sourced fact, keyed to the numbered list in Sources. Inference my reasoning, in no source. Contested the field genuinely disagrees. Single study real but unreplicated. Vendor framing true, but authored by someone selling the thing.
The last one matters more here than you'd expect. Two of the six central sources were written by people at Tableau.
Unit One
Run the experiment on yourself
~3 h 15 m · 4 blocks · threshold 1 opens here
Evidence this unit produces
Your own accuracy ladder across four encodings, and Decision Record #1: one real chart you own, decomposed.
Most people meet this material as a ranked list they are asked to believe. That's the half-crossing — you can recite "position beats angle" for years without it changing a single decision, because a list is not evidence, it's the residue of evidence.
So Unit One starts at the other end. You will sit the 1984 experiment before you read the 1984 paper. Twenty-four stimuli, four encodings, quick visual judgments, no measuring. Then you'll compute your own log-error midmeans and put your ladder next to theirs.
Expect your ladder to be noisier than the published one and possibly out of order in places. Six trials per encoding against their fifty-one subjects times ten ratios S1 is not a fair fight. That gap is the unit's real subject.
JudgeSit Round 1 in the Lab: 24 stimuli, four encodings.40 m
Go to the Lab tab and run Round 1 before reading anything else. The instructions there mirror the original: identify which of the two marked values is smaller, then make a quick visual judgment of what percentage the smaller is of the larger. Do not measure with a finger, a pencil, or a screenshot tool. Do not go back and revise.
The four types are position along a common scale, length in a divided bar, angle in a pie, and circular area in a bubble chart. Heer & Bostock unified all four into one response format so they could be compared directly S2; Cleveland & McGill ran position-length and position-angle as two separate experiments and explicitly warned that the means from one should not be compared with the means from the other S1. You're following the later format.
Round 1 complete — my ladder is on the screen.
ReadCleveland & McGill 1984, sections 1 through 4.5.1 h
Full text, free scan (PDF) — Journal of the American Statistical Association 79(387), September 1984, pp. 531–554. S1
Read for three things and let the rest wash past:
Section 2 — the ten elementary tasks, and their care about what they are not claiming. Note that they say plainly the ten are not distinct, not exhaustive, and overlapping; that a circle shown to a viewer has an area but also a diameter and a circumference.
Section 3 — the ordering, and the honest paragraph about where it came from. It is not derived. It is assembled from psychophysics, their own tinkering, and prior experiments, and then tested.
Section 4 — the experimental design and the analysis. This is the part you just lived. You'll return to it in Unit Two.
One sentence to hold onto, from their Section 6: they had no proof that laboratory results work in the field, only that it appeared plausible. The founding paper of evidence-based chart design is more careful about its own limits than most of what has been built on top of it. S1
Read sections 1–4.5.
BuildDecision Record #1 — the whole artifact, crudely, today.1 h 15 m
Pick one chart you are actually responsible for. Not a hypothetical, not a famous bad chart from the internet. Something that will be looked at this week by someone whose decisions you care about.
Fill in the record in the Records tab. It is deliberately crude — five fields, no template, no scoring. By Unit Four you will find it embarrassing, and that is the measurement.
This is the whole task in miniature: what value is encoded, in what channel, what judgment the reader makes, what you predict the error looks like, and what you'd change. The capstone is this, done well, with numbers.
Decision Record #1 written.
WriteWhere my ladder disagrees with theirs — and one reason it might.20 m
In the Records tab, under Unit 1 note. Two paragraphs, maximum.
Resist the two easy answers ("I'm just bad at this" and "the ranking is wrong"). The interesting answers are about the instrument: how many trials you ran, how the values happened to fall, whether you were rushing, whether one encoding's stimuli happened to draw easier ratios. Cleveland & McGill found the error depended on the true percentage itself S1 — six trials cannot average that out.
Unit 1 note written.
Unit Two
The measure, not the ranking
~3 h 50 m · 4 blocks · threshold 1 closes here
Evidence this unit produces
A short piece called "how the estimator changed my answer": your own ladder recomputed three ways, with the ranking changes marked.
Here is the thing that separates crossing this threshold from mimicking it. Anyone can repeat the ranking. Almost nobody can say why the accuracy measure is log₂(|judged − true| + ⅛) and not |judged − true|, or why the summary is a midmean and not a mean. And if you can't defend the measure, you don't have evidence — you have a hierarchy you're quoting.
This is the liminal stretch of the course. Somewhere in the middle of the derivations you will feel that this is pedantry standing between you and the useful part. That feeling is the crossing. The useful part is this: the ranking is a summary statistic, and summary statistics have choices in them.
DeriveFive exercises on the error measure — including one that is wrong.1 h 15 m
Exercise 1 · worked example
A subject judges 60% when the true value is 45%. Compute the log absolute error.
Nothing hidden. Note the units: a difference of 1.0 on this scale is a doubling of absolute error. Cleveland & McGill chose base 2 because average relative errors tended to change by factors of less than ten S1 — base 10 would have compressed everything into a narrow band.
Exercise 2 · completion problem
A subject judges 47% when the true value is 47%. The absolute error is 0. What is log₂(0)?
Fill in the consequence, then the fix:
Confidence before you check:
log₂(0) is undefined — it runs to negative infinity. Perfect answers would blow up the scale. Adding ⅛ before the log floors the measure at log₂(0.125) = −3, so an exact hit scores −3 rather than nothing at all.
Cleveland & McGill's own reason: the ⅛ was added to prevent a distortion of the scale at the bottom end, because absolute errors in some cases got very close to zero S1. Note what this quietly does — it makes the difference between a 0-point error and a 1-point error much smaller than the difference between a 10-point and a 20-point error. The measure cares about relative error, and stops caring near perfection.
Exercise 3 · independent
Cleveland & McGill report that in the position-length experiment, the larger of the two length values was 1.32 log units above the smallest of the three position values, and the smaller length value was 0.51 log units above the largest position value S1.
Convert both to plain factors, and state the range in one sentence a colleague would understand.
Confidence:
21.32 = 2.5. 20.51 = 1.4. Their own sentence: average errors for length judgments were 40% to 250% larger than those for position judgments S1.
Notice how much weaker that is than "bar charts are better than pie charts." It's a range, it's about one experiment's specific stimuli, and the low end — 40% — is the kind of difference a slightly better label could swamp.
Exercise 4 · independent
In the position-angle experiment the gap between angle and position was 0.97 log units, with a 95% bootstrap interval of (0.79, 1.15) S1. Separately, 88% of the large errors — those above 4 log units — occurred on angle judgments, at a rate 7.3 times that of position S1.
These two numbers describe the same data. Why would you quote the second one to a product team and the first one to a statistician?
Confidence:
The mean difference (a factor of about 1.96) describes the typical reader. The large-error rate describes the tail — how often someone is badly, decision-changingly wrong. A product team is usually insuring against the tail; the mean is nearly irrelevant to them. A statistician wants the interval because it tells them whether the effect is real at all.
Inference The second framing is also the more honest one for most business charts, and it is almost never the one that gets quoted. Cleveland & McGill reported both. The folklore kept the ordering and dropped both numbers.
Exercise 5 · find the break
A colleague drops this into a design doc. Every sentence sounds informed. Three of them are wrong. Find them.
Cleveland & McGill measured a 0.97 log-unit gap between angle and position, so pie charts are about 97% less accurate than bar charts. They also found position errors of 1.40 against length errors of 2.72, so length is roughly twice as bad. And since their 51 subjects split into technical and non-technical groups, with the technical group performing better, the gap should be smaller for our engineering audience.
Confidence:
0.97 log units is not 97%. It is a factor of 20.97 = 1.96 — angle judgments carried roughly double the error S1. The unit was silently dropped, and "97% less accurate" also inverts the direction: the measure is error, so more is worse, not less.
The two figures come from different experiments. 1.40 and 2.72 are position-length; 0.97 is position-angle. Cleveland & McGill state it would be inappropriate to compare means across the two S1. Within position-length the gap is 1.32 log units, a factor of 2.5 — not "roughly twice."
The expertise claim is invented. They reported that they did not detect any difference between the technical and non-technical groups, and treated the subjects as one homogeneous sample on the grounds that these are very basic tasks carried out in everyday activities S1. This is citation invention in miniature — a hypothesis becoming a fact by being attributed.
The third one is the dangerous one, because it is the sentence a reader is least likely to check and the one most likely to shape a decision.
All five exercises attempted before revealing.
ReadHeer & Bostock 2010 — the replication, and what moved.45 m
Read Experiments 1A and 1B closely; skim the Mechanical Turk performance analysis unless the methodology interests you.
The ranking held. Position still significantly outperformed length. S2
Types 1 and 2 came closer together than in the original, which they attribute to a smaller display shrinking the distance effect. S2
Angle did not underperform length, though theory predicted it should. They say so plainly. Hold that; Unit Three is built on it. S2
Rectangular areas were judged worst when the two rectangles were square — aspect ratio 1:1. Which means squarified treemap algorithms are optimizing toward the hardest case, and readers benefit from the algorithm's failure to fully succeed. S2Single study
Practical numbers worth stealing: gridlines want at least 8 pixels of separation, and pushing chart height past about 80 pixels bought no accuracy on a 0–100 scale. S2Single study
Also worth noticing as a matter of craft: without a qualification task, over 10% of their responses were unusable, and they attribute that to confusion rather than cheating S2. Any time you put a chart in front of people to test it, most of what you're measuring at first is whether they understood the question.
Read Experiments 1A and 1B.
BuildRecompute your own ladder three ways.1 h 30 m
In the Lab, under your Round 1 results, there is an estimator switch. It recomputes your ladder under three summaries: midmean (what they used), plain mean, and median. It also lets you drop the ⅛.
Work through this deliberately:
Record your ranking under each of the three summaries. Note every place two encodings swap.
Find the single trial that moves the most under the switch, and look at it again. It will usually be one bad guess.
Turn off the ⅛ and see which cell breaks or distorts.
Deliverable, in Records: a piece called how the estimator changed my answer. Three or four paragraphs. It should end with a sentence you'd be willing to say out loud in a design review about how much confidence a ranking built on six trials deserves.
Inference The midmean — the mean of the middle two quartiles — was chosen because their log errors showed frequent outliers, mild skew, and clumping at multiples of five, since people like round numbers S1. With six trials per type, this implementation drops the highest and lowest and averages the middle four. That is a much blunter instrument than theirs and you should treat your own ladder accordingly.
Ladder recomputed three ways; piece written.
RecallCold check — no scrolling back.20 m
Retrieval · Unit 1 material
Without reopening Unit One or the Records tab: what channel does the chart in your Decision Record #1 use, and what judgment did you say its reader makes? Write it from memory first, then go compare.
Confidence:
Open Records and compare. If the two match, the record was doing real work. If your memory is vaguer than what you wrote, the record was doing the work for you — which is fine, but means you haven't internalized it yet. If your memory is sharper than the record, the record is underwritten; go fix it now.
Cold check done.
Unit Three
The task, not the chart
~4 h 25 m · 5 blocks · threshold 2
Evidence this unit produces
Round 2 of the experiment — the same charts, a different question, your own error moving. Plus Decision Records #2, #3 and #4, now written task-first.
Everything to this point could still collapse into the folk version: rank the channels, pick a high one, ship. Unit Three is where that stops working, and it stops working because of a 1987 result that the folklore simply lost.
Simkin & Hastie showed 40 subjects simple bar charts, divided bar charts and pie charts, and asked two different kinds of question about them. For a comparison judgment — this segment against that one — the expected order held: position, then length, then angle. For a proportion-of-the-whole judgment, angle in a pie chart was as accurate as position in a simple bar chart, and more accurate than length in a divided bar. S3
Same charts. Same eyes. Different question. Different ranking.
So the ranking is not a property of the chart. It is a property of the pairing of a chart with a decoding task — and the task is set by what the reader wants to know, which is a thing you have to go find out rather than assume.
Read for the mechanism, not just the result. Their argument is that people arrive at a chart with schemata — generic expectations about what kind of message a chart type carries — and that those expectations govern which elementary process gets run. A pie chart tells your visual system it is being asked about parts of a whole, so it recruits an anchoring-and-adjustment process suited to that, and does well. Ask it a comparison question instead and the machinery is mismatched.
Provenance note: 40 undergraduates on a CRT, one paper, 1987 Single study. Spence & Lewandowsky later found pie charts performing well for comparing combinations of proportions S9, so the direction has support, but treat the specific finding as one solid result rather than a settled law.
Read Simkin & Hastie.
JudgeRound 2 in the Lab — same stimuli, new question.25 m
Round 2 shows you pies and divided bars again. The question changes: instead of "what percent is the smaller of the larger," you'll be asked "what percent of the whole is the marked segment."
Before you start, the Lab will ask you to predict which encoding will come out ahead this time, and how confident you are. Do that honestly — the prediction is the instrument, not a formality.
Inference Twelve trials cannot replicate Simkin & Hastie, and yours may well come out the other way. What you're buying is not a result. It's the experience of the same chart behaving differently under a different question, which is much harder to un-know than a sentence about it.
Round 2 complete; prediction logged before running.
DeriveWhich task is this? Six real questions.45 m
For each question below, name the elementary perceptual task it demands, then name a chart that would serve it and one that would sabotage it. This is method-choosing, not method-executing — the harder and more useful skill.
"Is support ticket volume trending up or down this quarter?"
"Which of our seven regions is the biggest?"
"Roughly what share of total spend is infrastructure?"
"Did the gap between plan and actual widen or narrow after March?"
"Which two of these forty accounts are outliers on both axes?"
"Is the West region bigger than the East region?"
Confidence:
Direction / slope. A line chart serves it. A table sabotages it — Cleveland & McGill's Figure 8 makes exactly this point: strip the ability to perceive slope and the nonlinear pattern becomes very hard to see S1.
Position along a common scale, plus ordering. A sorted dot chart or bar chart serves it. A choropleth map sabotages it — that's shading, rank 6, and the state areas distort it besides S1.
Proportion of the whole. This is the Simkin & Hastie case: a pie is defensible here and a divided bar is worseS3. A sorted bar chart with no total shown answers a different question than the one asked.
Length — a vertical distance between two curves — which is the trap. Cleveland & McGill were emphatic: judging differences between curves is so bad they didn't bother running subjects, and the brain reads minimum distance between curves rather than vertical distance S1. Plot the difference directly as its own series.
Position along two common scales simultaneously, plus a detection task. A scatterplot serves it. Two separate bar charts destroy it.
Position along a common scale — but note this is a two-value comparison, and the honest answer may be that a sentence beats every chart. Ehrenberg's point, which Cleveland & McGill concede: if conveying numbers precisely were the only goal, tables would be better S1.
If you got 3 and 4 the same way round as the folklore would have them, that's the threshold biting.
Six questions answered before revealing.
BuildDecision Records #2, #3, #4 — task first.2 h
Three more charts you own. The order of the fields inverts: state the reader's judgment before you name the channel. If you can't state the judgment, you have found something more important than a chart problem — go ask the reader.
Records #2 and #3 give you the field prompts. Record #4 gives you a blank box. Write it in whatever shape the chart actually needs. If the scaffold was doing real work, you'll feel its absence; that's the point of removing it.
Records #2–#4 written.
WriteOne chart I would now defend that I would have replaced last month.15 m
In Records. One paragraph. The point is to catch the threshold in the act: find a case where the ranking alone would have told you to change something, and the task analysis tells you to leave it alone.
If you can't find one, say so — and say what that suggests about how you were choosing charts before.
Written.
Unit Four
Three rankings, one of them tested
~3 h 50 m · 4 blocks · threshold 3
Evidence this unit produces
A provenance table for your own encoding vocabulary: every channel you actually use, marked measured, extended by reasoning, or folklore.
The ranking you've been working with is the quantitative one. In 1986 Jock Mackinlay built APT, a program that designed charts automatically, and to do that he had to make the ranking machine-readable. Two things came out of it.
First, a pair of criteria that are still the cleanest formulation anyone has: a graphical language is expressive if it states all of the facts and only the facts, and effective if the human visual system can decode it readily. Expressiveness is a truth condition; effectiveness is a perception condition. Most bad charts fail the first and get argued about on the second. S4
Second, and this is the threshold: Mackinlay said outright that the Cleveland & McGill ranking does not address non-quantitative information, which involves different perceptual tasks and different task rankings — noting that color sits at the bottom of the quantitative ranking and is a very effective way to encode nominal sets. S4
So he produced three rankings: quantitative, ordinal, nominal. The one everybody quotes is the quantitative one. And here is the part that almost never travels with the table: Cleveland & McGill empirically verified the basic properties of the quantitative ranking. The ordinal and nominal columns are Mackinlay's extension — reasoned from Bertin and from psychophysics, not measured. S4Inference
ReadMackinlay 1986, and then Kosara's demolition.1 h 15 m
The three columns as they are conventionally transcribed:
#
Quantitative
Ordinal
Nominal
1
Position
Position
Position
2
Length
Density (value)
Color hue
3
Angle
Color saturation
Texture
4
Slope
Color hue
Connection
5
Area
Texture
Containment
6
Volume
Connection
Density (value)
7
Density (value)
Containment
Color saturation
8
Color saturation
Length
Shape
9
Color hue
Angle
Length
10
Texture
Slope
Angle
11
Connection
Area
Slope
12
Containment
Volume
Area
13
Shape
Shape
Volume
Look at where length sits. Second for quantitative data, eighth for ordinal, ninth for nominal. If you encode a nominal category with bar length you are using a channel that is nearly the worst available for that data type and silently asserting an ordering the data doesn't have — an expressiveness violation, not merely an ineffective one. S4
Kosara traces the belief that we read pie charts by central angle back to a 1926 study in which participants were asked which cue they used and just over half said central angle — and that, he argues, is how it became established as fact. S6
His own work with Skau then tested it: they built pie variants that isolated arc, angle and area, and varied donut inner radius from filled pie to thin outline. Both studies point to angle being the least important cue, and to donut charts being about as accurate as pies. S5Contested
Sit with what that does to the chain. Cleveland & McGill's position-angle experiment used a pie chart as the stimulus for the angle task S1. If readers aren't primarily using angle to read pies, then the experiment measured something real — pies are worse than bars for that comparison — but the label on the finding may be wrong. The result survives; the mechanism doesn't.
Read Mackinlay and Kosara.
DeriveExpressiveness before effectiveness — four datasets.45 m
Independent · no worked example this time
For each, type the data (quantitative / ordinal / nominal), name a channel you'd choose, and say which criterion rules out your second choice — expressiveness or effectiveness.
Twelve engineering teams, each with an owner's name.
Severity of incidents: sev-1 through sev-4.
Monthly recurring revenue by account, spanning three orders of magnitude.
Deploy status per service: passing, failing, not configured.
Confidence:
The specific channels are arguable; what isn't arguable is the shape of the reasoning. (1) is nominal — color hue is second in that column and bar length is ninth, and length additionally asserts an order that team names don't have, which is an expressiveness failure. (2) is ordinal — density or saturation rank high there and preserve order; hue ranks fourth for ordinal and, more importantly, doesn't carry order, so it fails expressiveness for a severity scale. (3) is quantitative — position, on a log scale, and the three-orders span is why area would be a disaster twice over: area is fifth, and Stevens' power law means area judgments carry a systematic bias with an exponent well below 1 S1S8. (4) is nominal with a conventional color mapping, and the trap is that "not configured" is a different kind of thing from "failing" — a missing state, not a value.
The discipline worth taking away: run expressiveness first. It's a yes/no question about truth and it eliminates most of the candidates before any perceptual argument starts. S4
Attempted before revealing.
BuildAudit your own encoding vocabulary.1 h 30 m
List every visual channel you have actually used in the last year — in dashboards, decks, docs, product UI. For each one, mark its provenance:
Measured — there is an experiment behind its position in the ranking, and you could name it.
Extended by reasoning — someone credible derived it from something measured. Mackinlay's ordinal and nominal columns live here. S4
Folklore — you believe it and cannot say where it came from.
Write it into Records as a table. Expect the folklore column to be the longest, and expect it to contain at least one rule you have enforced on other people.
Inference This is the block that carries the threshold. Not the rankings themselves — the habit of asking, of any design rule you're about to apply, whether anyone ever measured it. That habit transfers to every technical field you work in, which is what makes it worth four hours.
Provenance table written.
WriteOne rule I've been applying that was never tested.20 m
In Records. Name it, name where you think you picked it up, and say what you'll do about it — which may legitimately be "keep applying it, and stop citing it as though it were evidence."
Written.
Unit Five
Rebuild something you own, then test it
~5 h · 4 blocks · no scaffolding
Evidence this unit produces
A rebuilt chart, a cold measurement of yourself against the original, and a one-page defense memo that someone could disagree with on the merits.
No new reading. No prompts. Take the worst chart in your decision records and fix it — and then do the thing almost nobody does, which is check whether the fix worked.
ChoosePick the target and state the win condition first.20 m
Before touching the design, write down what would count as an improvement and what would count as a wash. If you can't specify that in advance, you will find an improvement no matter what you build.
Target and win condition written.
RebuildRedesign it, task-first.2 h
Constraints, all of which you now have reasons for: expressiveness before effectiveness; the reader's actual judgment before the channel; and a stated cost for every place you deliberately chose a lower-ranked channel, because there are legitimate reasons to and "it looked better" is occasionally one of them if you say so out loud.
Redesign built.
TestRound 3 — the cold re-run, at least five days later.40 m
The Lab has a third round: the same four encodings, fresh stimuli. It stays locked until five days after you finished Round 1.
The lock is deliberate. Running it now would measure how well you remember Round 1; running it cold measures whether anything durable changed. You can override the lock, and if you do, the Lab will label the result as contaminated rather than pretend otherwise.
Inference Interpret the comparison carefully. Improvement here is practice, not perception — your visual system did not get better in a week. What it plausibly measures is whether you've learned to stop trusting the encodings that deserve less trust.
Round 3 run cold.
DefendThe memo.1 h
One page. Three paragraphs and a table.
The reader and the judgment. Who looks at this, and what are they trying to decide.
The encoding and its cost. What you chose, what you rejected, and the estimated or measured error difference — with the source or the honest statement that you're extrapolating.
What would change your mind. The result that would make you rebuild it again.
Plus the provenance table for every claim in it: measured, extended, or folklore.
If the memo contains a sentence you couldn't defend in front of someone who has read Cleveland & McGill, cut it. That's the exit criterion for the whole course.
Memo written.
Track BOptional: re-encode this course's own progress meter.+2 h
The meter at the top of this page is an argument. It renders your progress in a deliberately terrible channel and re-encodes upward as you go, ending on a dot against a common scale — which is Cleveland & McGill's own §5.3 progression from shaded patch map to framed rectangle to dot chart, run on you instead of on a map of murder rates. S1
It is also, right now, unevaluated. Nobody tested it. It is exactly the kind of thing this course teaches you to distrust.
So: audit it. Is the sequence defensible? Does the framed-rectangle stage actually behave like position-along-nonaligned-scales, or is it just a bar with a box around it? Is the whole conceit an expressiveness violation — encoding an ordinal thing (progress through five units) in a quantitative channel? Then rebuild it and say why yours is better.
A course that teaches a discipline it does not practice is worth less than one that admits where it is exempting itself. This block exists to close that gap.
Meter audited and rebuilt.
The instrument
Proportional judgment
Three rounds. Round 1 replicates the comparison task across four encodings. Round 2 changes the question to proportion-of-the-whole. Round 3 is a cold re-run, locked for five days after Round 1.
Rules, from the original: make a quick visual judgment. Do not measure — not with a finger, not with a pencil, not by screenshotting and counting pixels. Do not revise an answer once submitted. Cleveland & McGill omitted tick marks and labels except at the extremes precisely so that subjects couldn't read values off an axis and divide S1; the stimuli here do the same.
Your work
Decision records & deliverables
Everything typed here is saved in this browser only. Use Export before you rely on it.
Sourcing
How this course was built
Every threshold rests on a primary source. Below, each source carries a read depth, because "cited" and "read" are different things and courses routinely blur them. Two of the six central sources come from authors at Tableau, which sells visualization software; that isn't disqualifying, it's disclosable.
Where a claim in this course rests on a source read only at abstract depth, the claim is marked in the text and stated narrowly.
ID
Source
Tier
Read depth
S1
Cleveland, W. S. & McGill, R. (1984). Graphical Perception: Theory, Experimentation, and Application to the Development of Graphical Methods. JASA 79(387), 531–554. PDF
1
Full text
S2
Heer, J. & Bostock, M. (2010). Crowdsourcing Graphical Perception. ACM CHI, 203–212. PDF
1
Full text
S3
Simkin, D. & Hastie, R. (1987). An Information-Processing Analysis of Graph Perception. JASA 82(398), 454–465. PDF
1
Abstract + method section
S4
Mackinlay, J. (1986). Automating the Design of Graphical Presentations of Relational Information. ACM TOG 5(2), 110–141. DOI
1
Abstract + quoted passages on the ranking extension
S5
Skau, D. & Kosara, R. (2016). Arcs, Angles, or Areas: Individual Data Encodings in Pie and Donut Charts. CGF 35(3), 121–130. Page
1
Abstract only
S6
Kosara, R. (2016). An Empire Built on Sand: Reexamining What We Think We Know About Visualization. BELIV. PDF
2
Partial — the pie-chart provenance argument
S7
Bertin, J. (1967). Sémiologie Graphique. Gauthier-Villars.
2
Not read. Cited here only as the acknowledged ancestor of S4, exactly as S1 and S2 cite it.
S8
Stevens, S. S. (1975). Psychophysics. Wiley. — the power law, p = kaα
2
Not read; reached through S1's discussion of exponents for length, area and volume
S9
Spence, I. & Lewandowsky, S. (1991). Displaying Proportions and Percentages. Applied Cognitive Psychology 5, 61–77.
1
Abstract only; used to corroborate direction, not magnitude
S10
Zeng & Battle (2023). A Review and Collation of Graphical Perception Knowledge for Visualization Recommendation. ACM CHI. DOI
3
Abstract; a route into the modern literature, not a load-bearing source
What was searched for and not found
A refutation search was run for each central claim. The material result: the ranking itself has replicated repeatedly, but its interpretation is contested in two specific places — whether angle is what readers use in a pie chart S5S6, and whether any single ranking survives contact with varying tasks and contexts S3S10. Both are taught in this course rather than smoothed over.
No serious challenge was found to the core position-beats-length-and-angle finding for comparison judgments. If you find one, that would be worth more than finishing the course.