Bayesian truth serum

Here’s a sneaky trick for extracting the truth from someone even when she’s trying to conceal it from you: Rather than asking her how she thinks or behaves, ask her how she thinks other people think or behave.

MIT professor of psychology and cognitive science Drazen Pelec calls this trick “Bayesian truth serum,” according to Tyler Cowen in Discover Your Inner Economist. The logic behind it is simple: our impressions of “typical” attitudes and behavior are colored by our own attitudes and behavior. And that’s reasonable. You should count yourself as one data point in your sample of “how people think and behave.”

Your own data is likely to influence your sample more strongly than other data points, however, for two reasons. First, because it’s much more salient to you, compared to your data about other people, so you’re more likely to overweight it in your estimation. And second, through a ripple effect — people tend to cluster with other people who think and act similarly to themselves, so however your sample differs from the general population, that’s an indicator of how you yourself differ from the general population.

So, to use Cowen’s example, if you ask a man how many sexual partners he’s had, he might have a strong incentive to lie, either downplaying or exaggerating his history depending on who you are and what he wants you to think of him. But his estimate of a “typical” number will still be influenced by his own, and by that of his friends and acquaintances (who, because of the selection effect, are probably more similar to him than the general population is). “When we talk about other people,” Cowen writes, “we are often talking about ourselves, whether we know it or not.”

Game Theory and Football: How Irrationality Affects Play Calling

Coaches and coordinators in professional football get paid a lot of money to call the right plays – not just the best plays for particular situations, but also unpredictable plays that will catch the other team off guard. It’s a perfect setup for game theory analysis!

As in other game theory situations, the best play depends in part on what your opponent does. Your running play is much more likely to succeed against a pass-prevent defense, but would be in trouble against a run-stuffing formation. If the defense can guess what you’re going to call, they can adjust accordingly and have an advantage. Even on 3rd down and long – a common passing situation – there’s value in calling a percent of running plays, because the defense is less likely to be geared toward stopping that. But as you do it more, the chance of catching the defense off guard gets smaller. There’s some optimal balance where the expected success of a surprising run is equal to the expected success of a more sensible (but anticipated) pass.

The goal is to stay unpredictable and exploit patterns where your opponent is using a sub-optimal combination. If a team notices that passing plays are working better, they’ll be more likely to call them. As the defense notices, they’ll shift away from their run-defense and focus more on defending passes. In theory, the two teams reach an equilibrium.

In practice, it doesn’t quite work that perfectly – human beings are making the decisions, and humans are both vulnerable to cognitive biases and notoriously bad at mimicking true unpredictability. Brian Burke, a fellow fan of combining sports with statistics, was poring over the play-calling data for second downs and noticed something odd:

There’s a strange spike in percent of running plays called at 2nd and 10! Tactically, 2nd and 10 isn’t all that different from 2nd and 9 or 11, so it’s strange to see such a difference. Why would they call so many more running plays in that particular situation?

The key is to realize that there are two ways a team tends to find itself facing a 2nd and 10 situation – runs that happen to go nowhere or any incomplete pass. Of those, incomplete passes are far more common. So in cases of 2nd and 10, it’s most often because the team just failed a passing play. That suggests two reasons coaches might be irrationally switching to running plays, even at the cost of sacrificing unpredictability:

(1) The hasty generalization bias (also called the small sample bias) and the recency effect are cognitive biases in which people overgeneralize from a small amount of data, especially recent data. Failed passes are very common (about 40% fail), so there’s no good reason for a coach to treat any single failed pass as evidence that they’d be better off switching to a running play. But the urge to overreact to the failed pass that just happened is strong, thanks to these two biases.

(2) People are terrible at generating unpredictability — when asked to make up a “seemingly-random” sequence of coin flips, we tend to use far more alternation between Heads and Tails than would actually occur in a real sequence of coin flips. So even if coaches weren’t overreacting to a failed pass, and they were simply trying to be unpredictable, they would still tend to switch to a running play after a passing play more often than random chance would dictate.

Indeed, when Brian separated the data by previous play, the alternation trend is clear — passes are more likely after runs, and runs are more likely after passes:

(My favorite team, the Baltimore Ravens, was pretty bad about this under the previous regime, Coach Billick)

Brian concludes:

Coaches and coordinators are apparently not immune to the small sample fallacy. In addition to the inability to simulate true randomness, I think this helps explain the tendency to alternate. I also think this why the tendency is so easy to spot on the 2nd and 10 situation. It’s the situation that nearly always follows a failure. The impulse to try the alternative, even knowing that a single recent bad outcome is not necessarily representative of overall performance, is very strong.

So recency bias may be playing a role. More recent outcomes loom disproportionately large in our minds than past outcomes. When coaches are weighing how successful various play types have been, they might be subconsciously over-weighting the most recent information—the last play. But regardless of the reasons, coaches are predictable, at least to some degree.

Coaches are letting irrational biases influence their play calling, pulling them away from the optimal mix. The result, according to Pro Football Reference stats, is less success on those plays. I wonder how well a computer could call plays using a Statistical Prediction Rule

Game theory and basketball

Ben Morris is a friend-of-a-friend of mine who recently competed in a contest sponsored by ESPN called “Stat Geek Smackdown,” in which the goal was to correctly predict as many of the NBA playoff games as possible. For each correct guess, a contestant received 5 points.

Heading into the final game between Miami and Dallas, Ben was in second place, trailing just 4 points behind a veteran stat geek named Ilardi. By most estimates, Miami had about a 63% chance of beating Dallas. But Ben realized that if he and Ilardi both chose Miami, then even if Miami won the game, Ilardi would still win the competition, because he and Ben would each get 5 points and the gap between their scores would remain unchanged. In order for Ben to win the competition, he would have to pick the winning team and Ilardi would have to pick the losing team.

So that created an interesting game theory problem: If Ben predicted that Ilardi would pick Miami, since they were more likely to win, then Ben should pick Dallas. But if Ilardi predicted that Ben would be reasoning that way, then Ilardi might pick Dallas, knowing that all he needs to do to win the competition is to pick the same team as Ben. But of course if Ben predicts that Ilardi will be thinking that way, maybe Ben should pick Miami…

What would you do if you were Ben? You can read about Ben’s reasoning on his excellent blog, Skeptical Sports, but here’s my summary. Ben essentially had two options:

(1) His first option was to play his Nash equilibrium strategy, which is a concept you might recall if you ever took game theory (or if you saw the movie “A Beautiful Mind,” although the movie botched the explanation). That’s the set of strategies (Ben’s and Ilardi’s) which gives each of them no incentive to switch to a new strategy as long as the other guy doesn’t. The Nash equilibrium strategy is especially appealing if you’re risk averse because it’s “unexploitable,” meaning that it gives you predictable, fixed odds of winning the game, no matter what strategy your opponent uses.

In this case — and you can read Ben’s blog for the proof — the Nash equilibrium is for Ben to pick Miami with exactly the same probability as Miami has of losing (0.37) and for Ilardi to pick Miami with exactly the same probability as Miami has of winning (0.63). (You might wonder how you should pick a team “with X probability,” but it’s pretty easy: just roll a 100-sided die, and pick the team if the die comes up X or lower.)

If you do the calculation, you’ll find that playing this strategy — i.e., rolling a hundred-sided die and picking Miami only if the die came up 37 or lower — would give Ben a 23.3% chance of beating Ilardi, no matter how Ilardi decided to play. Not terrible odds, especially given that this approach doesn’t require Ben to make any predictions about Ilardi’s strategy. But perhaps Ben could do better if he were able to make a reasonable guess about what Ilardi would do.

(2) That leads us to option two: Ben could abandon his Nash equilibrium strategy, if he felt that he could predict Ilardi’s action with sufficient confidence. To be precise, if Ben thinks that Ilardi is more than 63% likely to pick Miami, then Ben should pick Dallas.

Here’s a rough proof. Call “p” the likelihood that Ilardi picks Miami, and “q” the likelihood that Ben picks Miami. Then we can assign probabilities to each of the outcomes in which Ben wins:

Since the two outcomes are mutually exclusive, we can add up their probabilities to get the total probability that Ben wins, as a function of p and q:

Probability Ben wins = .37p + .63q – pq

Just to illustrate how Ben’s chance of winning changes depending on p, I plugged in three different values of p to create three different lines: For the black line, p=0.63. For the red line, p < 0.63 (to be precise, I plugged in p=0.62, but any value of p<0.63 will create an upward sloping line). For the blue line, p > 0.63 (to be precise, I plugged in p=0.64, but any value of p>0.63 will create a downward sloping line).

If p = .63, that renders Ben’s chance of winning constant ( .233) for all values of q. In other words, if Ilardi seems to be about 63% likely to pick Miami, then it doesn’t matter how Ben picks, he’ll have the same chance of winning (23.3%) as he would if he played his Nash equilibrium strategy.

If p > .63, Ben’s chance of winning decreases as q (his probability of choosing Miami) increases. In other words, if Ben thinks there’s a greater than 63% chance that Ilardi will pick Miami, then Ben should pick Miami with as low a probability as possible (i.e., he should pick Dallas).

If p < .63, Ben’s chance of winning increases as q (his probability of choosing Miami) increases. In other words, if Ben thinks there’s a lower than 63% chance that Ilardi will pick Miami, then Ben should pick Miami with as high a probability as possible (i.e., he should pick Miami).

So what happened? Ben estimated that Ilardi would pick Miami with greater than 63% probability. That’s mainly because most people aren’t comfortable playing probabilistic strategies that require them to roll a die —  people will simply “round up” in their mind and pick the team that would give them a win more often than not. And Ben knew that if he was right about Ilardi picking Miami, then Ben would end up with a 37% chance of winning, rather than the 23.3% chance he would have had if he stuck to his equilibrium strategy.

So Ben picked Dallas. As he’d predicted, Ilardi picked Miami, and lucky for Ben, Dallas won. This one case study doesn’t prove that Ilardi reasoned as Ben expected, of course. Ben summed up the takeaway on his blog:

Of course, we shouldn’t read too much into this: it’s only a single result, and doesn’t prove that either one of us had an advantage.  On the other hand, I did make that pick in part because I felt that Ilardi was unlikely to “outlevel” me.  To be clear, this was not based on any specific assessment about Ilardi personally, but based my general beliefs about people’s tendencies in that kind of situation.

Was I right? The outcome and reasoning given in the final “picking game” has given me no reason to believe otherwise, though I think that the reciprocal lack of information this time around was a major part of that advantage.  If Ilardi and I find ourselves in a similar spot in the future (perhaps in next year’s Smackdown), I’d guess the considerations on both sides would be quite different.

I feel ya, Gureckis

I’m feeling a deep sense of camraderie right now with Todd Gureckis, a psychologist at NYU. That’s because a couple of weeks ago, senator Tom Coburn (R-OK) released a report titled, “Under the Microscope,” scrutinizing the funding decisions of the National Science Foundation and complaining about what he felt was a waste of taxpayer money on many frivolous research projects — one of which was Gureckis’. “Armed with a $1 million grant from NSF,” Coburn wrote, “researchers at Indian (sic) University-Bloomington and New York University analyzed baby names to determine trends in parents’ naming decisions.”

The paper in question, co-authored with Rob Goldstone, is called, “How You Named Your Child: Understanding The Relationship Between Individual Decision Making and Collective Outcomes.” Gureckis was surprised at Coburn’s criticism, and responded on his website:

“The Coburn report makes it seem as though this research spent money to determine the frequency and popularity of names… Had those developing this report actually looked the research paper they were criticizing, they would know that we were not specifically interested in baby names except in so far as they offer a unique opportunity for studying such the impact of social influence on decision making. We all know that iPhones are popular but the underlying reasons for this cultural success is distorted by the role that advertising budgets and existing computer technologies play in determining which ideas win out and which die off in the consumer marketplace. In contrast, the popularity of names is more organically determined by processes of social influence (there is no company out there trying to convince you to name you child something in particular). Baby names thus represent an important and relatively “pure” empirical test of theories of cultural transmission and social influence in large groups.”

Now of course, I’m not an NSF-funded researcher being criticized for frivolity. But the reason I felt so much camraderie with Gureckis after reading about his situation was because this sort of thing happens to me all the time — I’ll bring up a particular case as a way of shedding light on a general principle, and the people I’m talking to focus on the particular case and ignore the general principle.

For example, I’ve tried a couple of times to start a discussion about the difficulties of measuring happiness, and I’ve begun by citing the fact that most parents claim to be very happy that they have children despite the fact that research shows parents are less happy, on average, than non-parents. So that points to this really interesting tension between two ways of measuring happiness (how satisfied are you when you consider your life overall, versus how happy do you feel on a moment-to-moment basis) that apparently can contradict each other, and raises the question of whether one is “wrong,” and if so, which?

At least, that’s the discussion I keep wanting to have. But I never get to, because the thread always turns into a debate about having children, with commenters testifying about how happy they are that they had kids and how the parenting-skeptics are missing out.

Asking for reassurance: a Bayesian interpretation

Bayesianism gives us a prescription for how we should update our beliefs about the world as we encounter new evidence. Roughly speaking, when you encounter new evidence (E), you should increase your confidence in a hypothesis H only if that evidence would’ve been more likely to occur in a world where H was true than in a world in which H was false — that is, if P(E|H) > P(E|not-H).

I think this is indisputably correct. What I’ve been less sure about is whether Bayesianism tends to lead to conclusions that we wouldn’t have arrived at anyway just through common sense. I mean, isn’t this how we react to evidence intuitively? Does knowing about Bayes’ rule actually improve our reasoning in everyday life?

As of yesterday, I can say: yes, it does.

I was complaining to a friend about people who ask questions like, “Do you think I’m pretty?” or “Do you really like me?” My argument was that I understood the impulse to seek reassurance if you’re feeling insecure, but I didn’t think it was useful to actually ask such a question, since the person’s just going to tell you “yes” no matter what, and you’re not going to get any new information from it. (And you’re going to make yourself look bad by asking.)

My friend made the valid point that even if everyone always responds “Yes,” some people are better at lying than others, so if the person’s reply sounds unconvincing, that’s a telltale sign that that they don’t genuinely like you/ think you’re pretty. “Okay, that’s true,” I replied. “But if they reply ‘yes’ and it sounds convincing, then you haven’t learned any new information, because you have no way of knowing whether he’s telling the truth or whether he’s just a good liar.”

But then I thought about Bayes’ rule and realized I was wrong — even a convincing-sounding “yes” gives you some new information. In this case, H = “He thinks I’m pretty” and E = “He gave a convincing-sounding ‘yes’ to my question.” And I think it’s safe to assume that it’s easier to sound convincing if you believe what you’re saying than if you don’t, which means that P(E | H) > P(E | not-H). So a proper Bayesian reasoner encountering E should increase her credence in H.

(Of course, there’s always the risk, as with Heisenberg’s Uncertainty Principle, that the process of measuring something will actually change it. So if you ask “Do you like me?” enough, the true answer might shift from “yes” to “no”…)

Why Imagined Indulgence Helps Us Diet

What makes decadent waffles so damn satisfying in the morning? Is it the optimal balance of crispy and soft textures? The fat in the whipped cream? The sugar content? It turns out that there’s a factor beyond the actual food: your frame of mind. A team of researchers at Yale just performed a clever study and found that you feel fuller and more sated if you believe you just ate something indulgent.

As with most psychology experiments, the study involved lying to people. Subjects were given a milkshake on two separate occasions but were told that one contained a whopping 620 calories and the other had a more sensible 140 calories. In reality, both shakes were the same – right in the middle at 380 calories.

Before and after each test, the researchers monitored the subjects’ ghrelin levels as a measure of how satisfied they were. Ghrelin – the hormone which triggers hunger – increases and spikes before meals, then drops off after people eat. If the calorie content were all that mattered, there would be no difference in reactions to the two shakes. But there was:

Results: The mindset of indulgence produced a dramatically steeper decline in ghrelin after consuming the shake, whereas the mindset of sensibility produced a relatively flat ghrelin response. Participants’ satiety was consistent with what they believed they were consuming rather than the actual nutritional value of what they consumed.

What should we make of the finding (besides a continued fascination with the placebo effect)? For one thing, it reinforces the notion that our stomachs are very crude sense organs which aren’t precise or accurate at judging how much food they need.

For anyone trying to achieve (or maintain) a healthy weight, the dynamic makes it tougher to diet. The conscious decision to eat ‘sensible’ food motivates our bodies to demand more calories. What a frustrating situation!

Patrick at Discoblog toys with a creative solution:

It definitely suggests some new approaches to dieting, like berating yourself for eating celery sticks in an effort to make them seem more luxurious and satisfying. But it’s not clear if lying to yourself is as effective as having other people lie to you. And believing that you are constantly eating poorly might have other psychological side effects, one supposes.

I agree, it probably doesn’t work as well to lie to yourself (and nobody will be able to convince me that celery sticks are fatty treats). But we can draw a useful tactic that doesn’t require deception. Instead of applying the study’s findings when we eat light food, keep it in mind when eating dessert. Next time you want a rich slice of cheesecake, look up how many calories it has! According to the study, focusing on the fact that the slice has 50% of your recommended calories will make you feel more satisfied eating less of it.

What I’d like to see is a study that’s honest about the number of calories but emphasizes different ingredients to foster that ‘indulgent’ mindset. Would our bodies react differently to drinking a “300-calorie fruit milkshake” compared to the same one described as a “300-calorie shake with bananas, heavy cream, vanilla extract, and pure cane juice”?

If that works, we can help our friends and families by focusing attention on the fattiest, sweetest, and tastiest part of a dish. Next time Julia is willing to make those delicious-looking blintzes again, sign me up. I can eat one as she tells me about the heavy cream that went into the homemade ricotta.

[UPDATE] I’m looking a little deeper into what exactly was being measured – the abstract and researchers said “Participants’ satiety was consistent with what they believed they were consuming rather than the actual nutritional value of what they consumed.” but news sources report this as well:

The study also didn’t find that the larger drop in ghrelin in those who drank the indulgent shakes was accompanied by a larger drop in hunger levels, a finding that the researchers couldn’t fully explain. “We may not have used a reliable measure of hunger,” says Crum. “My sense is that hunger levels should have changed.”

I had assumed that the participants’ satiety was the same as their remaining hunger – but those two quotes seem at odds at first glance.

A crab canon for Douglas Hofstadter

(published at 3 Quarks Daily)

Since it first came out in 1979, Douglas Hofstadter’s Pulitzer Prize-winning book “Gödel, Escher, Bach: An Eternal Golden Braid” has widened the eyes of multiple generations of nerdy kids, and I was certainly no exception. The book draws all sorts of parallels between music, art, math, and computer science, ultimately shaping them into a bold thesis about how consciousness arises from self-reference and recursion. It’s also a very playful book, full of puzzles, puns, and imagined dialogues between Achilles and a tortoise which weave in and out of the main chapters, illustrating the concepts therein.

One of those dialogues, titled “Crab Canon,” seems puzzling when you begin reading it – sprinkled with seeming non-sequitors, the word choice a bit awkward and off-kilter. Then shortly after the halfway point, when you start to see recent lines repeated, in reverse order, you realize: the whole dialogue is a line-level palindrome. The first line is the same as the last, the second line is the same as the second-to-last, and so on. But because Hofstadter chooses his sentences carefully, they often have different meanings when they reoccur in the reverse order. So, for example, the following bit of dialogue in the first half…

Tortoise: Tell me, what’s it like to be your age? Is it true that one has no worries at all?
Achilles: To be precise, one has no frets.
Tortoise: Oh, well, it’s all the same to me.
Achilles: Fiddle. It makes a big difference, you know.
Tortoise: Say, don’t you play the guitar?

… becomes this bit of dialogue in the second half:

Achilles: Say, don’t you play the guitar?
Tortoise: Fiddle. It makes a big difference, you know.
Achilles: Oh, well, it’s all the same to me.
Tortoise: To be precise, one has no frets.
Achilles: Tell me, what’s it like to be your age? Is it true that one has no worries at all?

Hofstadter does “cheat” a bit, by allowing himself to vary punctuation (for example, “He often plays, the fool” reoccurs later in a new context as “He often plays the fool”). Nevertheless, it’s an impressive execution of a clever conceit.

Thumbing through Gödel, Escher, Bach again recently, I came across the Crab Canon and was struck with the desire to attempt a similar feat myself: a line-level palindrome that tells a linear story, that is, a story in which a series of non-repeating events occur.

It’s maddeningly difficult trying to come up with lines that make sense, that in fact make a different sense, in both directions. And because all of the lines interlock — each relying simultaneously on the line preceding it, the line following it, and its mirror-image line — changing any one line in the poem tends to set off a ripple effect of necessary changes to all the other lines as well.

I eventually figured out a few crucial tricks, like relying on ambiguous pronouns (“they,” “their”) and using images that carry a different meaning depending on what’s already happened. Below is the final result – my own canon, in homage to the book that dazzled my teenaged self years ago:

SEASIDE CANON, for Douglas Hofstadter

by Julia Galef

~
The ocean was still.
In an empty sky, two gulls turned lazy arcs, and
their keening cries echoed
off the cliff and disappeared into the sea.
When the child, scrambling up the rocks, slipped
out of her parents’ reach,
they called to her. She was already
so high, but those distant peaks above —
they called to her.  She was already
out of her parents’ reach
when the child, scrambling up the rocks, slipped
off the cliff and disappeared into the sea.
Their keening cries echoed
in an empty sky. Two gulls turned lazy arcs, and
the ocean was still.

~

RS#36: Why should we care about teaching the humanities?

Episode #36 of the Rationally Speaking podcast is out, and this one’s a lively debate between me and Massimo about the value of humanities departments in universities. While I don’t deny the huge amount of enjoyment we get from arts and literature, I express skepticism about many of the typical justifications for requiring humanities courses. Those justifications strike me as either (1) overly vague and subjective (“the humanities make you a complete person”) or (2) making contrived claims about the practical benefits of studying the arts (“the humanities build critical thinking skills”) to which I usually want to reply, “If that’s your goal, there are much more direct ways to pursue it than studying literature.”

Rationally Speaking #36: Why should we care about teaching the humanities?

“Stand back everyone, I’ve been trained for situations like this…”

Last night my friends and I ended up talking about real-life situations in which our math skills serendipitously came in handy. And I got to reminisce about my one exciting “Thank goodness I paid attention in math class!” moment:

I was a high school kid, working as a summer intern at the Corcoran Gallery of Art (this was back in my “I want to be a museum curator” phase). In the sales office one Friday afternoon, I overheard a conversation between my two bosses:

Boss 1: “The new ticket collector didn’t keep track of child and adult ticket sales separately this week. All she sent us is the total number of tickets and total revenue. 953 tickets, $9,050 revenue. But the accountant wants us to record child and adult sales separately.”

Boss 2: “Sigh. All right, I guess we’ll have to go get the pile of ticket stubs and sort them all. What a pain in the ass…”

Me (gasps, runs over): “Wait! We don’t need to sort ticket stubs! We already have all the information we need to solve this!”

Boss 1: “We do? How?”

Me: “We need to set up a system of equations! Okay, let’s call A the number of adult tickets and C the number of child tickets. How much does each one cost?”

Boss 1: “Adult tickets are $10, child tickets are $6.”

Me: “Okay! So we know that the total number of tickets is 953 so we can write A + C = 953. And we know the total revenue is $9050 so we can write $10A + $6C = $9050. So we have two equations,  two variables:

A + C = 953
10A + 6C = 9050

And now we just solve:
10 (953 – C) + 6C = 9050
4C = 480
C = 120. Therefore A = 953 – 120 = 833.
So that’s the answer — we sold 120 child tickets, and 833 adult tickets.”

My bosses were delighted with the “cool trick” I had used. And I like to think my 7th-grade math teacher would have been tickled pink if she’d seen that go down. How often do you get to actually apply your word-problem skills in real life?

Nothing quite that math-textbook perfect has happened since. Though I keep hoping that someday I’ll overhear someone saying, “My friend and I wanted to meet up for lunch tomorrow, so we agreed to each leave our apartments at noon and walk towards each other’s place until we met. He walks at a rate of 4 mph and I walk at a rate of 3 mph. If only there was some way to figure out where and when we would meet so that I could make a lunch reservation…”

The Game Theory of Story Endings

Do happy endings really make you as happy if you see them coming a mile away? When we watch a trashy action flick or a fluffy romantic comedy, aren’t the conflicts less interesting because we know it’ll all end happily ever after? Someone has to bite the bullet and write a sad ending to give plausibility to the threat of unhappiness. It’s disincentivized because sad endings are more challenging and risk upsetting the audience, but someone has to do it.

Steven E. Landsburg muses about this in The Armchair Economist:

I am intrigued by the market for movie endings. Movie-goers want two things in an ending: They want it to be happy and they want it to be unpredictable. There is some optimal frequency of sad endings that maintains the right level of suspense. Yet the market might fail to provide enough sad endings.

An individual director who films a sad ending risks short-term losses, as word gets around that the movie is “unsatisfying.” It is true that there are long-term gains, as viewers are kept off their guard for future movies. Unfortunately, most of those gains may be captured by other directors, because movie-goers remember only that the murderer does sometimes catch up with the heroine in the basement, and do not remember that it happens only in movies with particular directors. Under these circumstances, no individual director may be willing to incur costs for his rivals’ benefit.

A solution is for directors to display their names prominently, so that viewers know when a movie was made by someone unpredictable. Viewers, however, may find it in their interests to retaliate by covering their eyes when the director’s name is shown.

If you can be associated more strongly with unpredictability, you reap more benefits. You’re also more strongly associated with the unhappy ending, which might turn audiences away.

One way to ease the blow of an unexpected sad ending is to make deaths triumphant, defiant, or heroic. Think of how Spock died in The Wrath of Khan (No, I’m not going to give a spoiler alert for a 30 year old movie). Sure, people die in Star Trek all the time – when Kirk, Spock, and fresh-faced, red-shirted Ensign Jimmy beam down to explore a planet for life, we all know one of them isn’t going to make it back. But to kill a main character is more significant. And it was done in a touching way. They got the unpredictability without upsetting their audience.

I genuinely respect Joss Whedon for his willingness to throw curve balls like this in his story lines. He’s developed a reputation for having sympathetic characters die, leave, or change sides – often without warning. Rather than watching Buffy, Firefly and Serenity thinking “So, how is it all going to work out this time?” we’re forced to think “Is it going to work out this time?”

TV Tropes has a name for all this – Anyone Can Die:

This is where no one is exempt from being killed, including the main characters (maybe even the hero). The Sacrificial Lamb is often used to establish the writer’s Anyone Can Die cred early on. However, if the Lamb’s death is a one-off with no follow-up, it’s just Killed Off for Real. To really be Anyone Can Die, the work must include multiple deaths, happening at different points in the story. Bonus points if the death is unnecessary and devoid of Heroic Sacrifice.

In game theory situations, reputation plays a large role. TV Tropes mentions building a ‘Anyone Can Die’ cred, which can be achieved through repeated interactions. In a TV series or multiple films by the same director, you get a feel for whether the good guys always prevail. But even within a single story, early and repeated signaling can make the remainder of the plot more intense. When a major character is killed off without it being a Heroic Sacrifice, that’s a powerful signal that anything can happen. The musical Into the Woods will always have a special place in my heart for mastering this dynamic.

But there’s another route. Historical dramas can increase society’s perception of “sadness plausibility” without anyone taking a hit for being a downer. Nobody’s going to feel unsatisfied that Titanic, The Great Escape, or Butch Cassidy and the Sundance Kid have sad endings. (Or if they do, they can take it up with reality for writing a depressing script. It’s not easy to keep those separate in our brains; we just get the overall sense that sometimes stories have sad endings. And that perception helps us enjoy all the other movies we watch.