Showing posts with label Behaviour. Show all posts
Showing posts with label Behaviour. Show all posts

Wednesday, June 2, 2021

New Published Articles

Dear readers,

I have taken a short break (as I am wont to do from time-to-time) to focus on some research. I am happy to now announce the fruit of this labour:

First, "Doping in Sports: A Compliance Conundrum" has been included as a chapter in the textbook The Cambridge Handbook of Compliance. It is chapter 65 of 69, which leads me to believe they followed the age-old adage, 'save your best for [almost] last.' (I will update the link with a version of the text if/when I can find one).

Second, "Stadium Giveaway Promotions: How Many Items to Give and the Impact on Ticket Sales in Live Sports" has been published (online) at the Journal of Sport Management. Avid readers will find it takes a familiar spin on the bobblehead topic I have been taking about ad nauseum

Thank you for all your support and stayed tuned for more research updates and (most importantly) future posts on Sports and Economics (in that order).

Jeffrey

Thursday, October 18, 2018

Final Predictions for the 2018 Cy Young Awards

As tradition would have it, I now offer you my predictions for this year's Cy Young award winners. Recall that, I have found some evidence that the Baseball Writers' Association of America (BBWA) vote not for the pitcher with the best overall season, but for the pitcher who may have finished the strongest. I suggest that this is caused by the recency effect, a behavioural bias wherein humans tend to remember events that happened more recently than events that happened further in the past. Below, I have summarised my results and offered some interesting insights into the trend as to who was MLB's best pitcher throughout the season.

Cy Young Predictions 
AL Cy Young winner:
  • Blake Snell - Tampa Bay Rays (64.3% chance of winning)
  • Chris Sale - Boston Red Sox (33.4)
  • Justin Verlander - Houston Astros (0.2%)
NL Cy Young Winner:
  • Jacob deGrom - New York Mets (~100% chance of winning)
  • Max Scherzer - Washington Nationals (<0.1%)


Below are graphical depictions of how the Cy Young race may have ended at different points in time. As the recency effect would suggest, the pitcher who had an exceptional performance closest to the voting data has a better chance of winning the award. 

In the American League, we see Chris Sale dominated the AL Cy Young voting race from late June right up to September. But as seen in years past, Chris Sale's chances diminished in the final month of the season and Blake Snell took the top position for the AL Cy Young award.  Note that Justin Verlander is not quite out of it... maybe we will have more great 'commentary' from his adoring wife Kate Upton? (here's the Safe-For-Work tweet when Justin Verlander's did not receive the 2016 AL Cy Young, more 'colourful' language used here).



In the National League, there is a zero percent chance someone named anything other than Jacob deGrom or Max Scherzer wins the NL Cy Young. In fact, my model predicts deGrom winning is all but a sure thing. While the race was close for the second half of the season, Max Scherzer will likely not win his third consecutive award. And he may have only himself to blame: if only Max Scherzer had decided to pitch in the final game of the season...

Thanks for reading!

Friday, June 1, 2018

Time out! Do some scores look worse than others?

TLDR: The outcome of (most) sports is determined not by the actual value of an individual's or team's performance, but by the relative performance.

Despite the fact that winners and losers are determined by the point differential, every sport summarises the score as the cumulative total points for each opponent. Athletes, coaches, fans, and pundits alike are then forced to calculate the point differential themselves in order to determine how close the game really is.

In the NBA, each team may call a time-out in attempt to alter the momentum of the game, particularly when the momentum is not in their favour or when their team is not in the lead. In the small window of time available to call for a time-out, if a team miscalculated how many points their team was trailing by, would they be more likely to call a time-out?

I find that teams are more likely to call a time out when they are down by a score spanning two tens places over a score spanning one tens place, even when the point differential is the exact same! For example, if a team is down 81-69, they are 10% more likely to call a time out than if they were down 83-71!

The outcome of (most) sports is determined not by the actual value of an individual's or team's performance, but by the relative performance. It matters not how many points (or goals, etc.) you have, but how many points you have in comparison to your opponent.

Despite the fact that winners and losers are determined by the point differential, every non-racing sport summarises the score as the cumulative total points for each opponent (please comment below if you have a counter-example). Athletes, coaches, fans, and pundits alike are then forced to calculate the point differential themselves in order to determine how close the game really is. Due to cognitive biases in the way we perceive numbers, we may then expect to see some interesting patterns in data pertaining to these cumulative vs relative scores.

The National Basketball Association (NBA) is a league where the average combined score of a single game is over 200 points and the average margin of victory is just 10 points. Each game can see multiple lead changes and notable swings in momentum between the two teams. During the course of the game, each team may call a time-out in attempt to alter the momentum of the game, particularly when the momentum is not in their favour or when their team is not in the lead.

But do certain game conditions entice teams to systematically choose when to call a time-out? Does trailing your opponent by certain values induce more time-outs than others? And what sort of mental arithmetic biases might we see when dealing with basketball scores?

To answer all of the above, I first began by collecting all of the play-by-play data for the 2016/2017 NBA regular season, conveniently provided by Reddit. I organise the data to the point where each observation is a single possession (generally, a possession begins every time the shot-clock is reset to 24 seconds). I collected data on the team, the opponent, the amount of time remaining in regulation, the quarter, and the current score. It is from the latter that I calculate the point differential.

Using a computer to calculate point differential, I am all but guaranteed to get the correct answer. However, using mental math, I am likely to make some mistakes. In fact, I first noticed that when the two scores span a difference of two tens places, e.g. 81-69, I was more likely to overestimate the difference: I erred by incorrectly thinking the losing team was further behind than 12 points. Yet, I often was able to accurately calculate the point differential if the scores span only one tens place, e.g. 83-71. In the small window of time available to call for a time-out, if one miscalculated how many points their team was trailing by, would they be more likely to call a time-out?

To test if others shared in my blunders, I ran a logit model to predict how many times a time-out is called based on the actual point differential and the tens-place point differential. Here I considered only the instances where the team calling the time out was a) trailing their opponent in score, and; b) the actual point differential was between 11 and 19 points. The latter limits the data to possessions where the team making the time-out decision is down by only scores spanning one or two tens places. I also control for the time remaining in the quarter and game, the team making the time-out call, and the score change since the last time-out/end of last quarter (also known as the 'run').

I find that teams are more likely to call a time out when they are down by a score spanning two tens places over a score spanning one tens place, even when the point differential is the exact same! Below is the graphical depiction this increase in probability of calling a time-out from the score spanning two tens places:

When down between 11 and 19 points teams are 10% more likely to call a time-out if the point differential spans two tens places than they are to call a time-out if the point differential only spans one tens place. Using the same example from above, if a team is down 81-69, they are 10% more likely (taking the overall number) to call a time out than if they were down 83-71!

I have shown there exists some evidence that certain scores induce more time-outs than others, namely when the point differential spans two tens places. I suggest that may be due to our tendency to systematically miscalculate the difference between two numbers which overstates the difference in certain circumstances. We may also see a similar effect when people are trying to calculate the time between two events, such as a layover at an airport: 7:35 to 10:10 is the same span of time as 5:00 to 7:35 despite that the former may look like a larger interval.

In the end, the same point differential and just a subtle change in the actual the scores can make all the difference in the world in how we perceive the margin of victory.

Thursday, November 16, 2017

Cy Young Award Winners - Recency Bias Matters!

For those who have not been following, I first said in September, then reiterated in October, that recency bias affects the results for the winner of Major League Baseball's (MLB's) best pitchers - also known as the Cy Young Awards. This is because the Baseball Writers Association of America (BBWAA) have a tendency to favour more recent pitching performances and you can read about why here.

With the commencement of the 2017 MLB season, on 15 November, the results of the Cy Young voting were publicised and my suspicions were confirmed:


AL Cy Young Winner

Player
My Predicted
Probability
Actual Cy Young
Vote Points
Corey Kluber
Cleveland Indians      
99.7% 204
Chris Sale
Boston Red Sox
0.3% 126


NL Cy Young Winner

Player
My Predicted
Probability
Actual Cy Young
Vote Points
Max Scherzer
Washington Nationals
89.3% 201
Clayton Kershaw
Los Angeles Dodgers
0.7% 126
Stephen Strasburg
Washington Nationals
6.4% 81

Wednesday, November 1, 2017

Anchoring in College Football Rankings: How some schools get Au-burned

TLDR:  "Past research has shown that colleges with a successful football team experience an influx of donations, student applications, enrollment, and student quality." 

However, subjective rankings of football teams can be biased through anchoring. "Evidence for anchoring can potentially be illustrated through the initial assignment a preseason ranking: these preseason rankings can have a major impact on a team's final ranking."

"Initial results suggest that a team's rank before the game explains an incredible 83% of team's ranks after the game."

As a new-comer to the USA, I have quickly found that college is synonymous with football. Every Saturday in Autumn, tradition would have it that groups of former students gather round in their college sweaters to watch their alma mater kick off against a rival school. Among the friendly banter comes a point of high contention: the ranking of each of school's football team.

Before the College Football Bowl Subdivision (FBS) season even begins, a group of independent sports writers known as the Associate Press (AP) release their preseason rankings of the top 25 schools. At the conclusion of the each week of games, these top 25 rankings are updated (note that  rankings after the 25th school are generally considered to be too noisy to be a great indicator of relative strength).

To uphold their reputation as an independent and objective news organization, we may expect that the AP could potentially have any team occupy the highest rank after each week. However, if we have learned anything from behavioural economics (not to mention specifically from my previous posts) we know that humans tend to anchor on reference points. If this were true, we would expect to see that each team's past ranking strongly affects their current ranking.

To test this theory, I collected the AP rankings and game results for 2016 FBS season from Sports-Reference.com. I then classified the top 25 rankings into one of four categories: 1 to 5, 6 to 10, and 11 to 25. All other teams were classified as Not Ranked. 

I then ran an ordered logistic regression which I used to predict how each team will be ranked following their game. In order to predict each team's new rank, I used only three pieces of information: their own rank before the game, their competitor's rank before the game, and the score differential between the two teams.

For the abstract mathematically inclined, the model is of the following form:
Rank Categoryi,t = f(Rank Categoryi,t-1, Rank Categoryj,t-1, Score Differential)

Surprisingly, my initial results suggest that a team's rank before the game explains an incredible 83% of team's ranks after the game!

In attempt to help visualize the results, I have created an interactive graph for you to play with. By altering these three variables - own rank, opponent rank, and score differential - the graph below will display the predicted AP ranking for the next week. All you need to do is fill in the green cells by choosing the ranking category for each team and a non-zero integer for the score differential! Then the predicted new rank will appear below!

As this is my first attempt at an interactive graph, I would love to hear feedback on it, especially if you are having any sort of difficulties with it.



Note that the rankings are actually based on valuable information. In other words, the AP does not randomly assign rankings but instead chooses to give a team a top rank if they are among the top teams in the nation. However, evidence for anchoring can potentially be illustrated through the initial assignment a preseason ranking: these preseason rankings can have a major impact on a team's final ranking. To demonstrate this, I reassigned the preseason ranking of every team and simulated the season to predict their final ranking. Using this methodology way we can test questions such as:
  • Would perennial top 5 Alabama still be in the top 5 if their preseason ranking was outside of the top 25? 
  • If an otherwise not ranked team instead had a top 5 preseason ranking, would they be able to stay in the top 5?
In short: yes, and yes.

For a longer explanation, I began by reassigning all preseason rankings as Not Ranked for all teams (but their opponents keep their true ranks) and simulated the season. In this world, as summarized in the table below, the model predicts only 4 teams could break into the top 25. Here we see that Alabama proves to be a true top 5 team and is heavily favoured by the model regardless of their preseason ranking.


In the second but-for world, I reassigned all of the teams as top 5. Unsurprisingly, many more teams are able to stay in the top 5 than were able to get into the top 25 as seen above. According to the model, even Auburn had an acceptable season and could have stayed in the top 5.



To contextualize this, at the end of the season the top four FBS teams get to play for the College Football Playoff National Championship (although these teams are ranked by a different set of people but they too may be influenced by the same anchoring effect). In 2016, these three games which make up the championship drew a total of nearly 65M viewers. Past research has shown that colleges with a successful football team experience an influx of donations, student applications, enrollment, and student quality.

In the end, as long as there is some anchoring in the FBS ranking system, the best FBS teams will not necessarily always get the opportunity to be the best team in the country. Unfortunately, these real life biases could hurt a school in the long term through missed private donations and missed high quality students.

Until then, may hopefully the best team win gets ranked higher.

Monday, October 2, 2017

Updated 2017 Cy Young Predictions


Cy Young Predictions 
As of October 2, 2017:
AL Cy Young winner:
  • Corey Kluber - Cleveland Indians (99.7% chance of winning)
  • Chris Sale - Boston Red Sox (0.3%)
NL Cy Young Winner:
  • Max Scherzer - Washington Nationals (89.3% chance of winning)
  • Stephen Strasburg - Washington Nationals (6.4%)
  • Gio Gonzalez - Washington Nationals (3.6%)
  • Clayton Kershaw - Los Angeles Dodgers (0.7%)

Sunday, October 1, 2017

Drive for Show, Putt for ... Loss Aversion?

TLDR: "The objective of golf is rather simple: complete the hole, round, and tournament in as few strokes as possible."

"However, ... golfers use very different strategies when putting based on how many strokes they have already taken relative to par [which]  can be explained by loss aversion."


"[...] We may expect to see some loss aversion on the putting green. Loss aversion would suggest that golfers care less about making an eagle or birdie putt than they care for making a par or bogey putt."

I find that in expectation "a golfer [makes] a putt for par 22 percentage points more often than if we were to put a ball on the exact same spot of the green and told him he's putting for a birdie."

Aside: Although golf is scored by the total number of strokes an individual required to finish the round, there are a number of odd terms to describe the performance of a golfer on each hole relative to the suggested number of strokes required. A few worth mentioning are:

Par - the suggested number of strokes for the given hole.
Eagle - two strokes less than par.
Birdie - one stroke less than par.
Bogey - one stroke more than par.
Double-Bogey - two strokes more than par.

The objective of golf is rather simple: complete the hole, round, and tournament in as few strokes as possible. However, each hole is given an explicit reference point to judge the quality of a performance on that hole, known as par.

A rational golfer would ignore par as a reference point. Instead, the rational golfer would rely on other pieces of information to inform herself how to approach the hole: distance to pin, location of hazards, the slope of the green, wind speed, etc. However, it has been shown that golfers use very different strategies when putting based on how many strokes they have already taken relative to par.

This phenomenon can be explained by loss aversion. Loss aversion tells us that humans tend to alter their actions to avoid losses more than they would alter their actions to obtain the equivalent gains. Stated alternatively, most people agree that losing $10 is more painful than gaining $10 is pleasing.

In the context of golf, we may expect to see some loss aversion on the putting green. Loss aversion would suggest that golfers care less about making an eagle or birdie putt than they care for making a par or bogey putt. Under this hypothesis, golfers would try much harder to avoid the loss of one stroke (i.e. not miss a putt for par) than they would try to gain one stroke (i.e. make the putt for birdie).

To test this hypothesis, I look at every putt from a 2016 Professional Golfers' Association (PGA) tournament (I apologize, I am only looking at the men's tournaments this time round). The data was collected and uploaded to Github by Brendan Sudol. The final set of data includes information on the tournament, the player, the hole, the par, the count of strokes by the player, and the distance to the pin before and after the shot.

I begin by simply calculating the percent of times each golfer makes a successful putt given the score he is currently lying (i.e. his score for the hole should the putt go in). As I hypothesized, golfers tend to focus to avoid a loss (i.e. not miss a putt for par) significantly more than they do to achieve a win (i.e. make the putt for birdie or eagle).


Now, you may question whether it is fair to look at such average success rates: it is not unreasonable to assume there is some sort of endogeneity in the data. For example, some golfers may specialize in driving the ball from the tee at the expense of quality putting. In this scenario, the golfers that can drive a ball well enough to to the green to set up an eagle putt may need an extra putt or two to save par. If this were true, we would almost necessarily see a low success rate for eagle putts relative to par putts.

For this reason, I calculate the conditional success rate by the score the player is putting for. By conditional, I insinuate controlling for variables such as the tournament, the player, the round within the tournament, the hole, and how far the ball is from the pin.  In other words, in the graph below, we can consider three all-but-identical putts, the only difference is in the golfer's score for the hole should he be successful.

As demonstrated in the graphic above, we would expect a golfer to make a putt for par 22 percentage points more often than if we were to put a ball on the exact same spot of the green and told him he's putting for a birdie.

If this sounds absurd, it is. At the end of the tournament, all that matters is the sum of the total strokes taken throughout each hole. Each individual, even that guy that drives the cart onto the green, is strictly better off by sinking a putt regardless of what the suggested par is for that hole. However, the concept of loss aversion is so innate to humans, even professionals cannot escape it.

Paraphrased by the authors of Scorecasting: "Golfers are so concerned with a loss that they are more aggressive in avoiding a bogey that they are in scoring a birdie."

Friday, September 1, 2017

Saving your best for last: the recency effect in Cy Young voting

TLDR: "[...] thanks to a cognitive bias in humans, September [baseball is] where pitchers can make up ground in the race for the title of their league's best pitcher - the Cy Young winner."

"[...] A strong performance in September is roughly twice as influential for a pitcher's Cy Young voting points as the exact same performance would have been in April."


September is known as the most important time in Major League Baseball (MLB): teams are jockeying for position in their divisions and each game begins to weigh more heavily on their chances of making the playoffs. But thanks to a cognitive bias in humans, September also becomes a time where pitchers can make up ground in the race for the title of their league's best pitcher - the Cy Young winner.

The Cy Young Awards are given each year to the pitcher who accumulates the most vote-points in both the American and National League (abbreviated AL and NL, respectively). Thirty members of the Baseball Writers Association of America (BBWAA) rank their top five pitchers in each league and the votes are weighted accordingly:


1st - 7 points
2nd - 4 points
3rd - 3 points
4th - 2 points
5th - 1 point

Since the voting takes place at the end of the regular season (yet before the beginning of the post-season), we might expect to see evidence of the recency effect, that is, the tendency for humans to more easily recall events that have occurred more recently than those which occurred in the distant past. In the context of Cy Young voting, we would expect to see the recency effect if voters tend to overlook pitcher performances that occurred earlier in the season and voted based on more recent performances.

To illustrate an example of the recency effect, look no further than last year's AL Cy Young voting (perhaps the recency effect to consider the most recent voting period). Rick Porcello's win over Justin Verlander was highly contested (especially by the runner-up's wife, Kate Upton) and the graph below may help to explain why.

I plot the cumulative game scores of the top 4 starting pitchers in the 2016 AL Cy Young voting race. A game score is a measure that indicates the dominance of a starting pitcher's performance with a baseline score of 50. Games scores above the 50-point baseline indicate stronger performances and game scores below 50 indicate poor pitching performances. For this exercise, I subtract the 50-point baseline from each game score to help illustrate the data. In the left panel below, I begin summing each individual's game scores from the commencement of the 2016 season and focus on the final two months of the regular season (plus the first two days in October). Here Rick Porcello (in Red Sox red) looks like a poor choice for the Cy Young. However, in the right panel, I begin summing the games scores from August 1st, 2016 until the end of October. Now the choice for the Cy Young becomes less obvious. (Note that the number of voting points each pitcher received are in parenthesis next to their name.)


To test my conjecture empirically, I collect information on all pitchers from the 2015 and 2016 seasons who:
  1. Started a minimum of 20 games, and;
  2. Finished in the top 10 in their respective league for the following categories:
    • Earned Run Average (ERA),
    • Wins, or;
    • Strikeouts.
I then calculate the game scores for each pitcher (recall that I do not add in the 50-point baseline to the game scores of each pitcher). All my data comes from Baseball-Reference.com. I then calculate the correlation between the BBWAA voting points and the average game score of a pitcher for a given month. The results are depicted graphically below.
As I expected, there appears to be evidence of the recency effect in Cy Young voting. A quality pitching performance in April does not correlate nearly as strongly as a quality pitching performance in September. A simple regression suggests that a strong performance in September is roughly twice as influential for a pitcher's Cy Young voting points as the exact same performance would have been in April. This is consistent with the recency effect hypothesis that the voters tend to forget the performances in early April and remember more recent performances in August and September when determining their choices for Cy Young Awards.

I am not the first to illustrate that the BBWAA does not always correctly award the best pitcher in baseball with the award for the best pitcher in baseball. But thanks to the recency effect, the BBWAA gives empirical evidence to the phrase "you are only as good as your last performance."

If you are enjoying what you are reading, I would love to hear from you! Please like, share, and comment below!



Predictions 
As of September 1, 2017, my simple model mentioned above predicts the following:
AL Cy Young winner:
  • Corey Kluber - Cleveland Indians (84.6% chance of winning)
  • Chris Sale - Boston Red Sox (15.4%)
NL Cy Young Winner:
  • Max Scherzer - Washington Nationals (76.3% chance of winning)
  • Gio Gonzalez - Washington Nationals (23.5%)
  • Clayton Kershaw - Los Angeles Dodgers (0.2%)

Tuesday, August 1, 2017

.299 hitters don't want to be discounted

TLDR: "The concept of psychological pricing and our perception of numbers ending in 99 is not unique to the grocery store. Undervaluing a baseball player with a batting average of .299 is consistent with this idea."

"[...] a player going into the final game of the season with a .299 batting average will aggressively chase the goal of hitting .300. [They] swing at 11 percentage points more pitches than the average batter [...] and swing at 63% of pitches on the last day of the season - all for the allure of the ending the season hitting .300"


Aside: In baseball, one of the most common metrics used to discuss the prominence of hitter is the batting average. The batting average is calculated as an individual's number of hits divided by the number of at bats and it is always represented as a three decimal number (e.g. if someone had 5 hits in 20 at bats, we would say he is hitting .250). In 2016 the MLB-wide batting average was .255.

How different is a .299 hitter from a .300 hitter? It turns out about $130 thousand per year. That is how much more money a Major League Baseball (MLB) player would expect to make on their next contract after just one more hit during the regular season.

Valuing .300 significantly more than .299 is consistent with the concept of psychological pricing - the reason why we always see a can of pop selling for $0.99 instead of $1.00. Humans tend to incorrectly perceive prices ending in .99 as meaningfully lower than those ending with one cent more. 

In the context of professional baseball, MLB players have to bargain for their wage from prospective teams as part of free agency. Part of the team's due diligence during contract negotiations requires mulling over the player's performance such as their batting average. (For an analogous discussion using the National Football League, see my earlier post). Previous research suggests MLB teams value a .300 hitter much more than .299 hitter (despite the negligible difference) and the former tends to receive a larger contract.

Alternatively, we might say that .300 is a reference point for which baseball players and general managers use as an anchor. The perceived value of a player is now measured against this standard of hitting .300: any player above this number are considered the masters of their craft and should be compensated accordingly.

In either case, we would expect a player sitting just shy of the .300 mark to show more aggression in their at bats to reach this achievement. One place we can look to test this idea is through the number of pitches these players are swinging at. Moreover, when the players know they are near their last chance to improve their batting average we would expect to see their aggression increase. Therefore, I chose to look at the number of pitches a player hitting .299 swings at on the final day of season.

I started by collecting all the pitches thrown on the final day of the regular seasons of 2014 to 2016 from Baseball Savant. Then I used Fan Graphs to identify which players begun the final day of the season hitting .299 and .300. (I term a player who goes into the final game hitting .299 a ".299 hitter", similar with .300 hitters). I then classify the reaction by the batter as either a swing, no swing, or neutral. Swings include putting the ball in play and swinging strikes - this is where the intention is to make contact for a hit. No swings are where the batter does not swing and includes balls and called strikes. Lastly, neutral reactions are when the outcome was already decided: intentional balls, wild pitches, pitch-outs, etc.

I remove the neutral reactions and I calculate the fraction of pitches that resulted in swings. I separate the data into three groups: .299 hitters, .300 hitters, and all hitters (regardless of batting average). The results are as follows:

With just one more hit needed to reach the .300-mark, .299 hitters are significantly more like to swing in their at bats. These .299 hitters swing at 53% of pitches, 11 percentage points more than a .300 hitter.

When we compare these .299 hitters against themselves, their own strategy completely flips on the last day of the season (using the 2016 data). Whereas they were swinging the bat on 38% of pitches for the majority of the season, these .299 hitters swing 63% of the time on the final day of the season.

The concept of psychological pricing and our perception of numbers ending in 99 is not unique to the grocery store. Undervaluing a baseball player with a batting average of .299 is consistent with this idea. Here I have demonstrated that a player going into the final game of the season with a .299 batting average will aggressively chase the somewhat arbitrary goal of hitting .300. 

Whether they are aware of the financial payoffs or they have set an internal reference point, .299 hitters swing at 11 percentage points more pitches than the average batter on the final day of the season. Moreover, while .299 hitters do not swing at 62% of pitches for the majority of the season, they swing at 63% of pitches on the last day of the season - all for the allure of the ending the season hitting .300!

If you like what you are reading, I would love to hear your feedback! Please like, share, and comment! 


Saturday, July 1, 2017

When the game is on the line, umpires make the safe call, not the right call

TLDR: "In [the example below], the probability of having an umpire invoke the Infield Fly rule is over 50% when the outcome of the at bat is relatively inconsequential [...] Conversely, in a tight game, the same batted ball is significantly less likely (~5%) to be called an Infield Fly."

"[T]hese results are consistent with the concept of Defensive Decision Making: [...] umpires may be choosing the more defensible action of not signaling for the Infield Fly instead of the action they may think as more appropriate."

In a previous post, I discussed how Major League Baseball (MLB) umpires tend to make errors more often when they have called two strikes or two balls in a row. I likened this to the Gambler's Fallacy, wherein the umpires mentally assigns a falsely low probability of witnessing (e.g.) three consecutive strikes, and therefore think the third pitch 'must' be a ball.

I offered another explanation - Defensive Decision Making -  where the umpire may be more likely to make the call that is more defensible in place of the call they see as best. When the hitter is advantaged in the count at three balls and zero strikes, we saw the umpire is more likely to call a close pitch a strike instead of correctly calling it a ball. The umpire may be appealing to a more defensible perception of fairness rather than the true call.

Continuing with this theme, others have suggested is that Defensive Decision Making is more evident when the stakes of the game increase. Previous studies have pointed to the perceived change in the way fouls and penalties are called in the playoffs versus the regular season for the National Basketball Association and National Hockey League, respectively. It has already been shown that sports viewers have an omission bias - fans tend to think a bad call by an official is worse than a non-call - and the defensive decision in this scenario would be the absence of excessive fouls/penalties. Superficial analyses of these situations reveal there is no statistical change in the number of fouls/penalties per game. These results are not surprising, given that these studies fail to address the change in strategy and style of play of the teams in the playoffs.

I attempt to uncover evidence consistent with the notion that Defensive Decision Making would occur more often in situations of heightened importance. In order to avoid the pitfalls of previous studies, I look at a situation where the officials must make a judgement call that is independent of team strategy - this is to say that the events requiring a subjective decision are as best as random. Therefore, I have chosen MLB's Infield Fly Rule, which is summarized below :

If all of the following conditions are met and the umpire signals an Infield Fly, the batter is automatically called out, regardless of the outcome. The conditions are as follows:
  1. First and second must be occupied with less than two out (third may or may not be occupied).
  2. The batter must hit a fair fly ball, which is not a line drive nor bunt, that;
  3. In the umpire's judgement, can be caught by an infielder with ordinary effort.
H/T to Close Call Sports which have some great examples of the rule if you are still confused.

Note that in (3) the focus is the umpire's judgement.

I have collected all the instances from the 2016 and 2017 MLB regular seasons (up to 3 May 2017) which meet the criteria of (1) and (2) as above from Baseball Savant's Statcast Search. This data provides the angle and distance each batted ball traveled under conditions (1) and (2) and whether the play was called an Infield Fly or otherwise. Below is a graphical depiction of the 7,001 such occurrences. Each point represents the batted ball of one at bat with red indicating an umpire signaled an Infield Fly.

I then assign each observation a value to indicate the importance of the at bat's outcome to the overall outcome of the game. This is known as the Leverage Index - it accounts for the inning, the number of outs, the runners on each base, and the difference in runs of each team. Higher values indicate the outcome of the at bat will have a larger effect on the outcome of the game.

Note that in criteria (2) above, the batted ball must be a "fair fly ball." Fly balls are determined by the height the ball travels during its flight. Also note that in criteria (3) the ball must be able to be "caught by an infielder with ordinary effort" which indicates it cannot travel too far from home plate (hence the name Infield Fly). If the umpires were acting truly independent of any biases, one would expect that only a ball's height and distance would have explanatory power of whether a batted ball is called an Infield Fly. Indeed it is true that these two characteristics explain most of the variation in the calls using a simple logit model:

prob(Infield Fly) = f(height, distance)

This model has a pseudo-R2 of 0.84 and correctly classifies the Infield Fly calls 99.2% of the time (i.e. the model and the umpires agree 99.2% of the time).

Next I add in the Leverage Index to the model. After experimenting with several specifications, I find that there is a significant negative correlation between the probability of the umpire calling for an Infield Fly and the importance of the at bat.

Below is a graphical depiction of how the probability of having the Infield Fly called declines as the leverage of the at bat increases when holding the height and distance of the batted ball constant. In this example, the probability of having an umpire invoke the Infield Fly rule is over 50% when the outcome of the at bat is relatively inconsequential to the overall outcome of the game. Conversely, in a tight game, the same batted ball is significantly less likely (~5%) to be called an Infield Fly.


I find that these results are consistent with the concept of Defensive Decision Making: the defensive decision in this scenario is to force the fielder to make the catch to get the out. Under this hypothesis, umpires may be choosing the more defensible action of not signaling for the Infield Fly instead of the action they may think as more appropriate.

I have shown there exists some evidence that Defensive Decision Making in one aspect of baseball, however there is no indication that Defensive Decision Making is limited to just the Infield Fly Rule. In fact, Defensive Decision Making may be prevalence in all sorts aspects of each and every game yet we lack the ability to test it empirically. Until then, this proves once again that there are many human biases that present themselves in everyday life, and umpires are no different the rest of us.

Thanks for reading, please offer you questions and comments below!

Thursday, June 1, 2017

Do MLB umpires want to be fair, or just perceived as fair?

TLDR: "A strike is 5.9 percentage points more likely to be incorrectly called a ball if the previous two pitches were also strikes."

"When faced with the prospect of making the same call three times in a row - thus creating a 'pattern' of calls - umpires make erroneous calls nearly six percentage points more often than in situations without obvious patterns."

In the average baseball game where the umpire is forced to make about 150 subjective calls, there are bound to be a few mistakes. But what if the umpire is systematically making erroneous calls?

As a Major League Baseball (MLB) umpire, one must be perceived as neutral. Since each pitch is a game between the pitcher and batter, making too many calls advantaging one or the other may be construed as favouritism. While umpires want to appear neutral, natural human biases may lead to a fear that patterns in their subjective calls are seen as anything but.

I want to uncover if there is a systematic bias to avoid these so-called 'patterns' in the called balls and strikes by MLB umpires. In order to tackle this question, I first collected all the calls by a home-plate umpire in the MLB from 2016 by Baseball Savant's Statcast Search. I am only concerned with the pitches that required the umpire to decide whether the pitch was either a strike or a ball. This excludes any time that the batter swung the bat or any time the pitcher threw a ball intentionally or threw a pitch into the dirt. What I am left with is 360,278 pitches for which an umpire had to make a subjective call, answering the question 'was the pitch a strike or a ball?'

The data also tracks the location of where the ball passed through the strike zone: this identifies whether the pitch was a 'true' strike, as per the MLB definition. Comparing the umpire's call to the true call, I find that the correct call was made 88.8% of the time. In other words, the error rate of umpires is 11.2%. The false-positive rate for strikes (i.e. called strikes that were actually balls) was 17.2%, which, although high, is similar to other's findings.

I then identify a few situations in the ball and strike calls to test my theory that umpires avoid patterns. The first pattern I identify I term X-X-Y, where X is either a ball or a called strike, and Y is the opposite (e.g. three consecutive pitches that were called strike, then another strike, then a ball). There are 17,481 times this pattern occurs in the data. Shockingly, the unconditional error rate of the Y in the sequence of X-X-Y is 14.4% - a 3.2 percentage-point increase from the overall error rate. This is the first piece of evidence that umpires avoid patterns by erring more often when the correct call would be the third consecutive strike or third consecutive ball than they err in another situation.

Next, I consider a different counterfactual to the X-X-Y pattern, which I call X-Y-Z. Again, the X is either a ball or a called strike and Y is the reciprocal. I let Z be either a ball or a strike (i.e. I allow for either a X-Y-X or a X-Y-Y pattern). If the umpire avoids a pattern of three of the same call, an X-Y pattern should have no impact on the call of the third pitch. This X-Y-Z pattern occurs 28,271 times in the 2016 data and the unconditional error rate of the Z pitch is 9.4%. This means that umpires make an erroneous call on the third pitch of a X-X-Y pattern 5 percentage points more often than that of a X-Y-Z pattern!

Up until this point, I have only considered unconditional probabilities (i.e. not considering other factors that may affect the error rate of the umpire). This includes the speed and type of pitch, the location of where the pitch crosses the plate, or even the inning of the game. In order to control for all these necessary characteristics, I run a logit model to predict an erroneous call. After playing with the specification, I settled on the following model:

prob(wrong call) = f( X-X-Y, pitch type, release speed, spin rate, location, inning, top/bottom of inning)

Pitch type is a set of dummies indicating the type of pitch thrown (four-seam fastball, curveball, etc.). Release speed is the velocity in which the ball is thrown measured in miles per hour, spin rate is the number of times the ball rotates while traveling from pitcher to catcher in revolutions per minute. Location is a set of dummies corresponding to the location of where the pitch crossed the plate (see this diagram for more information), and top/bottom of inning is a dummy indicating whether the game is in the top or the bottom of an inning. Lastly, X-X-Y is a dummy that indicates the pitch is the Y in a X-X-Y sequence.

After running the model, I calculate the marginal effect of the X-X-Y dummy. After controlling for pitch characteristics, being in a X-X-Y situation increases the probability of a wrong call on the third pitch by 5.9 percentage points over a X-Y-Z situation. This can be stated alternatively as a strike is 5.9 percentage points more likely to be incorrectly called a ball if the previous two pitches were also strikes.

Below is a graphical depiction of all the aforementioned probabilities.

If umpires did not care about the perceived fairness, the error rate would not be different for a given situation over another. This effect may be a conscious or subconscious decision to keep the balance of calls which favour the pitcher and batter. Another explanation is that umpires suffer from the Gambler's Fallacy and falsely assign a low probability to the event of three consecutive strikes or three consecutive balls. In the Gambler's Fallacy scenario, the umpire would think that the likelihood of three strikes in a row is so small, the third pitch 'must' be a ball.

In summary, when faced with the prospect of making the same call three times in a row - thus creating a 'pattern' of calls - umpires make erroneous calls nearly six percentage points more often than in situations without obvious patterns.

Please feel free to offer comments, questions, or even suggestions for future work!