Wednesday, September 9, 2026

Too little information

     I've seen a few things on social media saying that the Republicans stole the 2024 presidential election, but didn't pay much attention to them until a recent post by Andrew Gelman pointed to an article in Votebeat by Carter Walker and Jessica Huesman.  The basic story:  Walter Mebane, a political scientist at the University of Michigan, has developed a model that is supposed to detect election fraud.  When he applied it to data from Pennsylvania, he got estimates ranging from about 25,000 to 200,000 "fraudulent" votes.  Trump's margin in the state was about 120,000, so the analysis suggests that voter fraud might have made the difference.  But other political scientists criticized his analysis and said that there was no evidence of significant voter fraud. Although Walker and Huesman found the critics more credible, that seemed to be based on personal impressions rather than evaluation of Mebane's model: they speak of "complex assumptions and unfamiliar methods."  So here is my attempt to state the assumptions in a reasonably simple fashion.   

The model uses only two pieces of information:  turnout rates and party support at the precinct level.  How can you get evidence of fraud from such limited information?  Mebane's model implements the ideas offered a paper by Peter Klimek et al.  They propose that very high turnout and lopsided support for one party may be a sign of fraud.  This seems reasonable--if one party has complete control over the voting or counting of votes, you would expect them to run up the score.  But you don't need complex methods to identify those cases of potential fraud--you can just look at the figures for turnout and votes.  Their second kind of fraud, which they call "incremental fraud," involves adding (or switching) a moderate number of votes in a large number of precincts--e. g., changing 60% turnout and 50% Republican vote to 70% turnout and 60% Republican vote.  But in this case, you can't just look at the numbers for individual districts, since they aren't unusual.  Instead, you compare the distribution of the total numbers--are there more 70%/60% districts in the voting records than there are in reality?  But then you need a standard for "reality"--that is, the true distribution of turnout and votes in different precincts.  Klimek et al. assume that it follows a normal distribution, and I think Mebane does too.**  But there's no compelling theoretical reason to think that the actual distribution of turnout and voting choices follows a normal distribution, or any other specific distribution.  Another justification for using the normal distribution as the standard would be if data from elections generally agreed to be fair almost always followed a normal distribution, but there doesn't seem to be any evidence of this kind.  So we can't be confident about the true distribution, and therefore can't say whether the observed distribution differs from it.  Even if there is "incremental fraud," there's no way to identify it from the data on turnout and voting choices alone.  Cases of extremely high turnout and lopsided support may result from fraud (although they may occur naturally), but Mebane identifies only nine precincts like that out of about 9,000, involving a total of about 2,000 votes, which is much smaller than the margin in Pennsylvania.     

  * Another part of the story was the use (and maybe misuse) that the Election Truth Alliance had made of Mebane's analysis, and how much responsibility he had for that, but I'll leave it aside and just consider the model.

**I say "I think" because I found his paper hard to follow and I wasn't really motivated to figure it out--the important point is that there's no justification for any assumptions about the true distribution.