> one problem I've always had with Silver is that he's very smug about this models but they often enough don't work out.
What does it mean for a probabilistic prediction to "not work out"? If I tell you that the odds of flipping 2 heads in a row with a fair coin is only 25%, and it happens, did my model "not work out"?
> instead of reflecting on his models to improve them (at least publicly,) he just tells people they don't understand probability or his models.
I really don't intend this in any sort of rude way, but I think you might be interpreting his response as dismissive because it's accurate...
It's difficult to say a model "doesn't work out" based on a single event, especially when it puts a significant probability on the less-likely option. For example, 538's 2016 forecast gave Trump a roughly 1 in 3 chance of winning: that's _higher_ than the odds of flipping 2 heads in a row.
Measuring the quality of a model is more complicated: one way would be applying the model to the same outcome multiple times, but this isn't possible for single-event forecasts. Another way to see how calibrated your predictions are is to aggregate multiple predictions and see how often they line up with reality: an event predicted to occur with X% probability should occur X% of the time.
Lucky for us, 538 has done exactly this analysis[1]! Naturally, it's internal, so take it with a grain of salt, but it looks like their predictions are fairly well-calibrated.
What does it mean for a probabilistic prediction to "not work out"? If I tell you that the odds of flipping 2 heads in a row with a fair coin is only 25%, and it happens, did my model "not work out"?
> instead of reflecting on his models to improve them (at least publicly,) he just tells people they don't understand probability or his models.
I really don't intend this in any sort of rude way, but I think you might be interpreting his response as dismissive because it's accurate...
It's difficult to say a model "doesn't work out" based on a single event, especially when it puts a significant probability on the less-likely option. For example, 538's 2016 forecast gave Trump a roughly 1 in 3 chance of winning: that's _higher_ than the odds of flipping 2 heads in a row.
Measuring the quality of a model is more complicated: one way would be applying the model to the same outcome multiple times, but this isn't possible for single-event forecasts. Another way to see how calibrated your predictions are is to aggregate multiple predictions and see how often they line up with reality: an event predicted to occur with X% probability should occur X% of the time.
Lucky for us, 538 has done exactly this analysis[1]! Naturally, it's internal, so take it with a grain of salt, but it looks like their predictions are fairly well-calibrated.
[1] https://projects.fivethirtyeight.com/checking-our-work/