On the second point - fair enough, though even under Bayes it's sometimes reasonable to want a single answer on account of you only get to actually do one thing.
If you have that prior and you maximize P(model|data) on solutions with a zero probability mass on either P(data|model) or P(model), you're screwing up multiplication.
Well, the point is that if you have a continuous-space, then the maximum-likelihood solution will have zero entries with positive probability, but the posterior probability of a zero entry is 0.
Question in title.
This is obviously subjective, but I figure there ought to be some "go-to" paper. Maybe I've even seen it once, but can't find it now and I don't know if there's anything better.
Links to multiple papers with different focus would be welcome. For my current purpose I have a preference for one that aims low and isn't too long.