I may as well repeat my thoughts on Newcomb's, decision theory, and so on. I come to this from a background in decision analysis, which is the practical version of decision theory.
You can see decision-making as a two-step, three-state problem: the problem statement is interpreted to make a problem model, which is optimized to make a decision.
If you look at the wikipedia definitions of EDT and CDT, you'll see they primarily discuss the optimization process that turns a problem model into a decision. But the two accept different types of problem models; EDT operates on joint probability distributions and CDT operates on causal models. Since the type of the interior state is different, the two imply different procedures to interpret problem statements and optimize those models into decisions.
To compare the two simply, causal models are just more powerful than joint probability distributions, and the pathway that uses the more powerful language is going to be better. A short attempt to explain the difference: in a Bayes net (i.e. just a joint probability distribution that has been factorized in an acyclic fashion), the arrows have no physical meaning--they just express which part of the map is 'up' and which is 'down.' In a causal model, the arrows have physical meaning--causal influence flows along those arrows only in directions with arrows, and so the arrows represent which direction gravity pulls in. One can turn a map upside down without changing its correspondence to the territory; one cannot reverse gravity without changing the territory.
Because there are additional restrictions on how the model can be written, one can get additional information out of reading the model.
I am currently learning about the basics of decision theory, most of which is common knowledge on LW. I have a question, related to why EDT is said not to work.
Consider the following Newcomblike problem: A study shows that most people who two-box in Newcomblike problems as the following have a certain gene (and one-boxers don't have the gene). Now, Omega could put you into something like Newcomb's original problem, but instead of having run a simulation of you, Omega has only looked at your DNA: If you don't have the "two-boxing gene", Omega puts $1M into box B, otherwise box B is empty. And there is $1K in box A, as usual. Would you one-box (take only box B) or two-box (take box A and B)? Here's a causal diagram for the problem:
Since Omega does not do much other than translating your genes into money under a box, it does not seem to hurt to leave it out:
I presume that most LWers would one-box. (And as I understand it, not only CDT but also TDT would two-box, am I wrong?)
Now, how does this problem differ from the smoking lesion or Yudkowsky's (2010, p.67) chewing gum problem? Chewing Gum (or smoking) seems to be like taking box A to get at least/additional $1K, the two-boxing gene is like the CGTA gene, the illness itself (the abscess or lung cancer) is like not having $1M in box B. Here's another causal diagram, this time for the chewing gum problem:
As far as I can tell, the difference between the two problems is some additional, unstated intuition in the classic medical Newcomb problems. Maybe, the additional assumption is that the actual evidence lies in the "tickle", or that knowing and thinking about the study results causes some complications. In EDT terms: The intuition is that neither smoking nor chewing gum gives the agent additional information.