Second of all, we still don't really know what "best" means, and it's entirely possible different methods are best for different people in complex ways.
And there's a worse confounding factor, which is that people tend to interpret instructions in terms of whatever prior model they have. (That's actually why I object so strenuously to a couple of aspects of your "oath" model -- they're not so much intrinsically harmful, as harmful to people with certain prior models.)
Testing the distinction between your method and mine would require pretty stringent behavioral control of subjects in a large experiment, because you'd need to validate that the subject actually considered each situation and consequence. (Writing those things out is a good way to verify it, which is why I think your success was actually a side-effect of the thinking you had to do in order to design and write your oaths.)
However, if you just grab a bunch of volunteers and tell them to do either your version or mine of that process, I predict that a substantial number will not actually follow the directions, and will simply tell themselves they've already thought it through enough after considering maybe 1 or 2 situations, and then proceed to do whatever it is they already do to initiate change effects, sprinkled with a bit of flavor from whatever method they're supposed to be testing.
This is a major confounding factor in testing any cognitive behavior model, be it a self-help technique, time management system, or anything else. People tend to process virtually all new inputs through whatever mental strategies they already have, and lop off the parts that don't fit.
All we can feasibly get is the intent-to-treat effect. Estimating actual treatment effects is possible but not practical.
Is that a fair summary of the parent?
It seems to me that this blog has just reached it's first real crisis.
Three people are announcing three apparently opposed beliefs with substantial real expected consequences and yet no-one has yet spoken, or it seems to me implied, the key slogan... "LETS USE SCIENCE!" or, as hubristic Bayesian wannabes, not invoked Bayes as an idol to swear by, but rather said "LETS USE HUMANE REFLECTIVE DECISION THEORY, THE QUANTITATIVELY UNKNOWN BUT QUALITATIVELY INTUITED POWER DEEPER THAN SCIENCE FROM WHICH IT STEMS AND TO WHICH OUR COMMUNITY IS DEVOTED".
IF RDS was applied to our current situation, people would be analyzing Yvain's, Davis' and Eby's proposals, working out exactly what their implications are, and trying to propose, in the name of SCIENCE, hypotheses which will distinguish between them, and in the name of BAYES, confidence estimates of their analyses and of the quality with which the denotations of their words have cleaved reality at the joints enabling an odds ratio of updating to be extracted from a single data point. People would be working out what features of which of the models used by Yvain, Davis and Eby constitute evidence against what other features. They would be trying to evaluate non-verbally, through subjectively opaque but known-to-be-informative processes vulnerable to verbal overshadowing, what relative odds to place on those different features of the models. Finally, they would be examining the expected costs entailed by experiments being proposed and selecting those experiments which promise to provide the most information for the least cost be performed. The cost estimate would include both the effort required to perform the experiments, probably best assessed with an outside view in most cases like these, and the dangers to the minds of the participants from possible adverse outcomes, taking into account, as well as possible, the structural uncertainty of the models.
I sincerely hope to see some of that in the comments section soon, either under this post or the "Applied Picoeconomics" post.