Showing posts with label Mind Projection Fallacy. Show all posts
Showing posts with label Mind Projection Fallacy. Show all posts

Saturday, April 18, 2015

The Fundamental Confidence Fallacy


The title of this post comes from an excellent recent paper (as far as I can tell, still in draft form) on misunderstandings of confidence intervals. The paper, 'The fallacy of placing confidence in confidence intervals', by R. D. Morey et al.1 is by almost exactly the same set of authors whose earlier paper on a very similar topic I criticized, before, but the current paper does a far better job of explaining the authors' position, and arguing for it.

The authors identify the fundamental confidence fallacy (FCF) as believing automatically that,
If the probability that a random interval contains the true value is X%, then the plausibility (or probability) that a particular observed interval contains the true value is also X%.

Monday, June 17, 2013

Extreme values: P = 1 and P = 0




There is a popular folk theorem among some Bayesians, to the effect that it is unacceptable for a probability to be 0 or 1. There's a simple motivation for this principle: as rationalists, we demand the opportunity for nature to educate us by blessing us with novel observations. No matter how confident we become in some proposition, it should always be possible for us to change our minds when strong enough evidence accumulates in favour of some alternative. As Karl Popper rightly observed, after all, a theory that is invulnerable to falsification is not much of a theory.

But what happens if P(H | I) becomes zero? How is the probability for the hypothesis, H, to be updated by new evidence? If P(H | I) is 0 then the numerator in Bayes' theorem, prior times likelihood,

P(H | I) × P(D | HI), 

is also 0, regardless how convincing the data, D, may be. No matter what happens, the outcome is unchanged: a nice round posterior.

Similarly, if P(H | I) is 1, then for the converse hypothesis, P(~H | I) is necessarily 0. Now, the denominator in Bayes' theorem is 

P(H | I) × P(D | HI) + P(~H | I) × P(D | ~HI)

and when the second term (everything after the plus sign) is zero, both numerator and denominator in Bayes' theorem are the same, producing the ratio 1, for all eternity.

I have sympathy with this motivation, therefore, but as a general rule, it is utter nonsense, resulting from forgetting one of the most basic facts about how inference works. The mathematics I have just described is all correct, but there are other ways for us to change our minds, and retain our rationality.

A recent, brief discussion at another website drew my attention to an article by Eliezer Yudkowsky, in which he also argues that 0 and 1 are not probabilities. The argument is a little different: the amount of evidence (the likelihood ratio expressed in log-odds form) needed to update an intermediate probability to 0 or 1 is infinite. This infinite certainty is an absurdity, he claims, unable to be represented with real numbers, and so 0 and 1 aren't probabilities.

Yudkowsky, as many readers will know, is a widely regarded thinker and writer on the topic of applied rationality, and I can recommend his writing most highly. The overlap between his broad philosophy and mine is, I would say, very large, with the main difference that in cases where I lack mastery of the theoretical apparatus, he very often does not. Yudkowsky knows and understands the mind-projection fallacy better than the vast majority (see for example his article of the same name, and this followup), but in this instance, he seems to have forgotten it. It is essentially the same error made by all who claim that probabilities equal to zero or one should not enter one's calculations.

A little thought experiment, then, before resolving the paradox. Let H be the hypothesis that in some five-day interval, at some location on the Earth, the sun will rise on each of the five mornings. Let D represent the observation of the sun rising on the first of the mornings in question. What is P(D | HI)? I humbly submit that it is 1. Is H, therefore, not an appropriate, well-formed hypothesis? Is D not a valid observation? Evidently, if probability theory is to have any power at all, it must be capable of supporting hypotheses such as H, and data as trivial as D. It is not conceivable to have such things automatically ruled out under our epistemology.

In general, it is perfectly legal for P(D | HI) (or, for that matter, a posterior, like P(H | DI)) to be 0 or 1, but here's that basic fact about probability that we have to keep in mind: a probability can not be divorced from the model within which it is calculated. A model may imply infinite certainty, without any person ever achieving that state (which would be impossible to encode in their brain, anyway). Our notation says something very important: P(D | HI), no matter what it is, is necessarily contingent upon the conjunction HI, which obviously depends on the truth of I. This is something we can never be absolutely certain of.

The all-important "I" that forms the foundation for every Bayesian calculation is usually said to stand for 'information' - all the relevant prior knowledge we have. Unfortunately, this creates a little trap that too many fall into, which is to forget that there is another component besides information needed before "I" is fully populated. "I" could just as easily stand for 'imagination.' To get Bayes' theorem to do any useful work for us, we have to specify a theoretical framework. We have to make certain assumptions, including specification of a full set of hypotheses against which is to H compete. To arrive at a candidate set of hypotheses, we must make a leap of the imagination. There is no possible criterion for judging whether or not all our assumptions are correct, and no way to know in advance whether we have chosen the 'correct' set of hypotheses. To think otherwise is just wishful thinking.

To think that the infinite confidence implied under some "I" represents the actual infinite confidence of some physical rational agent is the mind-projection fallacy. Instead, a probability is a model of the confidence a rational agent would have if "I" was known to be true. That this confidence might need to be modelled using a non-numeric concept such as infinity is merely an uncomfortable (though often highly convenient) mathematical fact.

And now we can see how it is that we can continue to accrue knowledge under the threat of the apparent epistemological cul de sac that is P = 1 or P = 0. To liberate ourselves from the straight jacket of "I", we simply need to recognize that what we now call "I" is itself merely a hypothesis in some broader hierarchical model. This is how model checking (wielding the analytical blade of model comparison) works, which, as I pointed out before, seems philosophically unpalatable to many, yet is in fact an essential ingredient in our inferential machinery. This is how we can come to look again at our theoretical framework and say 'hold on, I should be working with a different hypothesis space.' Novel theories and scientific revolutions would be impossible without this flexibility.

Some see this need in Bayesian epistemology to make assumptions in "I" that can't be established with certainty as a severe weakness, but it isn't - at least not one that can be avoided (no matter how many black belts we hold in the ancient art of self deception). We can always extend the scope of our hypothesis space so that some of our assumptions become themselves random variables in  a wider inferential context, but to have all of them take on the role of hypotheses under test would require an infinitely deep hierarchy of models. In the example above, where H was a hypothesis about the sun rising, one might argue that a more sophisticated model would account for the possibility, however small, that my sensation of the sun rising was mistaken. Indeed, this is correct, and would prevent the likelihood function going to 1. Sooner or later, though, I'm going to have to introduce a definitive statement - one that supposes something to be definitely true - in order to avoid the intractable quagmire of infinite complexity.

The early frequentists (and some still, in private communication with me), claimed that this subjectivity of Bayesian probability is its downfall, but in reality, it is impossible to learn in a vacuum. No kind of inference is possible without assumptions. Part of the beauty of Bayesian learning is that we make our assumptions explicit. The frequentists, of course, also make assumptions (see Yudkowsky, for example), but by refusing to acknowledge them, like the fabled ostrich sticking its head in the sand, they eliminate the possibility to examine whether or not they are reasonable, to understand their consequences, or to correct them when they are manifestly wrong.




Friday, March 8, 2013

What is Randomness?



Random variables play an important part in the vocabulary of probability theory. I think there's a lot of confusion, though, about what randomness actually is. A few days ago, I found an expert statistician trying to distinguish between mere statistical fluctuation and actual changes in the causal environment. Another example that always bugs me comes from computer science, and is the ubiquitous insistence that a deterministic algorithm can not produce random numbers, but only pseudo-random numbers.   

Each of these examples commits a fallacy. The first may be an isolated slip-up, or else a deliberate attempt to gloss over technicalities with sloppy language (I may even be guilty (gasp!) of either of these myself, on occasion), but the second is almost universal within the entire profession of computer scientists, which represents a significant sample of the world's technically disciplined. If any one of those computer scientists understood what randomness is, they would recognize immediately that the need to distinguish between random and pseudo-random is entirely fictional. 

Its not that I have anything against computer scientists. Information technology, after all, is what this blog is all about: systematically processing knowledge, and modern society owes its existence to computer science. Nor do I think that computer scientists are excessively prone to the fallacy I'm talking about. In fact, the statistical literature, compiled by those who, of all people, should have dedicated considerable effort to understanding this topic, is crammed with instances, of which my initial example is representative. As another example, the current Wikipedia entry on randomness contains a very confused section that starts with the statement: 'Randomness, as opposed to unpredictability, is an objective property.'

So what is the problem? Lets look at pseudo-random numbers, first. The point about pseudo-random numbers is that they are produced in a clever way to 'replicate' a random variable - they come appropriately distributed and they appear uncorrelated, which is to say that even knowing the nth, (n-1)th, ... numbers in a list, it will be impossible to predict what the (n+1)th number will be. They are considered to be not truly random, however, because they are produced by mechanical operations of a computer on fixed states of its circuits. If we only knew the states of the circuit and the operations, then we would know the number that will come out next. To say that this prevents us from describing the numbers as random, however, is an instance of the mind-projection fallacy, as I have discussed before.

The mind projection fallacy consists of assuming that properties of our model of reality necessarily correspond to properties of reality. We often talk about a phenomenon being random. This makes it tempting to conclude that randomness is a property of the phenomenon itself, but this assumes too much, and also demands that the events we are talking about take place in some kind of bubble, where physics doesn't operate.

When I dip my hand into an urn to draw out a ball whose colour can't be predicted, the colour that comes out is rightly considered to be a random variable. But we are not talking about some quantum-mechanical wavefunction that collapses the moment the first photon from the ball hits my retina (or the moment the nerve impulse reaches my visual cortex, or any of a million other candidate moments). We are talking about a process with  real and definite cause and effect relationships. The layout of the coloured balls inside the urn, the trajectory of my hand into the urn, and the exact moment I decide to close my hand uniquely determine the outcome of the draw. Randomness is not the occurrence of causeless events, but is a consequence of our incomplete information. It's not a property of the balls in the urn, but a property of our prior state of knowledge.

It might be that at microscopic scales, quantum stochastic variability really does emerge from an absence of causation, but there are 2 important points to note in relation to this discussion. Firstly, good scientists recognize the need for agnosticism on this front - there really isn't enough evidence yet to decide one way or the other (BBC Radio 4's excellent 'In Our Time' has an episode, entitled 'The measurement problem,' with an interesting discussion on the topic). Secondly, the vast majority of cases where the concept of randomness is applied concern macroscopic phenomena, where classical mechanics is a perfectly adequate model. For these reasons, the only sensible general usage of the word 'random' is when referring to missing information, rather than as a description of uncaused events. That wikipedia article I quoted from, in apparent recognition of this, later cites Brownian motion and chaos as examples of randomness, thereby contradicting the earlier quote (though inexplicably, the two are identified as separate classes of random behaviour).

Getting back to the urn, if I knew precisely the coordinates (relative to my hand) and colours of the balls inside, the colour of the extracted sphere would not be a a surprise, and therefore wouldn't be a random variable. But under the standard drawing conditions, in which these are not known, it is a random variable. Similarly, knowing the state and operations of a deterministic computer algorithm would render its output non-random (provided I have the computational resources elsewhere (and the inclination) to replicate those operations), but this does not affect the randomness of its output when we don't know these things.

And finally, how can there be a distinction between statistical fluctuations and changes in the causal environment? If I repeat the experiment with the urn and get a different coloured ball the second time, which is that, sampling variability, or a difference in causes? Both, of course. What could sampling variability (at the macroscopic scale) be the result of, if not mechanical differences in the evolution of the experiment? If I toss a coin 4 times and get 4 heads in a row, in one sense, that's a statistical fluke, but in another, its the inevitable result of a system obeying completely deterministic mechanical laws. All that decides the level at which we find ourselves discussing the matter is our degree of awareness of the states and operations of nature.

Ok, so we live in a world polluted by some sloppy terminology, but does it really matter? I think it does. Every statistical model is an attempt to describe some physical process. As long as we systematically deny the action of physics on any aspects of these processes, then we close off access to potentially valuable physical insight. This is actually the problem that the discussed attempt to separate sampling fluctuation and causal variation was trying to address, but this half-hearted formulation serves only to postpone the problem until some later date.



Saturday, October 27, 2012

Parameter Estimation and the Relativity of Wrong




Its not enough for a theory of epistemology to consist of mathematically valid, yet totally abstract theorems. If the theory is to be taken seriously, there has to be a demonstrable correspondence between those theorems and the real world - the theory must make sense. It has to feel right.

This idea has been captured by Jaynes in his exposition of and development upon Cox's theorems1, in which the sum and product rules of probability theory (formerly, and often still, considered as axioms themselves) were rigorously derived. Jaynes built up the derivation from a small set of basic principles, which he called desiderata, rather than axioms, among which was the requirement quite simply for 'qualitative correspondence with common sense.' If you read Cox, I think it is clear that this is very much in line with his original reasoning.

Feeling right, and a strong overlap with intuition are therefore crucial tests of the validity and consistency of our theory of probability (in fact, of any theory of probability), which, let me reiterate, is really the theory of all science.  This is one of the reasons why a query I received recently from reader Yair is a great question, and a really important one, worthy of a full blown post (this one) to explore. This excellent question was about one theory being considered closer to the truth than another, and was phrased in terms of the example of the shape of the Earth: if the Earth is neither flat nor spherical, where does the idea come from that one of these hypotheses is closer to the truth? They are both false after all, within the Boolean logic of our Bayesian system. How can Bayes' theorem replicate this idea (of one false proposition being more correct than another false proposition), as any serious theory of science surely ought to?

Discussing this issue briefly in the comments following my previous post was an important lesson for me. The answer was something that I thought was obvious, but Yair's question reminded me that at some point in the past I had also considered the matter, and expended quite some effort getting to grips with it. Like so many things, it is only obvious after you have seen it. There is a story about the great mathematician, G. H. Hardy: Hardy was on stage at some conference of mathematics giving a talk. At some point, when he was saying 'it is trivial to show that.....,' he ground to a halt, stared at his notes for a moment, scratched his head, then walked off the stage absent-mindedly, into another room. Because of his greatness, and the respect the conference attendees had for him, they all waited patiently. He returned after half an hour, to say 'yes, it is trivial,' before continuing with the rest of his talk, exactly as planned.

Yair's question also reminded me of an excellent little something I read by Isaac Asimov concerning the exact same issue of the shape of the Earth, and degrees of wrongness. This piece is called 'The Relativity of Wrong2.' It consists of a reply to a correspondent who expressed the opinion that since all scientific theories are ultimately replaced by newer theories, then they are all demonstrably wrong, and since all theories are wrong, any claim of progress in science must be a fantasy. Asimov did a great job of demonstrating that this opinion is absurd, but he did not point out the specific fallacy committed by his correspondent. In a moment, I'll redress this minor shortcoming, but first, I'll give a bit of detail concerning the machinery with which Bayes' theorem sets about assessing the relative wrongness of a theory.

The technique we are concerned with is model comparison, which I have introduced already. To perform model comparison, however, we need to grasp parameter estimation, which I probably ought to have discussed in more detail before now.

Suppose we are fitting some curve through a series of measured data points, D (e.g. fitting a straight line or a circle to the outline of the Earth), then in general, our fitting model will involve some list of model parameters, which we'll call θ. If the model is represented by the proposition, M, and I represents our background information, as usual, then the probability for any given set of numerical values for the parameters, θ, is given by

(1)

If the model has only one parameter, then this is simple to interpret: θ is just a single number. If the model has two parameters, then the probability distribution P(θ) ranges over two dimensions, and is still quite easy to visualize. For more parameters, we just add more dimensions - harder to visualize, but the maths doesn't change.

The term P(θ | MI) is the prior probability for some specific value of the model parameters, our degree of belief before the data were obtained. There are various ways we could arrive at this prior, including ignorance, measured frequencies, and a previous use of Bayes' theorem.

The term P(D | θMI), known as the likelihood function, needs to be calculated from some sampling distribution. I'll describe how this is most often done. Assuming the correctness of θMI, then we know exactly the path traversed by the model curve. Very naively, we'd think that each data point, di, in D must lie on this curve, but of course, there is some measurement error involved: the d's should be close to the model curve, but will not typically lie exactly on it. Small errors will be more probable than large errors. The probability for each di, therefore, is the probability associated with the discrepancy between the data point and the expected curve, di - y(xi), where y(x) is the value of the theoretical model curve at the relevant location. This difference, di - y(xi), is called a residual.

Very often, it will be highly justified to assume a Gaussian distribution for the sampling distribution of these errors. There are two reasons for this. One is that the actual frequencies of the errors are very often well approximated as Gaussian. This is due to the overlapping of numerous physical error mechanisms, and is explained by the central limit theorem (a central theorem about limits, rather than a theorem about central limits (whatever they might be)). This also explains why Francis Galton coined the term 'normal distribution' (which we ought to prefer over 'Gaussian,' as Gauss was not the discoverer (de Moivre discovered it, and Laplace popularized it, after finding a clever alternative derivation by Gauss (note to self: use fewer nested parentheses))).

The other reason the assumption of normality is legitimate is an obscure little idea called maximum entropy. If all we know about a distribution is its location and width (mean and standard deviation), then the only function we can use to describe it, without implicitly assuming more information than we have, is the Gaussian function.

Here's what the normal sampling distribution for the error at a single data point looks like:

(2)

For all n d's in D, the total probability is just the product of all these terms given by equation (2), and since ea×eb = ea+b, then


(3)

This, along with our priors, is all we typically need to perform Bayesian parameter estimation.

If we start from ignorance, or if for any other reason, the priors are uniform, then finding the most probable values for θ simply becomes a matter of maximizing the exponential function in equation (3), and the procedure reduces to the method of maximum likelihood. Because of the minus sign in the exponent, maximizing this function requires minimizing Σ[(di - y(xi))2/2σi2]. Furthermore, if the standard deviation, σ, is the same for all d, then we just have to minimize Σ[di - y(xi)]2, which is the least squares method, beloved of physicists.

Staying, for simplicity, with the assumption of a uniform prior, then it is clear that when comparing two different fitting models, the one that achieves smaller residuals will be the favoured one, according to probability theory. (See, for example, equation (4) in my article on the Ockham Factor.) P(D | θMI) is larger for the model with smaller residuals, as just described.

The whole point of this post was to figure out how to quantify closeness to truth. The residuals we've just been looking at are how wrong the model is: d, the data point is reality, y(x) is the model, the difference between them is the amount of wrongness of the model, which we wanted to quantify. And by Bayes' theorem, more wrongness leads to less probability, exactly as desired.

Within a system of only two models, 'flat Earth' v's 'spherical Earth,' there is no scope for knowing that both models are actually false, but even working with such a system, we would probably keep in mind the strong potential for a third, more accurate model  (e.g. the oblate spheroid that Asimov discussed). Such mindfulness is really a manifestation of a broader 'supermodel.' In the two-model system, 'spherical Earth' is closer to the truth because it manifests much smaller residuals. It is also closer to the truth than 'flat Earth,' even after the third model is introduced, because its residuals are still smaller than those for 'flat Earth.' 'Oblate spheroid' will be even closer to the truth in the 3 theory system, but spherical and flat will still have non-zero probability - strictly, we can not rule them out completely, thanks to the unavoidable measurement uncertainty, and so the statement that we know them to be false is not rigorously valid.

I promised earlier to identify the fallacy perpetrated by Asimov's misguided correspondent. I have already discussed it a few months ago. It is the mind-projection fallacy, the false assumption that aspects of our model of reality must be manifested in reality itself. In this case: if wrongness (relating to our knowledge of reality) can be graded, then so must 'true' and 'false' (relating to reality itself) also be graded. There are two ways to reason from here: (1) truth must be fuzzy, or (2) our idea of continuous degrees of wrong must be mistaken.

The idea that all models that are wrong are necessarily all equally wrong, as expressed in the letter with which poor Asimov was confronted, is fallacious in the extreme. Wrong does not have this black/white feature. 'Wrong' and 'false' are not the same. Of course, a wrong theory is also false, but if I'm walking to the shop, I'd rather find my location to be wrong by half a mile than by a hundred miles.

We can say that a theory is less wrong (i.e. produces smaller residuals), without implying that is is more true. 'True' and 'false' retain their black-and-white character, as I believe they must, but our knowledge of what is true is necessarily fuzzy. This is precisely why we use probabilities. As our theories get incrementally less wrong and closer to the truth, so the probabilities we are allowed to assign to them get larger.

There often seems to be a kind of bait-and-switch con trick going on with many of the world's least respectable 'philosophies.' The 'philosopher' makes a trivial but correct observation, then makes a subtle shift, often via the mind-projection fallacy, to produce an equivalent-looking statement that is both revolutionary, and utter garbage. In post-modern relativism (a popular movement in certain circles), we can see this shifting between right/wrong and true/false. The observation is made that all scientific theories are ultimately wrong, then hoping you won't notice the switch, the next thing you hear is that all theories are equally wrong. They can't seem to make their minds up, however, which side of the fallacy they are on, because the next thing you'll probably hear from them is that because knowledge is mutable, then so are the facts themselves: truth is relative to your point of view and the mood you happen to be in, science is nought but a social construct. Part of the joy of familiarity with scientific reasoning is the clarity of thought to see through the fog of such nonsense.







[1] 'Probability, Frequency, and Reasonable Expectation,' R. T. Cox, American Journal of Physics 1946, Vol. 14, No. 1, Pages 1-13. (Available here.)


[2] 'The Relativity of Wrong,' Isaac Asimov, The Skeptical Inquirer, Fall 1989, Vol. 14, No. 1, Pages 35-44. (Download the text here.) (And no, I don't find it spooky that both references are from the same volume and number.)



Friday, June 29, 2012

The Mind Projection Fallacy


There is an argument that I have noted a couple of times when scientific colleagues with religious beliefs have tried to explain to me how they reconcile these seemingly contradictory things. Reality, they say, is divided into two classes of phenomena: the natural and the supernatural. Natural phenomena, they say, are the things that fall into the scope of science, while the supernatural lies outside of science’s grasp, and can not be addressed by rational investigation. This is completely muddle-headed, and seems to me to be based on an example of something called the mind projection fallacy.

A similar argument also crops up occasionally when advocates of alternative medicine try to rationalize the complete failure of their favorite pseudoscientific therapy to provide any evidence of efficacy in rigorous trials.

The very word ‘supernatural,’ at its heart, though, is one of those utterly self-defeating terms, like ‘free will’ and ‘alternative medicine,’ completely devoid of meaning and philosophically bankrupt. What is this free will that people keep going on about? Is it the freedom to break the laws of physics? No, and since every particle in your brain obeys the laws of physics, you are not free to make non-mechanistic decisions, so can we please shut up about free will? (Granted, I am using a restricted meaning of the term ‘free will.’)

And what is alternative medicine? Medicine is the use of interventions that are known to work in order to lessen the effects of disease. If it doesn’t work, or is not known to work, then its not medicine, full stop. There is no alternative. Lets please shut about alternative medicine.

What is the supernatural? Nature is by definition everything that exists and happens. What is outside nature is therefore necessarily an empty set.

The etymology of the word ‘supernatural’ is the result of an error of thinking. This error is the mind projection fallacy: falsely assuming that the properties of one’s model of reality necessarily exhibit correspondence with the actual properties of reality. The following dictionary definition of ‘supernatural’ was quoted to me in a recent discussion of the term (reportedly from Webster’s):

Supernatural. [adjective:] 1. of, pertaining to, or being ‘above or beyond what is natural or explainable by natural law’. 2. of, pertaining to, or attributed to God or a deity. 3. of a superlative degree; preternatural. 4. pertaining to or attributed to ghosts, goblins, or other unearthly beings; eerie; occult.

Number 1, is where we have to focus. Numbers 2 and 4 are, in origin at least, based on erroneous application of number 1, while number 3 is just weird. ‘Of a superlative degree’? That’s not supernatural, by any reasonable standard. ‘Preturnatural’? This word has two meanings (according to Dictionary.com): one is ‘supernatural’ (wow, that’s helpful) and the other is ‘exceptional or abnormal.’ Finding a hundred euros on the pavement would be both exceptional and abnormal, but again, not supernatural unless we are willing to debase the meanings of words to the level of uselessness.

So what about the primary meaning of supernatural, ‘above or beyond what is natural or explainable by natural law.’ The first part poses a problem, since there is no supplied procedure for determining what is natural, other than the obvious definition: ‘whatever is not supernatural.’ Now I’m aware that all word definitions are ultimately circular, but this is a case where the radius of curvature is clearly far to small to represent any useful addition to the language. The second part stipulates ‘explainable by natural law,’ which succumbs to exactly the same objection, but I strongly suspect that many people have failed to see this exactly because they have committed the mind projection fallacy – in this case, conflation of natural law with our description of it. Natural law is the set of principles, whatever they may be, that determine how real phenomena evolve. If a phenomenon is real, then it would be explainable by natural law. If a phenomenon is not real, then what is the point in debating whether or not it is supernatural? I feel, however, that too many people think that natural law is some set of equations, like E = mc2, written down in text books – but this is merely our model of natural law. I see no other convincing way to account for the appearance of this phrase in the quoted dictionary definition, than to assume that natural law is being commonly confused with known science in this way, since there seems to be no other good reason to postulate that a phenomenon is not governed by natural law (I know, the exact word was ‘explainable,’ but I think it is hard to rescue the situation by invoking this subtle difference).

By this common understanding of ‘supernatural,’ the photoelectric effect would have been supernatural in and prior to 1904, but natural before the end of 1905. A strange state of affairs, you might think.

Of course, word meanings don’t have to stick exactly to their original literal meanings, and anybody is free to apply the word ‘supernatural’ to any putative phenomenon they wish: gods, ghosts, whatever (as long as they are clear in what they are doing), but I argue firstly, that this is a misnomer, as nothing can be literally beyond nature (supernatural) and secondly, that use of this misguided word leads to horrendous confusions, such as those allowing highly educated and otherwise rational people to claim that religious phenomena (or homeopathy or chi) are by definition supernatural, and therefore by definition beyond the scrutiny of science.

The mind projection fallacy also raises its head in science, all too often, such as in quantum physics, and, in my opinion, in thermodynamics. It has also had very substantial consequences in the development and application of probability theory. Since scientific method generally strives to avoid fallacious reasoning, I feel that it is important to get well acquainted with this particular mental glitch, and to recognize some of the fields in which it still extends a corrupting influence.

Looking at thermodynamics, the famous second law, stating that the entropy of a closed system tends to increase, is often explained by the experts as resulting from our inability to distinguish between individual molecules (or other particles). My conviction, however, is that the mechanical evolution of an ensemble of such particles is unchanged if we are granted a means to identify them after they have evolved. The real reason for the second law seems to be that the proportion of possible initial microstates that result in non-increased entropy is very tiny (and such states appear, therefore with vanishing probability), but polluted with the standard language of the discipline, one can find it hard to grasp this. Why is this standard language an example of the mind projection fallacy? Because there is something unknown to us (the identities of the particles), and the fact of it being unknown is attributed as the cause of the physical evolution of the system, and therefore a physical property of the system. It is not, though, it is a property of our knowledge of the system.

With regard to quantum mechanics, I am open minded on the matter of whether or not nature evolves deterministically or non-deterministically. As I try to be a good scientist, I wait for the evidence to favour one hypothesis strongly over the other before casting my judgement. As much as we know about quantum mechanics already, that evidence is not yet in. I am, however, highly uncomfortable and skeptical about the possibility of something operating without a causal mechanism, yet exhibiting clear tendencies. It seems I am not alone in this, as several other well regarded thinkers have apparently shared this view, notably among them, the exceptional theoretical physicists Albert Einstein, Louis de Broglie, David Bohm, John Bell, and Edwin Jaynes. Imagine my delight, therefore, when I discovered the following passage by Jaynes in a book of conference proceedings1, articulating magnificently (and far better than I ever could) many of my own long-felt misgivings about the language of quantum mechanics:

The current literature on quantum theory is saturated with the Mind Projection Fallacy. Many of us were first told, as undergraduates, about Bose and Fermi statistics by an argument like this: “You and I cannot distinguish between the particles; therefore the particles behave differently than if we could.” Or the mysteries of the uncertainty principle were explained to us thus: “The momentum of the particle is unknown; therefore it has a high kinetic energy.” A standard of logic that would be considered a psychiatric disorder in other fields, is the accepted norm in quantum theory. But this is really a form of arrogance, as if one were claiming to control Nature by psychokinesis.

Whether or not the position and momentum of a particle (as related in the most familiar version of the Heisenberg principle) are truly ‘undetermined’ or merely unknowable to us, I am unsure, but there is a commonly encountered assumption that these two possibilities must be the same, and this results from the mind projection fallacy. It is indeed a mighty challenge to reconcile quantum phenomena with a fully deterministic mechanics. Some have succeeded, but it remains a challenge to pin down whether or not this is nature’s way. Many, however, follow Bohr and assert that there can be no underlying mechanism behind quantum phenomena. Let me quote Jaynes again (this time from ‘Probability theory: the logic of science’):

the ‘central dogma’ [of quantum theory]… draws the conclusion that belief in causes, and searching for them, is philosophically naïve, If everybody accepted this and abided by it, no further advances in understanding of physical law would ever be made… it seems to us that this attitude places a premium on stupidity.


The field in which the mind projection fallacy has had its most significant practical consequences is perhaps probability theory, which is a colossal shame, as probability is the king of theories: the meta-theory that decides how all other theories are obtained.

If you scan through the articles I have posted here on probability, you’ll observe that most if not all make use of Bayes’ theorem. It is an incredibly important and useful part of statistical reasoning, and represents the core of how human knowledge advances. It is also derived simply, as a trivial rearrangement of two of the most basic principles of probability theory: the product and sum rules. Yet, for a significant portion of the 20th century, when statistical theory was undergoing explosive development, Bayes’ theorem was rejected by the majority of authorities and practitioners in the field. How on Earth could this have come about? The mind projection fallacy, of course.

Because the theory models real phenomena in terms of probabilities, it was assumed that these probabilities must be real properties of the phenomena. Yet Bayes’ theorem converts a prior probability into a posterior probability by the addition of mere information. And since merely changing the amount of information cannot affect the physical properties of a system, then Bayes’ theorem must simply be wrong. QED.

The property that probability was thought to correspond to was frequency. For example, a coin has a 50% probability to land heads up because the relative frequency with which it does so is one half. For this reason, the orthodox school of statistical thought has become known as frequentist statistics.

One of the most extraordinary scientists of the 20th century, Ronald Fisher, for example, was one of the people who dominated the development of statistical theory during his lifetime. In his highly influential book, ‘The design of experiments,’2 he gave three reasons for rejecting Bayes’ theorem, foremost of which is:

… advocates of inverse probability [Bayes’ theorem] seem forced to regard probability not as an objective quantity measured by observable frequencies….

Clearly, he meant that the impossibility to reconcile Bayes’ theorem with the view of probability as a physical property of real objects (the frequencies with which different events occur) made it impossible to accept the theorem. (His other two reasons are just as bad.) It was Fisher’s deeply held objection to the logical foundations of probability theory that led him to do some of the most important work developing and popularizing the frequentist significance tests, which, as I have argued in detail here and here, are a poor method for assessing data. 

Another influential textbook, by Harald Cramér3, asserts that ‘any random variable has a unique probability distribution.’ Again, assuming that the probability is something objective and immutable, a physical property. The randomness is assumed to be necessarily a property of the system under study, rather than a statement of our lack of information – our inability to predict it before hand. To instantly recognize the ridiculousness of both Fisher’s and Cramér’s views, consider that I have just tossed a coin, which has landed, and I am asking you to assess the probability that the face of the coin pointing up is the one depicting the head: your only rational answer is 0.5, and it is the correct answer, for you. For me though, the correct answer is 1, because I am looking at the coin, and I can see the head facing up. Same physical system, different probabilities, dependent on the available information.

On the subject of probabilities as physical properties of the systems we study, I can again quote Jaynes, who has summarized the situation beautifully:

It is therefore illogical to speak of verifying [the Bernoulli urn rule, a law for determining probabilities] by performing experiments with the urn; that would be like trying to verify a boy’s love for his dog by performing experiments on the dog.

We can easily identify other instances of the mind projection fallacy in probability reasoning, some of which I have already discussed in earlier posts. For example, the error of thinking discussed in Logical v’s Causal Dependence, consisting of the belief that the expression P(A|B) can only be different from P(A) if B exerts a causal effect on A (an error that has made it into a number of influential textbooks on statistical mechanics) seems to arise from the conviction that a probability is an objective property of the system under study. If B changes the probability for A, then according to this belief, B changes the physical properties of A, and must therefore be, at least partially, the cause of A.

Another instance is to be found in The Raven Paradox, and consists of the belief that whether or not a particular piece of evidence supports a hypothesis is an objective property of the hypothesis, or the real system to which the hypothesis relates. In that post, we examined the supposition that observation of a sequence of exclusively black ravens supports the hypothesis that all ravens are black. We discovered an instance where such observations actually support the opposite hypothesis, illustrating that the relationship between the hypothesis and the data is entirely dependent on the model we chose. To think otherwise was shown to lead to disturbing and indefensible conclusions about ravens. 





[1] 'Maximum Entropy and Bayesian Methods,' edited by J. Skilling, Kluwer Publishing, 1989

[2] 'The Design of Experiements,' R.A. Fisher, Oliver and Boyd, 1935

[3] 'Mathematical Methods of Statistics,' H. Cramér, Princeton University Press, 1946