Friday, September 13, 2013

Is Rationality Desirable?




Seriously, why all this fuss about rationality and science, and all that? Can we be just as happy, or even more so, being irrational, as by being rational? Are there aspects of our lives where rationality doesn't help? Might rationality actually be a danger?

Think for a moment about what it means to desire something.

To desire something trivially entails desiring an efficient means to attain it. To desire X is to expect my life to be better if I add X to my possessions. To desire X, and not to avail of a known opportunity to increase my probability to add X to my possessions, therefore, is either (1) to do something counter to my desires, or (2) to desire my life not to get better. Number (2) is strictly impossible – a better life for me is, by definition, one in which more of my desires are fulfilled. Number (1) is incoherent – there can be no motivation for anyone to do anything against their own interests. Behaviour mode (1) is not impossible, but it can only be the result of a malfunction.

Let’s consider some complicating circumstances to check the robustness of this.
  1. Suppose I desire a cigarette. Not to smoke a cigarette, however, is clearly in my interests. There is no contradiction here. Besides (hypothetically) wanting to smoke something, I also have other goals, such as a long healthy life, which are of greater importance to me. To desire a cigarette is to be aware of part of my mind that mistakenly thinks this will make my life better, even though in expectation, it will not. This is not really an example of desiring what I do not desire, because a few puffs of nicotine is not my highest desire – when all desires that can be compared on the same dimension are accounted for, the net outcome is what counts. Neither is it an example of acting against my desires if I turn down the offer of a smoke, for the same reason.

  2. Suppose I desire to reach the top of a mountain, but I refuse to take the cable car that conveniently departs every 30 minutes, preferring instead to scale the steep and difficult cliffs by hand and foot. Simplistically, this looks like genuinely desiring to not avail of an efficient means to attain my desires, but in reality, it is clearly the case that reaching the summit is only part of the goal, another part being the pleasure derived from the challenging method of getting there.   

Despite complications arising from the inner structure of our desires, therefore, for me to knowingly refuse to adopt behaviour that would increase my probability to fulfill my desires is undeniably undesirable. Now, behavior that we know increases our chances to get what we desire has certain general features. For example, it requires an ability to accumulate reliable information about the world. It is not satisfactory to take a wild guess at the best course of action, and just hope that it works. This might work, but it will not work reliably. My rational expectation to achieve my goal is no better than if I do nothing. Reliability begins to enter the picture when I can make informed guesses. I must be able to make reliable predictions about what will happen as a result of my actions, and to make these predictions, I need a model of reality with some fidelity. Not just fidelity, but known fidelity - to increase the probability to achieve my goals, I need a strategy that I have good reasons to trust.

It happens that there is a procedure capable of supplying the kinds of reliable information and models of reality that enable the kinds of predictions we desire to make, in the pursuit of our desires. Furthermore, we all know what it is. It is called scientific method. Remember the reliability criterion? This is what makes science scientific. The gold standard for assessing the reliability of a proposition about the real world is probability theory – a kind of reasoning from empirical experience. Thus the ability of science to say anything worthwhile about the structure of reality comes from its application of probability theory or any of several approximations that are demonstrably good in certain special cases. If there is something that is better than today’s science, then better is the result of a favorable outcome under probabilistic analysis (since 'better' implies 'reliably better'), thus, whatever it is, it is tomorrow’s science.

So, if I desire a thing, then I desire a means to maximize my expectation to get it, so I desire a means to make reliable predictions of the outcomes of my actions, meaning that I desire a model of the world in which I can justifiably invest a high level of belief, thus I desire to employ scientific method, the set of procedures best qualified to identify reliable propositions about reality. Therefore, rationality is desirable. Full stop.

We cannot expect to be as happy by being irrational as by being rational. We might be lucky, but by definition, we cannot rely on luck, and our desires entail also desiring reliable strategies.

Items (A) to (D), below, detail some subtleties related to these conclusions.


(A) Where’s the fun in that?

Seriously? Being rational is always desirable? Seems like an awfully dry, humorless existence, always having to consult a set of equations before deciding what to do!

What this objection amounts to is another example (ii), from above, where the climber chooses to take the difficult route to the top of the mountain. What is really meant by a dry existence is something like elimination of pleasant surprises, spontaneity, and ad-hoc creativity, and that these things are actually part of what we value.

Of course, there are also unpleasant surprises possible, and we do value minimizing those. The capacity to increase the frequency of pleasant surprises, while not dangerously exposing ourselves to trouble is something that, of course, is best delivered through being rational. Being in a contained way irrational may be one of our goals, but as always, the best way to achieve this is by being rational about it. (I won’t have much opportunity to continue my pursuit of irrationality tomorrow, if I die recklessly today.)   

(B) Sophistication effect

To be rational (and thus make maximal use of scientific method, as required by a coherent pursuit of our desires) means to make study of likely failure modes of human reasoning (if you are human). This reduces the probability of committing fallacies of reasoning yourself, thus increasing the probability that your model of reality is correct. But, there is a recognized failure mode of human reasoning that actually results from increased awareness of failure modes of reasoning. It goes like this: knowing many of the mechanisms by which seemingly intelligent people can be misled by their own flawed heuristic reasoning methods makes it easy for me for hypothesize reasons to ignore good evidence, when it supports a proposition that I don’t like – “Oh sure, he says he has seen 20 cases of X and no cases of Y, but that’s probably a confirmation bias.”

Does this undermine my argument? Not at all. This is not really a danger of rationality. If anything, it is a danger of education (though one that I confidently predict that a rational analysis will reveal to be not sufficient to argue for reduced education). What has happened, in the above example is of course itself a form of flawed reasoning, it is reasoning based on what I desire to be true, and thus isn't rational. It may be a pursuit of rationality that led me to reason in this way, but this is only because my quest has been (hopefully temporarily) derailed. Thus my desire to be rational (entailed trivially by my possession of desire for anything) makes it often desirable for me to have the support of like-minded rational people, capable of pointing out the error, when even the honest quest for reliable information leads me into a trap of fallacious inference.

(C) Where does it stop?

The assessment of probability is open ended. If there is anything about probability theory that sucks, this is it, but no matter how brilliant the minds that come to work on this problem, no way around it can ever be found, in principle. It is just something we have to live with - pretending it's not there won't make it go away. What it means, though, is that no probability can be divorced from the model within which it is calculated. There is always a possibility that my hypothesis space does not contain a true hypothesis. For example, I can use probability theory to determine the most likely coefficients, A and B, in a linear model used to fit some data, but investigation of the linear model will say nothing about other possible fitting functions. I can repeat a similar analysis using say a three-parameter quadratic fit, and then decide which fitting model is the most likely using Ockham’s razor, but then what about some third candidate? Or what if the Gaussian noise model I used in my assessment of the fits is wrong? What if I suspect that some of the measurements in my data set are flawed? Perhaps the whole experiment was just a dream. These things can all be checked in essentially the same way as all the previously considered possibilities (using probability theory), but it is quite clear that the process can continue indefinitely.

Rationality is thus a slippery concept: how much does it take to be rational? Since the underlying procedure of rationality, the calculation of probabilities, can always be improved by adding another level, won’t it go on forever, precluding the possibility to ever reach a decision?

To answer this, let us note that to execute a calculation capable of deciding how to achieve maximal happiness and prosperity for all of humanity and all other life on Earth is not a rational thing to do if the calculation is so costly that its completion results in the immediate extinction of all humanity and all other life on Earth.

Rationality is necessarily a reflexive process, both (as described above) in that it requires analysis of the potential failure modes of the particular hardware/software combination being utilized (awareness of cognitive biases), and in that it must try to monitor its own cost. Recall that rationality owes its ultimate justification to the fulfillment of desires. These desires necessarily supersede the desire to be rational itself. An algorithm designed to do nothing other than be rational would do literally nothing - so without a higher goal above it, rationality is literally nothing.

Thus, if the cost of the chosen rational procedure is expected to prevent the necessarily higher-level desire being fulfilled, then rationality dictates that the calculation be stopped (or better, not started). Furthermore, the (necessary) desire to employ a procedure that doesn't diminish the likelihood to achieve the highest goals entails a procedure capable of assessing and flagging when such an occurrence is likely.

(D) Going with your gut feeling

On a related issue, concerning again the contingency (the lack of guarantee that the hypothesis space actually contains a true hypothesis) and potential difficulty of a rational calculation, do we need to worry that the possible computational difficulty and, ultimately, the possibility that we will be wrong in the end will make rationality uncompetitive with our innate capabilities of judgment? Only in a very limited sense.

Yes, we have superbly adapted computational organs, with efficiencies far exceeding any artificial hardware that we can so far devise, and capable of solving problems vastly more difficult than any rigorous probability-crunching machine that we can now build. And yes, it probably is rational under many circumstances to favor the rough and ready output of somebody’s bias-ridden squishy brain over the hassle of a near-impossible, but oh-so rigorous calculation. But under what circumstances? Either, as noted, when the cost of the calculation prohibits the attainment of the ultimate goal, or when rationally evaluated empirical evidence indicates that it is probably safe to do so.

Human brain function is at least partially rational, after all. Our brains are adapted for and (I am highly justified in believing) quite successful at making self-serving judgments, which, as noted, is founded upon an ability to form a reliable impression of the workings of our environment. And, as also noted, the degree of rigor called for in any rational calculation is determined by the costs of the possible calculations, the costs of not doing the calculations, and the amount we expect to gain from them.

This is not to downplay the importance of scientific method. Let me emphasize: a reliable estimate of when it is acceptable to rely on heuristics, rather than full-blown analysis, can only come from a rational procedure. The list of known cognitive biases that interfere with sound reasoning is unfortunately rather extensive, and presumably still growing. The science informs us that rather often, our innate judgement exhibits significantly less success than rational procedure. 



Sunday, July 28, 2013

Forward Problems




Recently, I unveiled a collection of mathematical material, intended partly as an easy entry point for people interested in learning probability from scratch. One thing that struck me, though, as conspicuously missing was a set of simple examples of basic forward problems. To correct this, the following is a brief tutorial, illustrating application of some of the most basic elements of probability theory.

Virtually anybody who has studied probability at school has studied forward problems. Many, though, will have never learned about inverse problems, and so the term 'forward problem' is often never even introduced. What are referred to as forward problems are what many identify as the entirety of probability theory, but in reality, this is only half of it. 

Forward problems start from specification of the mechanics of some experiment, and work from there to calculate the probabilities with which the experiment will have certain outcomes. A simple example is: a deck of cards has 52 cards, including exactly 4 aces, what is the probability to obtain an ace on a single draw from the deck? The name for the discipline covering these kinds of problems is sampling theory.

Inverse problems, on the other hand, turn this situation around: the outcome of the experiment is known, and the task is to tease out the likely mechanics that lead to that outcome. Inverse problems are the primary occupation of scientists. They are of great interest to everybody else, as well, whether we realize it or not, as all understanding about the real world results from analysing inverse problems. Solution of inverse problems, though, is tackled using Bayes' theorem, and therefore also requires the implementation of sampling theory.

To answer the example above (drawing an ace from a full deck), we need only one of the basic laws: the Bernoulli urn rule. Due to the lack of any information to the contrary, symmetry requires that each card is equally probable to be drawn on a single draw, and you can quickly verify that the answer is P(Ace) = 1/13. Here is another example, requiring mildly greater thought: 


The Birthday Paradox

Not really a paradox, but supposedly many find the result surprising. The problem is the following:

What is the minimum number of people needed to be gathered together to ensure that the probability that at least two of them share a birthday exceeds one half?

It doesn't look immediately like my definition of a forward problem, but actually it is several - to get the solution, we'll calculate the desired probability for all possible numbers of people, until we pass the 50% requirement. A 'birthday' is a date consisting of a day and a month, like January 1st - the year is unimportant. Assume also that each year has 365 days. To maximize our exposure to as many basic rules as possible, I'll demonstrate 2 equivalent solutions.

Solution 1

Having at least 1 shared birthday in a group of people can be fulfilled in several possible ways. There may be exactly 2 matching birthdays, or exactly 3 matching, or exactly 4, and so on. We don't want to calculate all these probabilities. Instead, we can invoke the sum rule to observe that


P(at least one shared)  =  1 - P(none shared) 
(1)

To get P(none shared), we'll use the Bernoulli urn rule again. If the number of people in the group is n, then we want the number of ways to have n different dates, divided by the number of ways to have any n dates. To have n different dates, we must select n objects from a pool of 365, where each object can be selected once at most. Thus we need the number of permutations, 365Pn, which, from the linked glossary entry, is


 
(2)

That's our numerator, the denominator is the total number of possible lists of n birthdays. Well, there are 365 ways to chose the first, and 365 ways to chose the second, thus 365 × 365 ways to chose the first two. Similarly 365 × 365 × 365 ways to chose the first three, thus 365n ways to chose n dates that may or may not be different. Taking the ratio, then


 
(3)

Thus, when we apply the sum rule, the desired expression is


 
(4)

And now we only need to crunch the numbers for different values for n, to find out where P first exceeds 50%. To calculate this, I've written a short python module, which is in the appendix, below (a spreadsheet program can also do this kind of calculation). Python is open-source, so in principle anybody with a computer and an internet connection (e.g. You) can use it. The graph below shows the output for n equal to 1 to 60:

The answer, contrary to Douglas Adams, is 23.


Solution 2

The second method is equivalent to the first, it still employs that trick in equation (1) with the sum rule, and exactly the same formula, equation (4), is derived, but this time we'll use the sum rule, the product rule, and the logical independence of different people's birthdays.

To visualize this solution more easily, lets imagine people entering a room, one by one, and checking each time whether or not there is a shared birthday. When the first person enters the room, the probability there is no shared birthday in the room is 1. When the second person enters the room, the probability that he shares his birthday with the person already there is 1/365,so the probability that he does not share the other person's birthday, from the sum rule, is (1 - 1/365).

The probability for no shared birthdays, with n = 2, is given by the product rule:


P(AB | I) = P(A | I)P(B | AI) 
(5)

We can simplify this, though. Logical independence means that if we know proposition A to be true, the probability associated with proposition B is the same as if we know nothing about A: P(B | I) = P(B | AI). Thus the probability that person B's birthday is on a particular date is not changed by any assumptions about person A's birthday.

So the product rule reduces in this case to


P(AB | I) = P(A | I)P(B | I)   
(6)

Then the probability of no matching birthdays with n = 2 is 1 × (1-1/365). For n = 3, we repeat the same procedure, getting as our answer 1 × (1-1/365) × (1-2/365).

For n people, therefore, we get


which re-arranges to give the same result as before:


 
(7)

As before, we subtract this from 1 to get the probability to have at least 1 match among the n people present.








Appendix



import numpy as np                                                      # package for numerical calculations
import matplotlib.pyplot as plt                                     # plotting package

def permut(n, k):                                                             # count the number of permutations, nPk
    answer = 1
    for i in np.arange(n, n-k, -1):
        answer *= i
    return answer


def share_prob(n):                                                           # calculate P(some_shared_bDay | n)
    p = []
    for i in n:
        p.append(1.0 - permut(365, i)/np.power(365.0, i))
    p = np.array(p)
    return p

def plot_prob():                                                                 # plot distribution & find 50% point
    x = np.arange(1.0, 61.0, 1.0)                                         # list the integers, n, from 1 to 60 inclusive
    y = share_prob(x)                                                          # get P(share|n)

    x50 = x[y>0.5][0]                                                          # calculate x, y coordinates where 50%
    y50 = y[y>0.5][0]                                                          # threshold first crossed

    fig = plt.figure()                                                              # create plot
    ax = fig.add_subplot(111)
    ax.plot(x, y, linewidth=2)

    # mark out the solution on the plot:
    ax.plot([x50, x50], [0, y50], '-k', linewidth=1)
    ax.plot([0, x50], [y50, y50], '-k', linewidth=1)
    ax.plot(x50, y50, 'or', markeredgecolor='r', markersize=10)
    ax.text(x50*1.1, y50*1.05, 'Smallest n giving P > 0.5:', fontsize=14)
    ax.text(x50*1.1, y50*0.95, '(%d, %.4f)' %(x50, y50), fontsize=14)

   # add some other information:
    ax.set_title('P(some_shared_birthday | n)', fontsize=15)
    ax.set_ylabel('Probability', fontsize=14)
    ax.set_xlabel('Number of people, n', fontsize=14)
    ax.axis([1, 60, 0, 1.05])





Saturday, July 20, 2013

Greatness, By Definition



The goals for this blog have always been two-fold: (1) to bring students and professionals in the sciences into close acquaintance with centrally important topics in Bayesian statistics and rational scientific method, and (2) to bring the universal scope and beauty of science to the awareness of as many as possible, both within and outside the field - if you have any kind of problem in the real world, then science is the tool for you.

Effective communication to scientists, though, runs the risk of being impenetrable for non-scientists, while my efforts to simplify make me feel that the mathematically adept reader will be quickly bored.

Helping to make the material on the blog more accessible, therefore, and as part of a very, very slow but steady plan to achieve world domination, here are two new resources I've put together, which I am very happy to announce:

  • Glossary - definitions of technical terms used on the blog
  • Mathematical Resource - a set of links explaining the basics of probability from the beginning. Blog articles and glossary entries are linked in a logical order, starting from zero assumed prior expertise.

Both of these now appear in the links list on the right-hand sidebar.

The new resources are partially complete. Some of the names of entries in the glossary, for example, do not yet correspond to existing entries. Regular updates are planned.

The mathematical resource is a near-instantaneous extension of the material compiled for the glossary, and is actually my glacially slow response to a highly useful suggestion made by Richard Carrier, almost one year ago. The material has been organized in what seems to me to be a logical order, and for those interested, may be viewed as a short course in statistics, delivering, I hope, real practical skills in the topic. Its main purpose, though, is to provide an entry point for those interested in the blog, but unfamiliar with some of the important technical concepts.

The glossary may also be useful to those already familiar with the topics. Terms are used on the blog, for example, with meanings different to those of many other authors. Hypothesis testing is one such case, limited by some to denoting frequentist tests of significance, but used here to refer more generally to any  ranking of the reliability of propositions. 

The new glossary, then, is an attempt to rationalize the terminology, bringing the vocabulary back in line with what it was always intended to mean, not to reflect some flawed surrogates for those original intentions. For the same reason, in some cases alternate terms are used, such as 'falsifiability principle', in preference to the more common 'falsification principle'.

Important distinctions are also highlighted. Morality, for example is differentiated from moral fact. Philosophers are found to be distinct from 'nominal philosophers'. Equally importantly, science is explicitly distanced from 'the activity of scientists'. As a result, morality, philosophy, and science are found to be different words for exactly the same phenomenon.

In a previous article, I warned against excessive reliance on jargon, so it's perhaps worth explaining how the current initiative is not hypocrisy. That article was concerned with over use of unnecessary jargon, which often serves as an impediment to understanding. Symptoms of this include (1) jargon terms replaced with direct (but less familiar) synonyms result in confusion, and (2) vocabulary replaced with familiar terms with inapplicable meanings goes unnoticed. By providing a precise lexicon, we can help to prevent exactly these problems, and others, thus whetting the edge of our analytical blade, and accelerating our philosophical progress.

On a closely related topic, there is a fashionable notion going along the lines that to argue from the definition of words is fallacious, as if definitions of words are useless. This is not correct: argument from definition is a valid form of reasoning, but one that is very commonly misused.

Eliezer Yudkowsky's highly recommendable sequence, A Human's Guide to Words, covers the fallacious application very well. His principal example is the ancient riddle: if a tree falls in a forest, where nobody is present to hear it, does it make a sound? Yudkowsky imagines two people arguing over this riddle. One asserts, "yes, by definition: acoustic vibrations travel through the air," the other responds, "no, by definition: no auditory sensation occurs in anybody's brain." These two are clearly applying different definitions to the same word. In order to reach consensus, they must agree on a single definition.

These two haven't committed any fallacy yet, each is reasoning correctly from their own definitions. But it is a pointless argument - as pointless as me arguing in English, when I only understand English, with a person who only speaks and understands Japanese. Fallacy begins, however, as in the following example.

Suppose you and I both understand and agree on what the word 'rainbow' refers to. One day, though, I'm writing a dictionary, and under 'rainbow,' I include the innocent looking phrase: "occurs shortly after rain." (Well duh, rain is even in the name.) So we go visit a big waterfall and see colours in the spray, and I say "look, it must have recently rained here." Con artists term this tactic 'bait and switch.' I can not legitimately reason in this way, because I have arbitrarily attached not a symbol to a meaning, but attributes to a real physical object. 

To show trivially that there is a valid form of argument from definition, though, consider the following truism: "black things are black." This is necessarily true, because blackness is exactly the property I'm talking about when I invoke the phrase "black things." It is not that I am hoping to alter the contents of reality by asserting the necessary truth of the statement, but that I am referring to a particular class of entities, and the entities I have in mind just happen to all be black - by definition. 

One might complain, "but I prefer to use the phrase 'black things' not just for things that are black, but also for things that are nearly black." This would certainly be perverse, but it's not in any sense I can think of illegal. Fine, if you want to use the term that way, you may do so. I'll continue to implement my definition, which I find to be the most reasonable, and every time you hear me say the words "black things," you must replace the words with any symbol you like that conveys the required meaning to you. Your symbol might be a sequence of 47 charcoal 'z's marked on papyrus, or a mental image of a yak, I don't care. 

Yes, our definitions are arbitrary, but arbitrary in the sense that there is no prior privileged status of the symbols we end up using, and not in the sense that the meanings we attach to those symbols are unimportant.

Here's an example from my own experience. Several times, I have tried unsuccessfully to explain to people my discovery that ethics is a scientific discipline. (By the way I'm not claiming priority for this discovery.) The objections typically go through 3 phases. First is the feeling of hand-waviness, which is understandable, given how ridiculously simple the argument is: 

them- No way, its too simple. You can't possibly claim such an unexpected result with such a thin argument.
me- OK, show me which details I've glossed over.
them- [pause...] All right, the argument looks logically sound, but I don't believe it - look at your axioms: why should I accept those? Those aren't the axioms I choose.
me- Those aren't axioms at all. I don't need to assume their truth. These are basic statements that are true by definition. If you don't like the words I've attached to those definitions, then pick you own words, I'm happy to accommodate them.
them- YOU CAN'T DO THAT! You're trying to alter reality by your choice of definitions....


And that's the final stumbling block people seem to have the biggest trouble getting over.

Definitions are important. If you think that making definitions is a bogus attempt to alter reality, then be true to your beliefs: see how much intellectual progress you can make without assigning meanings to words. The new <fanfare!> Maximum Entropy Glossary </fanfare!> is an attempt to streamline intellectual progress. If you engage with anything I have written on this blog, then you engage with meanings I have attached to strings of typed characters. In some important cases, I have tried to make those meanings clear and precise. If you find yourself disagreeing with strings of characters, then you are making a mistake. If you disagree with the way I manipulate meanings then we can discuss it like adults, confident that we are talking about the same things.



Tuesday, June 25, 2013

Crime and Punishment




There’s been an idea circulating for some time that retributive justice is morally and logically founded upon the fact that we possess a thing called free will - some assumed weird mechanism that disconnects human behaviour from the normal cause-and-effect based evolution of nature. If, after all, human actions were really ‘just’ the result of mechanistic microscopic processes, then whatever we do would be entirely determined by the laws of physics and the configuration of our environment. And if this were really so, then whatever somebody does is a consequence of the fact that they could not have willfully done otherwise, in which case there is no sense in which a person can be blamed for doing wrong. And if culpability can not be established, then doesn't the validity of punishment look suspect? So prevalent is this idea that it forms a major part of contemporary legal philosophy.

Not only is this idea of free will completely nonsensical, but the connection between it and the justice of retribution is totally unfounded. Vengeance, after all is really just an expression of anger. Is anger rational? Is it a reliable, systematic producer of well judged behaviour? Or is it merely a crude and ancient heuristic moderator of human interaction that in a modern, enlightened era, we could do with much less of?

There is simply no logical link between culpability and the righteousness of retributive punishment,  which somehow ‘repays a debt to society.’ Try to derive this principle logically, and you will find it impossible without directly assuming the desired outcome among the required premises. 

What we must see instead is that, in line with more agreeable consequentialist moral philosophies, the only appropriate consideration when assigning juridical interventions is: what actions will lead to a better society for us, and for our children to grow up in? In this case, the problem justifying enforced treatment (e.g. imprisonment) upon somebody who ‘couldn't have acted any other way’ disappears completely. The enforced treatment is only indirectly determined by the person’s actions, and is wholly derived from what we would like the world to look like in the future. The relevance of past behaviour is limited to the extent to which it serves as a predictor of future behaviour. What are traditionally viewed as punishments - justice administered for the satisfaction of the victims - become more properly viewed as treatments, designed to minimize the cost for society of a person’s demonstrated antisocial tendencies. 

The desire for revenge against a person who has committed wrongs against us is likely to be at least partly due to population genetics, naturally selected for self-preserving behaviour (it is advantageous for me to create an environment in which another’s bad behaviour toward me makes life uncomfortable for them), but the idea linking this concept of justice to free will seems to be far more memetic than genetic: it is a matter of culture.

The concept that free will is necessary and sufficient to entail the punishment of moral failing seems to date back to Aristotle, in Nicomachean Ethics. I’m no scholar of Aristotle, but to me its not clear whether for him the appropriateness of blame has a consequentialist or an absolutist foundation - are praise and blame desirable because they make certain modes of future behaviour more likely, or because they try to balance what has happened in the past?

If I had to speculate on the reason for the cultural success of the notion specifically linking retribution to free will, I’d guess that it was found to come in very useful when dictators wrestled with the seemingly contradictory goals of being loved, yet being utterly feared.

How can you be brutally violent against your enemies, while remaining admired by the remaining population? One way would seem to be to claim that violence against certain people is morally just, even necessary. “It made me cry to do that to him, but his crimes left me no choice.” Such pious adherence to absolute moral principle, even when it demands the most unpleasant actions, might even elevate a thug to saintly status, bringing joyous tears to the eyes of his devoted followers.

In the course of time, it may be that neuroscience, experimental psychology, and the social sciences will come to the conclusion that a better society is generally one in which people’s innate desire for vengeance is somewhat fulfilled (I doubt this, as I’ll explain shortly), but this would not undermine the principle that treatment of criminals should be determined on purely consequentialist grounds. If it happened to be that this desire was so strong, and so innate that no amount of cultural evolution could remove it, and that the frustration of unplacated victims of crime was so intense as to threaten civil unrest, then a retributive element may need to be restored, but the ultimate reasoning would be the rational evaluation of different courses of action, and selection in favour of those strategies determined to be in society’s best interests.

The debate between absolutist and consequentialist moral philosophies has been going on for a long time: consequentialism goes at least as far back as Machiavelli, around 500 years ago. Absolutism goes much further back, and persists still. This is really quite surprising - its not a difficult problem to solve. All morality is manifestly consequentialist, no matter what we might profess. 

Wait a moment, ‘thou shalt not kill.’ It doesn't get much more absolutist than that does it? No, it doesn't  But just how absolutist is that exactly?

For starters, no society implements principles like this in the strict absolutist way. Christians believe that this basic rule, ‘thou shalt not kill’ was handed to them by their personal deity: thou shalt not kill means that killing is absolutely wrong, under all circumstances - no exceptions allowed. Its never stopped Christian nations going to war when they felt like it. It never prevented Christian inquisitors burning people at the stake when the winter nights were dark and cold. All assumed absolutist principles have always been tacitly appended with a host of additional clauses beginning with the word ‘Unless...’ This is pure consequentialism.

Well, maybe those people adding their arbitrary ‘unless’ clauses were simply bad moralists. Thou shalt not kill is a good rule after all, right? Yes, typically. But what if the person who you are invited to consider killing has a strong ambition to kill you at the earliest convenient moment? Or alternatively, what if that person suffers intolerably, with no hope of improvement, ever? Killing can not be said to be categorically wrong under all circumstances - it all depends on the consequences.

Finally, absolutist versions of morality, in the sense that the content of the principle, “X is wrong,” takes precedence over the actual likely outcomes of performing X, are actually demonstrably incoherent. Lay aside the problem of what could possibly be the source of any absolute moral principle. Suppose for a moment that such principles really are set by some divine entity. What then? These moral laws are obviously not physical laws, since we have the capacity to systematically deviate (if we didn’t, they wouldn’t be called moral laws in the first place). Thus, somewhere in the process of our minds, decisions are made about whether or not to follow a particular moral principle at a particular time. If we believe that Godzilla will roast us alive for eternity if we fail to follow the rules, then those predicted consequences are what guide our behaviour. Moral decisions are always the result of a consequentialist evaluation of the options.

Going a little beyond the standard terminology, then, morality is absolute, but with only one rule: “whatever actions are revealed by a rational analysis to be most likely to bring me closer to achieving my goals are the actions I should implement.” This is exactly as I demonstrated in an earlier article on scientific morality. Furthermore, it illustrates that the founding principles of that argument, (1) goodness does not exist outside minds and (2) morality is doing what is good, are both properly basic: they are necessarily correct, and our knowledge of them is not contingent upon empirical observations.

Lets get back to the potential role of retribution in an advanced consequentialist morality. The extent to which the will to see wrongdoers punished is genetically innate, as opposed to culturally transmitted, is certainly an interesting question, and one whose investigation would no doubt require some ingenious experimental protocols. But I strongly suspect that the innateness of these feelings is limited to an extent that can easily be overruled by rationality, allowing vengeance to be effectively eliminated from all consideration in the problem of dealing with criminals. There are several reasons for this suspicion.

Firstly, if we look at the portion of the population most commonly found expressing anger, I’m fairly sure it'll be small children. Anger is, we all recognize, a childish emotion. We grow out of it. We learn (with great relief to most, I presume) to control it, and when as adults we occasionally succumb to emotional outbursts, we typically feel silly afterwards. As advanced society has developed, we have continually learned, oh so painfully slowly, that anger and resentment typically achieve little except the propagation of more anger and resentment. 

Secondly, there seems to be considerable evidence showing that the traditional practices of retributive justice have failed miserably. This paper, for example, argues strongly that imprisonment is ineffective at reducing the frequency and intensity of crime, and that alternative treatments such as education achieve greater reductions of recidivism. Another article summarizes some of its findings: "Research into specific deterrence shows that imprisonment has, at best, no effect on the rate of reoffending and often results in a greater rates of recidivism." The utilitarian advantages of a more rational approach seem to be there for the taking.

Thirdly, whatever memetic components there are, supporting any in-built tendency to desire vengeance, they can, by definition, be overcome by changing our culture.

Fourthly, religious leaders throughout history seem to have made artful use of the philosophy of free will in order to bolster acceptance of their reign of terror (hell doesn’t seem very fair, if all your actions are fixed by the way God set up the boundary conditions, and so damnation only gains a veneer of coherence if we have free will - a notion that evidently has to extend to the mortal plane, in order to justify certain historical hobbies of the major religions). This suggests that the hard-wired machinery of anger was, stripped of any socially conditioned props, insufficient to sustain the required levels of violence in our ever increasingly sophisticated culture.

When it comes to figuring out how to deal with crime, therefore, it is irrational to decide based on a shortsighted lust to see a criminal's debt repaid through suffering. Instead, we must look to scientific data to decide what courses of action minimize the costs to society. We must seek to understand what treatments will cost-effectively turn today's rule breakers into tomorrow's contributors to society, and what measures will economically eliminate the desire and the opportunity to commit crimes in the first place. 




Monday, June 17, 2013

Extreme values: P = 1 and P = 0




There is a popular folk theorem among some Bayesians, to the effect that it is unacceptable for a probability to be 0 or 1. There's a simple motivation for this principle: as rationalists, we demand the opportunity for nature to educate us by blessing us with novel observations. No matter how confident we become in some proposition, it should always be possible for us to change our minds when strong enough evidence accumulates in favour of some alternative. As Karl Popper rightly observed, after all, a theory that is invulnerable to falsification is not much of a theory.

But what happens if P(H | I) becomes zero? How is the probability for the hypothesis, H, to be updated by new evidence? If P(H | I) is 0 then the numerator in Bayes' theorem, prior times likelihood,

P(H | I) × P(D | HI)

is also 0, regardless how convincing the data, D, may be. No matter what happens, the outcome is unchanged: a nice round posterior.

Similarly, if P(H | I) is 1, then for the converse hypothesis, P(~H | I) is necessarily 0. Now, the denominator in Bayes' theorem is 

P(H | I) × P(D | HI) + P(~H | I) × P(D | ~HI)

and when the second term (everything after the plus sign) is zero, both numerator and denominator in Bayes' theorem are the same, producing the ratio 1, for all eternity.

I have sympathy with this motivation, therefore, but as a general rule, it is utter nonsense, resulting from forgetting one of the most basic facts about how inference works. The mathematics I have just described is all correct, but there are other ways for us to change our minds, and retain our rationality.

A recent, brief discussion at another website drew my attention to an article by Eliezer Yudkowsky, in which he also argues that 0 and 1 are not probabilities. The argument is a little different: the amount of evidence (the likelihood ratio expressed in log-odds form) needed to update an intermediate probability to 0 or 1 is infinite. This infinite certainty is an absurdity, he claims, unable to be represented with real numbers, and so 0 and 1 aren't probabilities.

Yudkowsky, as many readers will know, is a widely regarded thinker and writer on the topic of applied rationality, and I can recommend his writing most highly. The overlap between his broad philosophy and mine is, I would say, very large, with the main difference that in cases where I lack mastery of the theoretical apparatus, he very often does not. Yudkowsky knows and understands the mind-projection fallacy better than the vast majority (see for example his article of the same name, and this followup), but in this instance, he seems to have forgotten it. It is essentially the same error made by all who claim that probabilities equal to zero or one should not enter one's calculations.

A little thought experiment, then, before resolving the paradox. Let H be the hypothesis that in some five-day interval, at some location on the Earth, the sun will rise on each of the five mornings. Let D represent the observation of the sun rising on the first of the mornings in question. What is P(D | HI)? I humbly submit that it is 1. Is H, therefore, not an appropriate, well-formed hypothesis? Is D not a valid observation? Evidently, if probability theory is to have any power at all, it must be capable of supporting hypotheses such as H, and data as trivial as D. It is not conceivable to have such things automatically ruled out under our epistemology.

In general, it is perfectly legal for P(D | HI) (or, for that matter, a posterior, like P(H | DI)) to be 0 or 1, but here's that basic fact about probability that we have to keep in mind: a probability can not be divorced from the model within which it is calculated. A model may imply infinite certainty, without any person ever achieving that state (which would be impossible to encode in their brain, anyway). Our notation says something very important: P(D | HI), no matter what it is, is necessarily contingent upon the conjunction HI, which obviously depends on the truth of I. This is something we can never be absolutely certain of.

The all-important "I" that forms the foundation for every Bayesian calculation is usually said to stand for 'information' - all the relevant prior knowledge we have. Unfortunately, this creates a little trap that too many fall into, which is to forget that there is another component besides information needed before "I" is fully populated. "I" could just as easily stand for 'imagination.' To get Bayes' theorem to do any useful work for us, we have to specify a theoretical framework. We have to make certain assumptions, including specification of a full set of hypotheses against which is to H compete. To arrive at a candidate set of hypotheses, we must make a leap of the imagination. There is no possible criterion for judging whether or not all our assumptions are correct, and no way to know in advance whether we have chosen the 'correct' set of hypotheses. To think otherwise is just wishful thinking.

To think that the infinite confidence implied under some "I" represents the actual infinite confidence of some physical rational agent is the mind-projection fallacy. Instead, a probability is a model of the confidence a rational agent would have if "I" was known to be true. That this confidence might need to be modelled using a non-numeric concept such as infinity is merely an uncomfortable (though often highly convenient) mathematical fact.

And now we can see how it is that we can continue to accrue knowledge under the threat of the apparent epistemological cul de sac that is P = 1 or P = 0. To liberate ourselves from the straight jacket of "I", we simply need to recognize that what we now call "I" is itself merely a hypothesis in some broader hierarchical model. This is how model checking (wielding the analytical blade of model comparison) works, which, as I pointed out before, seems philosophically unpalatable to many, yet is in fact an essential ingredient in our inferential machinery. This is how we can come to look again at our theoretical framework and say 'hold on, I should be working with a different hypothesis space.' Novel theories and scientific revolutions would be impossible without this flexibility.

Some see this need in Bayesian epistemology to make assumptions in "I" that can't be established with certainty as a severe weakness, but it isn't - at least not one that can be avoided (no matter how many black belts we hold in the ancient art of self deception). We can always extend the scope of our hypothesis space so that some of our assumptions become themselves random variables in  a wider inferential context, but to have all of them take on the role of hypotheses under test would require an infinitely deep hierarchy of models. In the example above, where H was a hypothesis about the sun rising, one might argue that a more sophisticated model would account for the possibility, however small, that my sensation of the sun rising was mistaken. Indeed, this is correct, and would prevent the likelihood function going to 1. Sooner or later, though, I'm going to have to introduce a definitive statement - one that supposes something to be definitely true - in order to avoid the intractable quagmire of infinite complexity.

The early frequentists (and some still, in private communication with me), claimed that this subjectivity of Bayesian probability is its downfall, but in reality, it is impossible to learn in a vacuum. No kind of inference is possible without assumptions. Part of the beauty of Bayesian learning is that we make our assumptions explicit. The frequentists, of course, also make assumptions (see Yudkowsky, for example), but by refusing to acknowledge them, like the fabled ostrich sticking its head in the sand, they eliminate the possibility to examine whether or not they are reasonable, to understand their consequences, or to correct them when they are manifestly wrong.




Wednesday, May 22, 2013

Signal and Noise




Here's a useful thought experiment (slightly reworded) from Ronald Fisher's well-known text book from 1925, 'Statistical methods for research workers':
In each of two nearly identical universes, agricultural researchers wanted to compare 2 fertilizers (I know, it sounds like bullshit). In each universe, similar protocols were performed (this really happened, I swear): 2 plots of land were each divided into two parts, and the different parts treated with the different fertilizers. The same crop plant was cultivated on each part of each plot, and the individual yields recorded. The yields, in tons, were:
Universe 1:

plot
fertilizer A
fertilizer B
1
20
28
2
23
32


Universe 2:

plot
fertilizer A
fertilizer B
1
20
28
2
23
41

In which universe is there stronger evidence for the advantage of fertilizer B over fertilizer A?

In universe 1, the average advantage is 8.5 tons per plot, while in universe 2, the advantage is 13 tons. Seems like the farmers in universe 2 can place greater justified faith in fertilizer B.

But that conclusion is a bit too quick. There's more to the strength of evidence than just the magnitudes of the averages. We also have to consider the quality of the evidence, the signal to noise ratio. In universe 1, the numbers for each fertilizer are tightly clustered, (20 is not much different from 23, and 28 is not much different from 32) supporting the idea that the experiment was well controlled - random factors probably contributed little to the outcomes. 

In universe 2, however, the experiment doesn't look as well controlled. There's a big difference between the 2 results for fertilizer B, indicating that there is much more going on than just the choice of fertilizer. There's apparently more noise in this case, and if there's that much noise, then maybe the outcome of the experiment is purely down to random chance.

To analyze the relationship between the signal and the noise, Fisher recommends a null-hypothesis significance test (well, he invented significance tests, after all). Modelling the two sets of samples (A and B) as drawn from the same normally distributed population (the null hypothesis), we can calculate the plausibility of the observed difference between their means under the null hypothesis, H0. If the observed data, D, is too implausible under H0, then, as tradition goes, H0 is rejected. The problem is, we only have a few samples from which to calculate the width of that normal distribution, so another distribution, Student's t-distribution, which accounts for the uncertainty of the standard deviation, is used instead. To get a p-value, we have to calculate a t-statistic, and integrate the t-distribution from that statistic out to infinity to obtain the desired implausibility of D under H0.

To get the t-statistic, we can first calculate an aggregate sample standard deviation, s:



where xi, mi, and ni are respectively yields, averages, and the numbers of samples, for fertilizer i.

The t-statistic for comparison of two means (where H0 states that the 2 means are the same) is then given by


Taking account of the number of degrees of freedom in the experiment, nA + nB - 2, equal to 2, tables or almost any mathematical software are consulted to perform the required integral, which gives directly the p-value. For universe 1, the p-value I get from this test (two-tailed integration) is 0.041 - quite significant, the null hypothesis is on shaky ground. For universe 2, however, where we noticed that there was apparently a far greater random component to the data, the p-value is 0.11. This is almost 3 times larger, meaning that the data are here more believable under H0. We have weaker grounds for supposing that the fertilizers perform any different in universe 2.

This rare foray into the realm of orthodox stats has been hopefully sufficient to illustrate the point about quality of information, in terms of signal and noise (don't ask for guarantees that I did everything correctly, I may never understand the mindset under which these significance tests make sense). What I don't like, though, about the null-hypothesis significance test (among other things) is that no alternate hypotheses are formulated or evaluated. Without comparing Hto any other H's, the whole process is frankly rather hollow. What's more, if H0 is rejected, then further machinery is required to figure out how large the effect is. What I want, generally, is a set of proper probabilities, worked out for a whole range of possibilities, including H0.

Just for a laugh, then, I'll work through an approximate model, that'll allow me to plot a continuous probability distribution for a whole range of values for the average difference between the yields for the 2 fertilizers. I'll use the t-statistic again, but I won't be integrating tails (at least, not until after I have a posterior distribution). This t-statistic will help me get around the ambiguity concerning the width of the noise distribution, when calculating the likelihood function.

We can test the significance of the mean of a single set of samples from a single population, relative to some hypothetical mean, μ0, using another formula for the t-statistic:

where s is the regular sample standard deviation. (For emphasis, the reason for the different formula is that we are doing a different test - looking at a single mean, <x>, as opposed to comparing 2 means.)

The likelihood function, P(D | HI),  is calculated by evaluating the t-distribution with this statistic and the number of degrees of freedom, n - 1, which is 1.

The single population we're looking at is the population of differences between fertilizer B and fertilizer A. The parameter to vary in order to generate the likelihood function is the hypothesis, μ0.

As a prior density, I'll use a normal distribution centered at zero. To assign a width, I'll set the standard deviation to 15 tons, which means that there is a very small probability (< 5%) that the absolute value of the difference in yields for the two fertilizers exceeds 30 tons. The posterior probability is then simply the normalized product of this prior and the likelihood function, from Bayes' Theorem.

The graph below shows the posterior probability density as a function of μB-A, for each of our universes. The curves illustrate clearly the impact of decreasing the signal to noise ratio in universe 2: though the peak is further to the right, the tail extends further to the left, due to the lower sensitivity of the experiment in that universe. Integrating the two curves from -∞ to 0, we see that the probability that fertilizer B is actually not better than fertilizer A is twice as large in universe 2 as it is in universe 1, which is similar to the result above, in terms of p-values. As the old saying goes: garbage in, garbage out. Precise inference demands a well-controlled experiment.



Its only human that quite often in the quest for knowledge we'll derive greater confidence from results like universe 2, rather than universe 1. It takes care not to be seduced by a greater overall difference, before taking time to consider how much of that difference is likely to be due to random fluctuations. Often, for brevity, we'll summarize an experiment by recording only the mean result (or in less formal circumstances subconsciously estimate the mean, and forget all the other details), but as we've seen, to draw good-quality inferences we need to note not only the mean but also the dispersion and the number of samples (roughly, confidence in a result scales1 according to SNR ×  n  ). How many figures quoted by politicians or newspapers (or anybody else with influence) lose their sting when we notice that no error bar has been provided? Context is all important. Sometimes, only moderately careful analysis is enough to overturn an intuitively appealing conclusion, and it's results like this that show the importance of a cultivated awareness of the mathematical machinery of rational inference.





[1]
'Why randomized controlled trials fail but needn't: 2. Failure to employ physiological statistics, or the only formula a clinician-trialist is ever likely to need (or understand!),' D.L. Sackett, CMAJ October 30, 2001 vol. 165 no. 9, link