Showing posts with label math. Show all posts
Showing posts with label math. Show all posts

Monday, October 31, 2011

Arguments from Probability

As with any logical proposition, one contradiction disproves the proposed rule. If each of 38 counterexamples has merely a 10% chance of being valid — an underestimate — then the probability that the Earth is billions of years old is less than 3.8%. In other words, the Earth must be young with a likelihood of greater than 96%.
(From the Conservapedia.)
It is a statistical certainty (p < 10-11) that there are innocent people being held at Guantanamo Bay.
(Sig on Slashdot.) These are a couple of cases I've seen recently of probability theory misused. The first one is easier to deal with, so let's start with that. The claim is that each counterexample has "merely a 10% chance of being valid — an underestimate". This claim is already problematic. Take this counterexample:
The Bible makes references to the dinosaurs. There is no explanation for this if dinosaurs supposedly lived hundreds of millions of years ago.
How would we calculate the odds that this claim is correct at greater or less than 10%? For the record: Biblical evidence for dinosaurs is razor-thin. It mostly rests upon Job 40, which reads (in the King James Version):
Behold now behemoth, which I made with thee; he eateth grass as an ox. Lo now, his strength is in his loins, and his force is in the navel of his belly. He moveth his tail like a cedar: the sinews of his stones are wrapped together. His bones are as strong pieces of brass; his bones are like bars of iron.

I simply ask the reader: do you think there is a greater or less than 10% chance this describes a dinosaur? Follow-up question: if dinosaurs — huge beasts larger than anything else that ever walked this earth — lived in Biblical times, do you think it's likely there would be one single reference?

It strikes me that the probability that this is true is less than 10% — much less, in my judgment. But ultimately the point is that this is a judgment call.

But there is a deeper problem with the Conservapedia's argument. It makes the implicit assumption that each counterexample is independent. Here's a thought experiment: suppose we roll a fair 6-sided die and try to determine which face is on the table (that is, the one we cannot observe). We observe each of the five other faces, and find they are 1, 2, 3, 4 and 5. For each one we say: we did not observe a 6, which means the odds of the bottom face being a 6 is 5/6. In total, then, the odds of the bottom face being a 6 is (5/6)^6 or about 33.5%. But this is obviously absurd, since we know that the bottom face must be a 6: we've seen all the other faces. The reason the argument fails is that our observations are not independent.

The same holds with Conservapedia's argument. Consider these two claims:

1. The Moon's orbit is a very strong counterexample: the moon is receding from the Earth at a rate[3] that would have placed it too close to the Earth merely four billion years ago, causing instability in its orbit, tidal catastrophes on Earth, and other problems that would have prevented the Earth and the Moon being as they are today. Additionally, the moon's orbit is becoming increasingly and unexpectedly eccentric, suggesting a lack of long-term stability,[4] which further disproves the theory of an Old Earth.

4. The planetary orbits in the Solar System — including Earth's — are unstable and unsustainable over the long periods claimed by Old Earth believers.

Let's ignore the question science behind these claims and take them at face value. The problem is that even if both are true, both may simply be indicative that orbits are unstable. If that claim has a 10% chance of being true, then both claims 1 and 4 would be true, but would not make the basic fact any more or less true. If we observed planets in other star systems that had unstable orbits, for example, we could toss in additional claims about them, but they would not be independent of the claims already made and thus would add no value to the overall question of whether the universe is old or young.

Very well. Let's consider the Guantanamo prisoner claim now. We can back into the assumptions made pretty easily. If the likelihood of an individual prisoner being guilty is g, then the likelihood that every prisoner is guilty is g^N where N is the number of prisoners. Currently N=171 if Wikipedia is reliable. The claim is that g^171 < 10-11, so this means g < 0.86.

That strikes me as low. But perhaps the calculation was made when we had more detainees. The maximum number was 775, in which case g < 0.968. Eh, maybe. But as with the Old Earth argument, one problem is that we are asked to believe a specific probability about the individual case, and then extrapolate it to the overall case. The argument is extremely sensitive to these probabilities: suppose that actually g = 0.995. Then the odds of a single innocent detainee would be only 58%. We have no idea what the probability g really is, though, so we are left making guesses. And these guesses strongly influence the outcome.

Independence is less of an issue here, but it's important to mention another factor. Not each detainee may have the same probability of being innocent. This can make a big difference. Suppose that there are 10 detainees, 9 of which are definitely guilty and 1 of which has a 50% chance of being guilty. Alternately, suppose there are 10 detainees who each have a 95% chance of being guilty. In both cases, a randomly chosen detainee has a 95% chance of being guilty. But the probability of having at least one innocent detainee is 50% in the first case and 40% in the second. It is easy to craft other examples that can push these probabilities as far apart as you like.

One thing I've ignored covering here is the question of whether these numbers are actually relevant. That's a big topic that could cover large swathes of the philosophy of science, but suffice for now to point out that probability has a somewhat different meaning in the Old Earth question than in, say, gambling. Either the Earth was created about 4.5 billion years ago or it was not: this is a question fundamentally unlike asking whether the next dice roll will be boxcars. The similarity between the two is that both deal with unknowns. We will never know, with 100% certainty, whether the age of Earth is about 4.5 billion years, so the best we can do is assemble all the evidence and weigh it. That's how science works.

On the Guantanamo question, I have similarly questions of relevance. Suppose the claim was true that at least one detainee was innocent with statistical certainty. How would that affect our policy? Would it help us determine which one? Would we be willing to let terrorists go in order to decrease the probability of detaining innocents? Such questions cut to the heart of any judicial policy, and they can't be answered by facile calculation.

Friday, July 1, 2011

Tau+3

I'm a few days late on this, but Tau Day (6/28) was earlier in the week. Tau is a new idea to replace pi (the ratio of the circumference of a circle to its diameter) with the more natural constant tau (the ratio of the circumference of a circle to its radius). Seems crazy, but there's a pretty good case to be made for this.

The problem is that pi is so entrenched, I can't see it ever being changed. But it's interesting nonetheless. If I were still in the habit of writing mathematical papers I'd try to insert tau here and there (appropriately footnoted) and see if it caught on.

Wednesday, December 22, 2010

Misconceptions in Math Education

This is worth a read.

The great misconception about mathematics -- and it stifles and thwarts more students than any other single thing -- is the notion that mathematics is about formulas and cranking out computations. It is the unconsciously held delusion that mathematics is a set of rules and formulas that have been worked out by God knows who for God knows why, and the student's duty is to memorize all this stuff. Such students seem to feel that sometime in the future their boss will walk into the office and demand "Quick, what's the quadratic formula?" Or, "Hurry, I need to know the derivative of 3x^2 - 6x +1." There are no such employers.

Mathematics is not about answers, it's about processes.

Monday, November 1, 2010

Prediction Time

In the Senate, I'm predicting a 7-seat gain for the GOP. Ignoring the seats with 10-point leads one way or the other, there are 8 seats in play, with 2 locked up for the GOP. Those 8 seats are: California (DEM+3), Colorado (GOP+4), Illinois (GOP+2), Nevada (GOP+4), Pennsylvania (GOP+5), Washington (GOP+1), West Virginia (DEM+3), and Wisconsin (GOP+7), where I'm giving the latest poll results in each case. Many of these polls are within the margin of error, but we can still calculate probabilities for each. Based on these numbers, we can find a joint probability distribution for various levels of GOP pickup.

GOP GainProbability
+2negligible
+30.04%
+40.6%
+54.6%
+617.6%
+734.1%
+830.9%
+910.9%
+101.3%

The median of the distribution is at the high end of a 7-seat GOP pickup. It's certainly pretty likely that we'll have an 8-seat pickup, but playing the odds I'll predict +7.

In the House, it's more complicated since there are a lot more races. Also, polls are sketchier. But let's use the latest Real Clear Politics map to make some estimates. If we plug in some reasonable probabilities corresponding to "likely GOP", "leaning GOP" and "tossup", we get this distribution for GOP seats (here I'm ignoring insignificant tails):

GOP SeatsGOP LeadProbability
Under 2230.07%
223100.06%
224120.10%
225140.17%
226160.28%
227180.45%
228200.69%
229221.01%
230241.45%
231261.99%
232282.65%
233303.40%
234324.22%
235345.06%
236365.86%
237386.55%
238407.07%
239427.38%
240447.44%
241467.24%
242486.80%
243506.18%
244525.41%
245544.58%
246563.74%
247582.95%
248602.25%
249621.65%
250641.17%
251660.80%
252680.53%
253700.34%
254720.21%
255740.12%
256760.07%
Over 2560.08%

The median of the distribution is for the GOP to have 240 seats, representing a 61-seat GOP pickup. There is a 96% chance the GOP will pick up between 50 and 76 seats, and a 90% chance they pick up between 52 and 69 seats. So I think we can count on the GOP having at least a 26-seat majority and it's not impossible for them to attain a 60-seat majority.

Of course, this will do them little good without the Senate or the White House. Those will have to wait until 2012.

Tuesday, September 7, 2010

Math Corner

John Derbyshire has another good one up in his August Diary. It's another probability question, this time about cards:

I have an ordinary deck of 52 playing cards. I shuffle it thoroughly. What is the probability that not one card is in its original position?

As often happens, the way to approach this is to look at the conjugate question: what is the probability that at least one card is in its original position? Suppose, for example, that one card is in its original position - we'll call this a stationary card, because it didn't move after shuffling. There are C(52,1) ways to pick this one card, and the other 51 cards can be arranged arbitrarily, so there are 51! ways to arrange them. However, we've done some double-counting here, because some of those 51! include arrangements that have a second stationary card.

It's worth going over this point in some detail, because this is an argument we're going to come back to again. To see what's happening here, it's useful to reduce the number of cards. So let's say there are only 4 cards, numbered 1, 2, 3 and 4. There are, of course, 4! = 24 possible arrangements of these cards. Let's look at all the arrangements with the "1" card in its original position:

1,2,3,4
1,2,4,3
1,3,2,4
1,3,4,2
1,4,2,3
1,4,3,2

Also, let's look at all the arrangements with the "2" card in its original position:

1,2,3,4
1,2,4,3
3,2,1,4
3,2,4,1
4,2,1,3
4,2,3,1

Notice something? The lists aren't distinct. Those first two entries are common to both lists. What we've done is double-count arrangements that contain at least two stationary cards.

We can subtract those back out pretty easily: there are C(52,2) ways to pick two cards, and then 50! arrangements of the other 50 cards. So we'll subtract C(52,2)50! arrangements.

But wait! We've removed too much, because both of the previous sets included arrangements that had at least three stationary cards. Going back to the 4-card example, take a look at the arrangement "1,2,3,4". We originally double-counted it, but then we double-removed it. So to count those arrangements we have to add back in arrangements with at least three stationary cards, and that's C(52,3)49!.

It should be no surprise at this point that this pattern continues. Now we've double counted arrangements with four stationary cards, so we subtract C(52,4)48! of those, at which point we need to add back the ones with five - there are C(52,5)47! of them - and so on. We end up with this many arrangements:

C(52,1)51! - C(52,2)50! + C(52,3)49! - C(52,4)48! + ... + C(52,51)1! - C(52,52)0!

There are 52! total arrangements of 52 cards, so to get the probability, we divide the above expression by 52!. This simplifies down to:

1/1! - 1/2! + 1/3! - 1/4! + ... + 1/51! - 1/52!

But this is the probability of having at least one stationary card, and we wanted the probability of having zero stationary cards, which is:

1 - 1/1! + 1/2! - 1/3! + 1/4! + ... - 1/51! + 1/52!

Reasoning the same way you can see that if you had n cards, the probability would be the first n+1 terms of this series. (It's an interesting fact that the above expression is the first n+1 terms of something called the Taylor series for 1/e, where e is the base of natural logarithms that you may dimly remember from high school or college. For more than 8 or so cards, the difference between the actual probability and 1/e is very small: less than 0.01%.)

Let's look at Derbyshire's second (related) problem:

I have a deck of n cards, numbered from 1 to n. I shuffle the deck thoroughly. Then I turn the cards over one by one. If the k-th card I turn over bears the number k, call that a "match." What is the probability that after going through the whole deck I shall have tallied m matches, where m is some number in the range from zero to n?

Based on the work we did before, this isn't hard at all.

Let's write the number of arrangements of n cards that have m matches (what I earlier called "stationary cards") as A(n,m). Then the probability p(n,m) that Derbyshire seeks is just A(n,m)/n!. Furthermore, we can reason about A(n,m) as follows: Suppose we select m cards and call them the matches. There are C(n,m) ways to make this selection. For each selection, there are A(n-m,0) arrangements of the remaining cards that have no matches. Any combination of a selection of m matches along with an arrangement of n-m cards that have no matches is equivalent to an arrangement of n cards with m matches. So A(n,m) = A(n-m,0)C(n,m).

Furthermore, since p(n,m) = A(n,m)/n!, we can write p(n,m) = A(n-m,0)C(n,m)/n! = p(n-m,0)C(n,m)(n-m)!/n! = p(n-m,0)/m!. Since our solution to the first problem gave us p(n,0) for every n, we now know how to calculate p(n,m) for every n and m.

Friday, July 23, 2010

The Tuesday Child Problem, Solved

At long last, the solution.

As in the case with flipping two coins, the problem revolves around figuring out what options have been eliminated by the statement of the problem. To review, the statement is that I have two children, one of whom is a born born on a Tuesday. The question is what is the probability that my other child is a boy?

First, I'll do this formally, to show that formal mathematics works just fine to solve problems like this. That is, you don't have to stoop to writing out the various possibilities and then tabulating them, as we did in the coin problem.

There is a formula in probability theory that comes in handy here:

(1) P(X|Y) = P(X and Y) / P(Y)

(The notation P(X|Y) means: the probability that X occurs given that Y has occurred, or in short-hand: the probability of X given Y.)

From (1), by the way, you can get another useful formula:

(2) P(X and Y) = P(Y) P(X|Y)

And since P(X|Y) = P(X) if X and Y are independent, that last formula can be written P(X and Y) = P(X) P(Y) in the case of independence.

Finally, since the "and" operation is commutative (i.e., the order of its operands is irrelevant), we have:

(3) P(Y and X) = P(X) P(Y|X) = P(Y) P(X|Y)

In this case, we want to determine:

P(both my children are boys | one child is a boy born on a Tuesday)

By (1), that's equal to P(both my children are boys and one child is a boy born on a Tuesday) / P(one of my children is a boy born on a Tuesday).

The first probability is clearly equivalent to P(both my children are boys and one of them was born on Tuesday), which by (2) is P(both my children are boys) P(one of my children was born on Tuesday). The first of these is easy to calculate: 1/4. The second is trickier, but the easy way to see it is to imagine that neither was born on Tuesday. Since the likelihood for either not being born on Tuesday is 6/7, the chance that neither was born on Tuesday is just 6/7 squared, or 36/49. That means the chance that at least one was born on Tuesday is 1 - 36/49 = 13/49. So that first factor in our equation above is (1/4)(13/49) = 13/196.

Now let's look at the second factor, P(one of my children is a born born on Tuesday). We can apply similar logic by imagining that neither is. The chance that a single child is a boy born on Tuesday is clearly 1/2 (boy) times 1/7 (born on Tuesday) = 1/14. So the chance that a single child is NOT a boy born on Tuesday is 13/14. Hence the chance that neither of two children is a boy born on Tuesday is 13/14 squared, or 169/196. The chance that at least one child is a boy born on Tuesday is then 1 - 169/196 = 27/196.

Plugging these values back into our original formula, we get the final probability as (13/196) / (27/196) = 13/27. And that is the answer.

Another way to do it (a bit less formally, but still perfectly accurate) is more like the two-coins problem. Let's write B for boy, G for girl, T for Tuesday, and !T for not-Tuesday. Then all possible outcomes for one child can be written like this:

G
BT
B!T

These are not equally likely, of course. P(G) = 1/2, P(BT) = 1/14, and P(B!T) = 3/7. But these add up to 1, and are mutually exclusive, so they must cover every possibility.

So for two children, there are nine possible outcomes:

G,G
G,BT
G,B!T
BT,G
BT,BT
BT,B!T
B!T,G
B!T,BT
B!T,B!T

Again, though, these are not equally likely. But their probabilities are easy to compute since the outcomes are independent. Furthermore, some of these outcomes are not possible since we know that one child is a boy born on Tuesday. So the only outcomes remaining are:

G,BT => (1/2)(1/14) = 1/28
BT,G => (1/14)(1/2) = 1/28
BT,BT => (1/14)(1/14) = 1/196
BT,B!T => (1/14)(3/7) = 3/98
B!T,BT => (3/7)(1/14) = 3/98

The sum of these is 1/28 + 1/28 + 3/98 + 3/98 + 1/196 = 27/196. Two boys show up with probability 3/98 + 3/98 + 1/196 = 13/196. The ratio is 13/27, as before.

Wednesday, July 7, 2010

Math Surprises

Some of the best math problems are the simplest to state. But they can be fiendishly difficult to understand, precisely because their domains are so commonplace. A good example is fairly well known: Suppose I have two coins. One of them is showing heads. What is the probability the other one is also showing heads?

The answer most people jump to is 1/2, since we've had it drilled into us that coins are independent, so the outcome of one shouldn't affect the other. But the phrasing of the question contains a subtle trick. By stating that one of the coins is showing heads, I have eliminated the possibility that both coins are showing tails. But there are actually three remaining possibilities: heads/tails, tails/heads, or heads/heads. In only one of these three are both coins heads, so the answer is in fact 1/3.

The trickiness of the phrasing comes from the fact that when we say that one of the coins is heads, we deliberately don't say which one. We could say equivalently that at least one is showing heads. The probability is the expected 1/2 once we specify ahead of time which coin it is.

Given this preliminary, consider the following problem, which has been making its way around the Web of late:

I have two children, one of whom is a boy born on a Tuesday. What is the probability the other one is a boy? (You can make all the expected simplifying assumptions: no twins, boys and girls born in equal proportions, each day of the week is equally likely, etc. It's a math problem, not an obstetrics exercise.)

I will post the answer in a few days.

Wednesday, June 23, 2010

World Cup Tiebreakers

John J. Miller finds it "lame" that there is (or was: the games are over) a possible outcome in which the U.S. and England, tied for 2nd and 3rd place in Group C, would have their positions determined by lottery.

He may not be aware of this, but the same is true in the American football playoff system. There is a series of tiebreakers, the last of which is "coin toss".

This isn't surprising: at some point in any round-robin system, it's possible that two or more teams may tie on any given number of measures. If that happens, something must be introduced to break the tie. You might say: use a one-game playoff, which works fine if two teams tie. But what if three teams tie? That's certainly a possible outcome, and further playoffs may do nothing to resolve the situation. So a coin toss, or lottery, is necessary just to come to a decision. It may be "lame", but it's unavoidable.

Happily, the U.S. and England each won their games and advanced to the knockout round without the need for any lotteries.

Monday, May 24, 2010

Martin Gardner, R.I.P.

Like many math-obsessed geeks who lived in the second half of the twentieth century, I read with poignant regret of the passing of Martin Gardner this past weekend.

My introduction to him was through his regular Scientific American column "Mathematical Games". During my formative years, personal computers were shifting from rarity to ubiquity, and many of his columns were well suited to computer experimentation: Conway's game of "Life", Mandelbrot's fractals, etc. I whiled away many a rainy (and not a few sunny) days writing BASIC code to bring those beautiful structures to the screen of my very own Apple IIc.

John Derbyshire writes his own tribute on the Corner. A great man, Gardner. He will be missed.

Friday, May 7, 2010

U.S. House of Reps Electoral Math

Jim Geraghty of National Review provides an invaluable breakdown of 99 Congressional seats that are in play this fall. Let's start with that and apply some math.

The current breakdown of the House is 253 Democrats, 178 Republicans, and 4 open seats. So either party needs 218 seats to secure a majority. According to Geraghty, all the Republican seats except one (Joseph Cao of Louisiana) are "safe." And 158 Democratic seats are also "safe." Let's assume that all of those safe seats remain unchanged (or at least that if one or two switch sides, they will be balanced by other switches so that the total numbers remain unchanged). What is the probability that the Republicans gain control of the House?

I assumed that a "code blue" in Geraghty's construction was an 80% chance of winning, a "code green" was 65%, a "code yellow" 50%, a "code orange" 35%, and a "code red" 20%. I also placed Joseph Cao in the "code orange" category.

The Republicans need to pick up 41 seats to get a majority. If Geraghy is right and my probabilities are reasonable, that gives them a 99.5% chance of doing so. Wow! It's worth looking at how the odds break down for some other possible splits.

Split# of GOP PickupsProbability
DEM +1 (or more)40 or les0.5%
GOP +1 (or more)41+99.5%
GOP +5 (or more)43+98.5%
GOP +11 (or more)46+93.7%
GOP +15 (or more)48+86.5%
GOP +21 (or more)51+68.0%
GOP +25 (or more)53+51.8%
GOP +31 (or more)56+27.8%

It's certainly looking very good for the GOP right now. Just to check myself, I re-ran the study with more conservative probabilities (thus reducing the likelihood of taking over any seat), and even then the GOP has an 81% chance to end up with a 15-seat (or better) majority. It's the GOP's race to lose.

Monday, March 8, 2010

Math Allegories in "Alice in Wonderland"

Fascinating article in the NY Times about Lewis Carroll's Alice in Wonderland.

In the mid-19th century, mathematics was rapidly blossoming into what it is today: a finely honed language for describing the conceptual relations between things. Dodgson found the radical new math illogical and lacking in intellectual rigor. In "Alice," he attacked some of the new ideas as nonsense — using a technique familiar from Euclid’s proofs, reductio ad absurdum, where the validity of an idea is tested by taking its premises to their logical extreme.

Some of Carroll's satire is simply wrong - de Morgan's discovery of complex numbers turned out to be integral to modern mathematics and, in a twist of irony, physics and even engineering. But it's still worth bearing in mind when reading Alice.

Tuesday, December 15, 2009

Slicing Pizza

When you get a pizza, it's usually not sliced exactly through the center. So some pieces are bigger than others. For over 40 years, mathematicians have been working on the problem of classifying exactly who gets more if two diners share alternate pieces of pizza from an unequally-cut pie.

This classification has now been found. It's not the most ground-breaking research. Deierman and Mabry won't become household names, and won't be nominated for a Fields Medal for this work. But I love this sort of thing because it shows how the application of some intelligence, hard work, and cleverness (not the same thing as intelligence) can simplify decidedly nontrivial problems.

The solution, btw, is simple enough to write down here:

  • If at least one cut passes through the center, then the two diners will always get equal amounts.
  • If you cut your pizza an even number of times (say, 4), then the two diners will always get equal amounts.
  • If you cut your pizza an odd number of times and that odd number is expressible as 4n+3 (say, 3, 7, etc.) then the person who eats the piece that contains the center of the pizza gets more.
  • If you cut your pizza an odd number of times and that odd number is expressible as 4n+1 (say, 5, 9, etc.) then the person who eats the piece that does not contain the center of the pizza gets more.

One cool consequence of this result is that, if you have a pizza with an odd number of unequal cuts, and you want to share it equally between two people, all you have to do is just cut it one more time (the new cut must pass through the common intersection of all the previous cuts). Then as long as the diners take alternate pieces, they will get equal amounts. That's a nontrivial, counterintuitive result, and with a practical application!

Thursday, August 13, 2009

To Sample or Not To Sample?

For over a decade the idea of using statistical sampling instead of a straight count for the U.S. Census has been hot in some circles. Broadly speaking, Democrats have favored the method, while Republicans have opposed it.

Democrats expect to gain from using sampling, since traditional methods tend to undercount Democratic constituencies: minorities, illegal immigrants, and transients. And for the same reasons, Republicans expect to lose. So their respective support and opposition aren't surprising. Naturally, their public reasoning cannot be so self-centered.

Republicans base their formal opposition in part on the U.S. Constitution, which states: "The actual Enumeration shall be made within three years after the first meeting of the Congress of the United States, and within every subsequent term of ten years." The word "Enumeration" contains an implicit demand that each person be counted. The 14th Amendment includes similar language: "Representatives shall be apportioned among the several States according to their respective numbers, counting the whole number of persons in each State..." (emphasis mine). However, I find this line of argument unpersuasive. While an "enumeration" can mean a "list or catalog", it can also mean a "reckoning or count". While the former definition would exclude sampling, the latter would not. The phrase "whole number of persons" might be more persuasive, except that that phrase is pretty clearly used to void the original text's treatment of slaves: "Representatives and direct Taxes shall be apportioned among the several States which may be included within this Union, according to their respective Numbers, which shall be determined by adding to the whole Number of free Persons... three fifths of all other Persons." So the meaning of the Amendment is that every person shall be counted equally, i.e. there will be no more three-fifths rule.

Some Republicans also object that sampling can make the census less accurate. This is a better argument, since if true it is devastating to the whole basis of sampling. Technical Report 537, by Brown, Eaton, Freedman, Klein, Olshen, Wachter, Wells and Ylvisaker of the Department of Statistics, U.C. Berkeley, goes into great detail on the legal issues, techniques, and potential problems with sampling. The possibility that sampling will reduce accuracy is raised: "Will proposed adjustments to the census take out more error than they put in?" Unfortunately, the question is left unanswered: even these professors are unsure whether sampling will satisfy its intended purpose. The paper also points out that one problem sampling is meant to address - undercounting of certain groups - could be solved without sampling: "Census figures could be scaled up to match the demographic analysis totals for subgroups of the national population defined by age, sex and race (Section V). The people in a demographic group who are thought to be missing from the census would be added back, in proportion to the ones who are counted—state by state, block by block."

On the pro-sampling side, one of the most vocal supporters is Dr. Kenneth Prewitt, director of the Census Bureau from 1998-2001. He argues that sampling will be more accurate than a direct count, although Tech Report 537 seems to cast some doubt on this. Another point he makes is that the federal government government already uses statistical data for policy-making all the time. This point, though, is easily refuted. The Census Act as amended in 1976 states:

...except for the determination of the population for purposes of apportionment of Representatives in Congress among the several States, the Secretary shall, if he considers it feasible, authorize the use of the statistical method known as "sampling" in carrying out the provisions of this title. [emphasis mine]


The emphasized portion was the basis of the U.S. Supreme Court's 1998 decision in "U.S. House of Representatives v. U.S. Department of Commerce, et al.," which upheld the principle that sampling should not be used for apportionment.

My own take on the debate is that the Constitution demands that an accurate count be made by any means. I do not believe that a direct count is required. The census is a critical part of our representative democracy, though (which is why it is covered in our minimalist Constitution at all), and it is important that it be simple and transparent. The appearance of bias or political machination would erode civic trust in the most representative of our federal legislative bodies. A direct count that was just slightly less accurate than some other technique (sampling or something else) might still be favored due to its transparency.

The concerns raised in Tech Report 537 lead me to conclude that any improvement in accuracy from sampling would not be sufficiently great or certain enough that we should use it for the census. But future developments could change this balance: direct counts could become far less accurate (for some reason) or indirect counts far more accurate.