f14-4
PDF · 9 pages · 83.2 KB
Open PDF file
Excerpt from the textbook Numerical Recipes in Fortran 77 (Cambridge University Press, 1986-1992), starting at page 622 of Chapter 14, Statistical Description of Data. It closes section 14.3 on the Kolmogorov-Smirnov test, then begins 14.4 on contingency tables for nominal variables, using chi-square, degrees of freedom, and Cramer's V and the contingency coefficient. It is a published book sample, not Phil's own work.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
622 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).but, because of its cumulative nature, a K–S test would require many data points in the notch
before signaling a discrepancy.
Second,weshouldnotethat,ifyouestimateanyparametersfromadataset(e.g.,amean
andvariance),thenthedistributionoftheK–Sstatistic Dforacumulativedistributionfunction
P(x)thatuses the estimated parameters is no longer given by equation (14.3.9). In general,
you will have to determine the new distribution yourself, e.g., by Monte Carlo methods.
CITED REFERENCES AND FURTHER READING:
von Mises, R. 1964, Mathematical Theory of Probability and Statistics (New York: Academic
Press), Chapters IX(C) and IX(E).
Stephens, M.A. 1970, Journalof the RoyalStatistical Society , ser. B, vol. 32, pp. 115–122.[1]
Anderson,T.W.,andDarling,D.A.1952, AnnalsofMathematicalStatistics , vol.23,pp.193–212.
[2]
Darling, D.A. 1957, Annals of Mathematical Statistics , vol. 28, pp. 823–838. [3]
Michael, J.R. 1983, Biometrika , vol. 70, no. 1, pp. 11–17. [4]
No´e, M. 1972, Annals of Mathematical Statistics , vol. 43, pp. 58–64. [5]
Kuiper,N.H. 1962, Proceedingsof theKoninklijkeNederlandseAkademie vanWetenschappen ,
ser. A., vol. 63, pp. 38–47. [6]
Stephens, M.A. 1965, Biometrika , vol. 52, pp. 309–321. [7]
Fisher, N.I., Lewis, T., and Embleton, B.J.J. 1987, Statistical Analysis of Spherical Data (New
York: Cambridge University Press). [8]
14.4 Contingency Table Analysis of Two
Distributions
Inthissection,andthenexttwosections,wedealwith measuresofassociation
fortwodistributions. Thesituationisthis: Eachdatapointhastwoormoredifferent
quantitiesassociatedwithit,andwewanttoknowwhetherknowledgeofonequantity
gives us any demonstrableadvantage in predicting the value of another quantity. In
manycases,onevariablewillbean“independent”or“control”variable,andanotherwill be a “dependent” or “measured” variable. Then, we want to know if the latter
variableisin fact dependent on or associated with the former variable. If it is, we
wanttohavesomequantitativemeasureofthestrengthoftheassociation. Oneoftenhears this loosely stated as the question of whether two variables are correlated or
uncorrelated , but we will reserve those terms for a particular kind of association
(linear, or at least monotonic), as discussed in §14.5 and §14.6.
Notice that, as in previous sections, the different concepts of significance and
strength appear: The association between two distributions may be very significanteven if that association is weak — if the quantity of data is large enough.
Itisusefultodistinguishamongsomedifferentkindsofvariables,withdifferent
categories forming a loose hierarchy.
•Avariableiscalled nominalifitsvaluesarethemembersofsomeunordered
set. For example, “state of residence” is a nominal variable that (in the
U.S.) takes on one of 50 values; in astrophysics, “type of galaxy” is a
nominalvariablewiththethreevalues“spiral,”“elliptical,”and“irregular.”
14.4ContingencyTableAnalysisofTwoDistributions 623Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).•Avariableistermed ordinalifitsvaluesarethemembersofadiscrete,but
ordered,set. Examplesare: gradein school,planetaryorderfromtheSun(Mercury = 1, Venus = 2, ...), number of offspring. There need not be
any concept of “equal metric distance” between the values of an ordinal
variable, only that they be intrinsically ordered.
•We will call a variable continuous if its values are real numbers, as
are times, distances, temperatures, etc. (Social scientists sometimes
distinguishbetween intervalandratiocontinuousvariables,butwedonot
find that distinction very compelling.)
A continuous variable can always be made into an ordinal one by binning it
into ranges. If we chooseto ignorethe orderingof the bins, then we can turn it into
a nominal variable. Nominal variables constitute the lowest type of the hierarchy,
and therefore the most general. For example, a set of severalcontinuous or ordinal
variables can be turned, if crudely, into a single nominal variable, by coarsely
binning each variable and then taking each distinct combinationof bin assignments
as a single nominal value. When multidimensional data are sparse, this is oftenthe only sensible way to proceed.
The remainder of this section will deal with measures of association between
nominalvariables. For any pair of nominal variables, the data can be displayed as
acontingency table , a table whose rows are labeled by the values of one nominal
variable, whose columns are labeled by the values of the other nominal variable,and whose entries are nonnegative integers giving the number of observed events
for each combination of row and column (see Figure 14.4.1). The analysis of
association between nominal variables is thus called contingency table analysis or
crosstabulation analysis.
We will introduce two different approaches. The first approach, based on the
chi-squarestatistic,doesagoodjobofcharacterizingthesignificanceofassociation,
but is only so-so as a measure of the strength (principally because its numerical
values have no very direct interpretations). The second approach, based on theinformation-theoreticconceptof entropy,saysnothingatallaboutthesignificanceof
association (usechi-squareforthat!),but is capableof veryelegantlycharacterizing
the strength of an association already known to be significant.
Measures of Association Based onChi-Square
Some notation first: Let Nijdenote the number of events that occur with the
first variable xtaking on its ith value, and the second variable ytaking on its jth
value. Let Ndenote the total number of events, the sum of all the N ij’s. Let Ni·
denote the number of events for which the first variable xtakes on its ith value
regardless of the value of y;N·jis the number of events with the jth value of y
regardless of x.S o w e h a v e
Ni·=/summationdisplay
jNij N·j=/summationdisplay
iNij
N=/summationdisplay
iNi·=/summationdisplay
jN·j(14.4.1 )
624 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).1. male
2. female
.......... . ....
. . .. . .. . .. . . 1.
red
# of
red males
N
11
# of
red females
N21# of
green females
N22# of
green males
N12# of
males
N1⋅
# of
females
N2⋅2.
green
# of red
N⋅1# of green
N⋅2total #
N
Figure 14.4.1. Example of a contingency table for two nominal variables, here sex and color. The row
andcolumnmarginals(totals) areshown. Thevariables are “nominal,”i.e.,theorderinwhichtheir values
are listed is arbitrary and does not affect the result of the contingency table analysis. If the ordering
of values has some intrinsic meaning, then the variables are “ordinal”or“continuous, ”and correlation
techniques ( §14.5- §14.6) can be utilized.
N·jandNi·are sometimes called the row and column totals ormarginals , but we
will use these terms cautiously since we can never keep straight which are the rows
and which are the columns!
Thenullhypothesisisthatthetwovariables xandyhavenoassociation. Inthis
case, the probability of a particular value of xgiven a particular value of yshould
be the same as the probability of that value of xregardless of y. Therefore, in the
null hypothesis,the expectednumberfor any N ij, which we will denote nij, can be
calculated from only the row and column totals,
nij
N·j=Ni·
Nwhichimplies nij=Ni·N·j
N(14.4.2 )
Notice that if a column or row total is zero, then the expected number for all the
entries in that column or row is also zero; in that case, the never-occurringbin of x
oryshould simply be removed from the analysis.
Thechi-squarestatisticisnowgivenbyequation(14.3.1),which,inthepresent
case, is summed over all entries in the table,
χ2=/summationdisplay
i,j(Nij−nij)2
nij(14.4.3 )
Thenumberofdegreesoffreedomis equaltothenumberofentriesinthetable
(productof its row size and column size) minus the numberof constraints that have
arisen from our use of the data themselves to determinethe nij. Each row total and
columntotal is a constraint,exceptthat this overcountsbyone,since thetotal ofthe
14.4ContingencyTableAnalysisofTwoDistributions 625Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).column totals and the total of the row totals both equal N, the total number of data
points. Therefore,if thetable is ofsize IbyJ, the numberofdegreesoffreedomis
IJ−I−J+1. Equation (14.4.3), along with the chi-square probability function
(§6.2),now givethe signi ficanceof an associationbetween the variables xandy.
Suppose there is a signi ficant association. How do we quantify its strength, so
that (e.g.) we can compare the strength of one association with another? The idea
here is to find some reparametrization of χ2which maps it into some convenient
interval,like0to1,wheretheresultis notdependentonthequantityofdatathatwe
happento sample,butratherdependsonlyontheunderlyingpopulationfromwhich
thedataweredrawn. Thereareseveraldifferentwaysofdoingthis. Twoofthemorecommon are called Cramer’s V and thecontingency coefficient C .
The formula for Cramer ’sVis
V=/radicalBigg
χ2
Nmin (I−1,J−1)(14.4.4 )
where IandJare again the numbers of rows and columns, and Nis the total
number of events. Cramer ’sVhas the pleasant property that it lies between zero
and one inclusive, equals zero when there is no association, and equals one only
when the association is perfect: All the events in anyrow lie in one uniquecolumn,and vice versa. (In chess parlance, no two rooks, placed on a nonzero table entry,
can capture each other.)
In the case of I=J=2, Cramer’sVis also referredto as the phistatistic.
The contingency coef ficient Cis defined as
C=/radicalBigg
χ2
χ2+N(14.4.5 )
It also lies between zero and one, but (as is apparent from the formula) it can never
achievethe upperlimit. While it can be used to comparethe strengthof associationof two tables with the same IandJ, its upper limit depends on IandJ. Therefore
it can never be used to compare tables of different sizes.
ThetroublewithbothCramer ’sVandthecontingencycoef ficient Cisthat,when
they take on values in between their extremes, there is no very direct interpretation
of what that value means. For example,you are in Las Vegas, and a friend tells youthat there is a small, but signi ficant, association between the color of a croupier ’s
eyesandthe occurrenceofredandblackonhisroulettewheel. Cramer ’sVis about
0.028,yourfriendtellsyou. Youknowwhattheusualoddsagainstyouare(becauseof the green zero and double zero on the wheel). Is this association suf ficient for
you to make money? Don ’t ask us!
SUBROUTINE cntab1(nn,ni,nj,chisq,df,prob,cramrv,ccc)
INTEGER ni,nj,nn(ni,nj),MAXI,MAXJ
REAL ccc,chisq,cramrv,df,prob,TINY
PARAMETER (MAXI=100,MAXJ=100,TINY=1.e-30) Maximum table size, and a small num-
ber. C USES gammq
Givenatwo-dimensionalcontingency tableintheformofanintegerarray nn(1:ni,1:nj) ,
thisroutine returns the chi-square chisq,the number ofdegrees offreedom df,thesignif-
icance level prob(small values indicating a significant association), and two measures of
association, Cramer’s V(cramrv) and the contingency coefficient C(ccc).
INTEGER i,j,nni,nnj
626 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).REAL expctd,sum,sumi(MAXI),sumj(MAXJ),gammq
sum=0 Will be total number of events.
nni=ni Number of rows
nnj=nj and columns.
do12i=1,ni Get the row totals.
sumi(i)=0.
do11j=1,nj
sumi(i)=sumi(i)+nn(i,j)sum=sum+nn(i,j)
enddo
11
if(sumi(i).eq.0.)nni=nni-1 Eliminate any zero rows by reducing the
number. enddo 12
do14j=1,nj Get the column totals.
sumj(j)=0.
do13i=1,ni
sumj(j)=sumj(j)+nn(i,j)
enddo 13
if(sumj(j).eq.0.)nnj=nnj-1 Eliminate any zero columns.
enddo 14
df=nni*nnj-nni-nnj+1 Correctednumberofdegreesoffreedom.
chisq=0.
do16i=1,ni Do the chi-square sum.
do15j=1,nj
expctd=sumj(j)*sumi(i)/sum
chisq=chisq+(nn(i,j)-expctd)**2/(expctd+TINY) Here TINYguarantees that
any eliminated row or column will notcontribute to the sum.enddo
15
enddo 16
prob=gammq(0.5*df,0.5*chisq) Chi-square probability function.
cramrv=sqrt(chisq/(sum*min(nni-1,nnj-1)))ccc=sqrt(chisq/(chisq+sum))
return
END
Measures of Association Based onEntropy
Consider the game of “twenty questions, ”where by repeated yes/no questions
youtry to eliminate all exceptone correctpossibility for an unknownobject. Better
yet, consider a generalization of the game, where you are allowed to ask multiple
choice questions as well as binary (yes/no) ones. The categories in your multiple
choicequestionsaresupposedtobemutuallyexclusiveandexhaustive(asare “yes”
and“no”).
The value to you of an answer increases with the number of possibilities that
it eliminates. More speci fically, an answer that eliminates all except a fraction pof
the remaining possibilities can be assigned a value −lnp(a positive number, since
p< 1). The purposeof the logarithmis to make the value additive,since (e.g.) one
question that eliminates all but 1/6 of the possibilities is considered as good as twoquestions that, in sequence, reduce the number by factors 1/2 and 1/3.
So that is the value of an answer; but what is the value of a question? If there
areIpossibleanswerstothequestion (i=1,...,I )andthefractionofpossibilities
consistent with the ith answer is p
i(with the sum of the pi’s equal to one), then the
valueofthe questionis theexpectationvalueofthevalueoftheanswer,denoted H,
H=−I/summationdisplay
i=1pilnpi (14.4.6 )
14.4ContingencyTableAnalysisofTwoDistributions 627Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).In evaluating (14.4.6), note that
lim
p→0plnp=0 ( 14.4.7 )
The value Hlies between 0 and lnI. It is zero only when one of the pi’s is one, all
theotherszero: Inthiscase,thequestionisvalueless,sinceitsanswerispreordained.
Htakesonitsmaximumvaluewhenallthe pi’sareequal,inwhichcasethequestion
is sure to eliminate all but a fraction 1/Iof the remaining possibilities.
The value His conventionallytermed the entropyof the distribution given by
thepi’s, a terminology borrowed from statistical physics.
Sofarwehavesaid nothingabouttheassociationoftwovariables;butsuppose
we are deciding what question to ask next in the game and have to choose between
two candidates, or possibly want to ask both in one order or another. Suppose that
one question, x, has Ipossible answers, labeled by i, and that the other question,
y,a sJpossible answers, labeled by j. Then the possible outcomes of asking both
questionsformacontingencytablewhoseentries N ij,whennormalizedbydividing
by the total number of remaining possibilities N, give all the informationabout the
p’s. In particular,we can make contact with the notation(14.4.1)by identifying
pij=Nij
N
pi·=Ni·
N(outcomesofquestion xalone)
p·j=N·j
N(outcomesofquestion yalone)(14.4.8 )
The entropies of the questions xandyare, respectively,
H(x)=−/summationdisplay
ipi·lnpi· H(y)=−/summationdisplay
jp·jlnp·j (14.4.9 )
The entropy of the two questions together is
H(x, y )=−/summationdisplay
i,jpijlnpij (14.4.10 )
Now what is the entropy of the question ygiven x(that is, if xis askedfirst)?
It is the expectation value over the answers to xof the entropy of the restricted
ydistribution that lies in a single column of the contingency table (corresponding
to the xanswer):
H(y|x)=−/summationdisplay
ipi·/summationdisplay
jpij
pi·lnpij
pi·=−/summationdisplay
i,jpijlnpij
pi·(14.4.11 )
Correspondingly, the entropy of xgiven yis
H(x|y)=−/summationdisplay
jp·j/summationdisplay
ipij
p·jlnpij
p·j=−/summationdisplay
i,jpijlnpij
p·j(14.4.12 )
628 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).We can readily prove that the entropy of ygiven xis never more than the
entropy of yalone, i.e., that asking xfirst can only reduce the usefulness of asking
y(in which case the two variables are associated !):
H(y|x)−H(y)=−/summationdisplay
i,jpijlnpij/p i·
p·j
=/summationdisplay
i,jpijlnp·jpi·
pij
≤/summationdisplay
i,jpij/parenleftbiggp·jpi·
pij−1/parenrightbigg
=/summationdisplay
i,jpi·p·j−/summationdisplay
i,jpij
=1−1=0(14.4.13 )
where the inequality follows from the fact
lnw≤w−1( 14.4.14 )
We nowhaveeverythingweneedtode fineameasureofthe “dependency ”ofy
onx, that is to say a measure of association. This measure is sometimes called the
uncertainty coefficient ofy. We will denote it as U(y|x),
U(y|x)≡H(y)−H(y|x)
H(y)(14.4.15 )
This measure lies between zero and one, with the value 0 indicating that xandy
have no association, the value 1 indicating that knowledge of xcompletely predicts
y. For in-between values, U(y|x)gives the fraction of y’s entropy H(y)that is
lost if xis already known (i.e., that is redundant with the information in x). In our
game of“twenty questions, ”U(y|x)is the fractional loss in the utility of question
yif question xis to be asked first.
If we wish to view xas the dependentvariable, yas the independentone, then
interchanging xandywe can of course de fine the dependencyof xony,
U(x|y)≡H(x)−H(x|y)
H(x)(14.4.16 )
If we want to treat xandysymmetrically, then the useful combination turns
out to be
U(x, y )≡2/bracketleftbiggH(y)+H(x)−H(x, y )
H(x)+H(y)/bracketrightbigg
(14.4.17 )
If the two variables are completely independent, then H(x, y )=H(x)+H(y),s o
(14.4.17) vanishes. If the two variables are completely dependent, then H(x)=
H(y)=H(x, y ),so(14.4.16)equalsunity. Infact,youcanusetheidentities(easily
proved from equations 14.4.9 –14.4.12)
H(x, y )=H(x)+H(y|x)=H(y)+H(x|y)( 14.4.18 )
14.4ContingencyTableAnalysisofTwoDistributions 629Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).to show that
U(x, y )=H(x)U(x|y)+H(y)U(y|x)
H(x)+H(y)(14.4.19 )
i.e.,thatthesymmetricalmeasureisjustaweightedaverageofthetwoasymmetrical
measures(14.4.15)and(14.4.16),weightedbytheentropyofeachvariableseparately.
Here is a program for computing all the quantities discussed, H(x),H(y),
H(x|y),H(y|x),H(x, y ),U(x|y),U(y|x),andU(x, y ):
SUBROUTINE cntab2(nn,ni,nj,h,hx,hy,hygx,hxgy,uygx,uxgy,uxy)
INTEGER ni,nj,nn(ni,nj),MAXI,MAXJ
REAL h,hx,hxgy,hy,hygx,uxgy,uxy,uygx,TINYPARAMETER (MAXI=100,MAXJ=100,TINY=1.e-30)
Given atwo-dimensional contingency table inthe formof aninteger array
nn(i,j),w h e re
ilabelsthe xvariableandrangesfrom1to ni,jlabelsthe yvariableandrangesfrom1to
nj,thisroutinereturnstheentropy hofthewholetable,theentropy hxofthe xdistribution,
the entropy hyof the ydistribution, the entropy hygxofygiven x,t h ee n t r o p y hxgyof
xgiven y, the dependency uygxofyon x(eq. 14.4.15), the dependency uxgyofxon y
(eq. 14.4.16), and the symmetrical dependency uxy(eq. 14.4.17).
Parameters: MAXIandMAXJdefine the maximum size of table; TINYis a small number.
INTEGER i,j
REAL p,sum,sumi(MAXI),sumj(MAXJ)
sum=0do
12i=1,ni Get the row totals.
sumi(i)=0.0
do11j=1,nj
sumi(i)=sumi(i)+nn(i,j)sum=sum+nn(i,j)
enddo
11
enddo 12
do14j=1,nj Get the column totals.
sumj(j)=0.
do13i=1,ni
sumj(j)=sumj(j)+nn(i,j)
enddo 13
enddo 14
hx=0. Entropy of the xdistribution,
do15i=1,ni
if(sumi(i).ne.0.)then
p=sumi(i)/sum
hx=hx-p*log(p)
endif
enddo 15
hy=0. and of the ydistribution.
do16j=1,nj
if(sumj(j).ne.0.)then
p=sumj(j)/sum
hy=hy-p*log(p)
endif
enddo 16
h=0.do
18i=1,ni Total entropy: loop over both x
do17j=1,nj and y.
if(nn(i,j).ne.0)then
p=nn(i,j)/sumh=h-p*log(p)
endif
enddo
17
enddo 18
hygx=h-hx Uses equation (14.4.18),
hxgy=h-hy as does this.
630 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).uygx=(hy-hygx)/(hy+TINY) Equation (14.4.15).
uxgy=(hx-hxgy)/(hx+TINY) Equation (14.4.16).
uxy=2.*(hx+hy-h)/(hx+hy+TINY) Equation (14.4.17).
returnEND
CITED REFERENCES AND FURTHER READING:
Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New
York: Wiley).
Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS-
X Advanced Statistics Guide (New York: McGraw-Hill).
Fano, R.M. 1961, Transmission of Information (NewYork: Wileyand MIT Press), Chapter 2.
14.5 Linear Correlation
We next turn to measures of association between variables that are ordinal
or continuous, rather than nominal. Most widely used is the linear correlation
coefficient . For pairs of quantities (xi,y i),i =1,...,N, the linear correlation
coefficient r(also called the product-moment correlation coef ficient, orPearson’s
r) is given by the formula
r=/summationtext
i(xi−x)(yi−y)
/radicalbigg/summationtext
i(xi−x)2/radicalbigg/summationtext
i(yi−y)2(14.5.1 )
where, as usual, xis the mean of the xi’s,yis the mean of the yi’s.
Thevalueof rliesbetween −1and 1,inclusive. Ittakesonavalueof 1,termed
“complete positive correlation, ”when the data points lie on a perfect straight line
withpositiveslope,with xandyincreasingtogether. Thevalue 1holdsindependent
of the magnitude of the slope. If the data points lie on a perfect straight line withnegative slope, ydecreasing as xincreases, then rhas the value −1; this is called
“complete negative correlation. ”A value of rnear zero indicates that the variables
xandyareuncorrelated .
When a correlation is known to be signi ficant, ris one conventional way of
summarizing its strength. In fact, the value of rcan be translated into a statement
aboutwhat residuals(rootmeansquaredeviations)areto beexpectedifthe dataare
fitted to a straight line by the least-squares method (see §15.2, especially equations
15.2.13–15.2.14). Unfortunately, ris a rather poor statistic for deciding whether
an observed correlation is statistically signi ficant, and/or whether one observed
correlation is signi ficantly stronger than another. The reason is that ris ignorant of
the individual distributions of xandy, so there is no universal way to compute its
distribution in the case of the null hypothesis.
Abouttheonlygeneralstatementthatcanbemadeisthis: Ifthenullhypothesis
is that xandyare uncorrelated, and if the distributions for xandyeach have
enough convergent moments ( “tails”die off suf ficiently rapidly), and if Nis large