Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Scheid and numerical / Numerical Recipes in Fortran

f14-4

PDF · 9 pages · 83.2 KB
Open PDF file

Excerpt from the textbook Numerical Recipes in Fortran 77 (Cambridge University Press, 1986-1992), starting at page 622 of Chapter 14, Statistical Description of Data. It closes section 14.3 on the Kolmogorov-Smirnov test, then begins 14.4 on contingency tables for nominal variables, using chi-square, degrees of freedom, and Cramer's V and the contingency coefficient. It is a published book sample, not Phil's own work.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
622 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).but, because of its cumulative nature, a K–S test would require many data points in the notch before signaling a discrepancy. Second,weshouldnotethat,ifyouestimateanyparametersfromadataset(e.g.,amean andvariance),thenthedistributionoftheK–Sstatistic Dforacumulativedistributionfunction P(x)thatuses the estimated parameters is no longer given by equation (14.3.9). In general, you will have to determine the new distribution yourself, e.g., by Monte Carlo methods. CITED REFERENCES AND FURTHER READING: von Mises, R. 1964, Mathematical Theory of Probability and Statistics (New York: Academic Press), Chapters IX(C) and IX(E). Stephens, M.A. 1970, Journalof the RoyalStatistical Society , ser. B, vol. 32, pp. 115–122.[1] Anderson,T.W.,andDarling,D.A.1952, AnnalsofMathematicalStatistics , vol.23,pp.193–212. [2] Darling, D.A. 1957, Annals of Mathematical Statistics , vol. 28, pp. 823–838. [3] Michael, J.R. 1983, Biometrika , vol. 70, no. 1, pp. 11–17. [4] No´e, M. 1972, Annals of Mathematical Statistics , vol. 43, pp. 58–64. [5] Kuiper,N.H. 1962, Proceedingsof theKoninklijkeNederlandseAkademie vanWetenschappen , ser. A., vol. 63, pp. 38–47. [6] Stephens, M.A. 1965, Biometrika , vol. 52, pp. 309–321. [7] Fisher, N.I., Lewis, T., and Embleton, B.J.J. 1987, Statistical Analysis of Spherical Data (New York: Cambridge University Press). [8] 14.4 Contingency Table Analysis of Two Distributions Inthissection,andthenexttwosections,wedealwith measuresofassociation fortwodistributions. Thesituationisthis: Eachdatapointhastwoormoredifferent quantitiesassociatedwithit,andwewanttoknowwhetherknowledgeofonequantity gives us any demonstrableadvantage in predicting the value of another quantity. In manycases,onevariablewillbean“independent”or“control”variable,andanotherwill be a “dependent” or “measured” variable. Then, we want to know if the latter variableisin fact dependent on or associated with the former variable. If it is, we wanttohavesomequantitativemeasureofthestrengthoftheassociation. Oneoftenhears this loosely stated as the question of whether two variables are correlated or uncorrelated , but we will reserve those terms for a particular kind of association (linear, or at least monotonic), as discussed in §14.5 and §14.6. Notice that, as in previous sections, the different concepts of significance and strength appear: The association between two distributions may be very significanteven if that association is weak — if the quantity of data is large enough. Itisusefultodistinguishamongsomedifferentkindsofvariables,withdifferent categories forming a loose hierarchy. •Avariableiscalled nominalifitsvaluesarethemembersofsomeunordered set. For example, “state of residence” is a nominal variable that (in the U.S.) takes on one of 50 values; in astrophysics, “type of galaxy” is a nominalvariablewiththethreevalues“spiral,”“elliptical,”and“irregular.” 14.4ContingencyTableAnalysisofTwoDistributions 623Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).•Avariableistermed ordinalifitsvaluesarethemembersofadiscrete,but ordered,set. Examplesare: gradein school,planetaryorderfromtheSun(Mercury = 1, Venus = 2, ...), number of offspring. There need not be any concept of “equal metric distance” between the values of an ordinal variable, only that they be intrinsically ordered. •We will call a variable continuous if its values are real numbers, as are times, distances, temperatures, etc. (Social scientists sometimes distinguishbetween intervalandratiocontinuousvariables,butwedonot find that distinction very compelling.) A continuous variable can always be made into an ordinal one by binning it into ranges. If we chooseto ignorethe orderingof the bins, then we can turn it into a nominal variable. Nominal variables constitute the lowest type of the hierarchy, and therefore the most general. For example, a set of severalcontinuous or ordinal variables can be turned, if crudely, into a single nominal variable, by coarsely binning each variable and then taking each distinct combinationof bin assignments as a single nominal value. When multidimensional data are sparse, this is oftenthe only sensible way to proceed. The remainder of this section will deal with measures of association between nominalvariables. For any pair of nominal variables, the data can be displayed as acontingency table , a table whose rows are labeled by the values of one nominal variable, whose columns are labeled by the values of the other nominal variable,and whose entries are nonnegative integers giving the number of observed events for each combination of row and column (see Figure 14.4.1). The analysis of association between nominal variables is thus called contingency table analysis or crosstabulation analysis. We will introduce two different approaches. The first approach, based on the chi-squarestatistic,doesagoodjobofcharacterizingthesignificanceofassociation, but is only so-so as a measure of the strength (principally because its numerical values have no very direct interpretations). The second approach, based on theinformation-theoreticconceptof entropy,saysnothingatallaboutthesignificanceof association (usechi-squareforthat!),but is capableof veryelegantlycharacterizing the strength of an association already known to be significant. Measures of Association Based onChi-Square Some notation first: Let Nijdenote the number of events that occur with the first variable xtaking on its ith value, and the second variable ytaking on its jth value. Let Ndenote the total number of events, the sum of all the N ij’s. Let Ni· denote the number of events for which the first variable xtakes on its ith value regardless of the value of y;N·jis the number of events with the jth value of y regardless of x.S o w e h a v e Ni·=/summationdisplay jNij N·j=/summationdisplay iNij N=/summationdisplay iNi·=/summationdisplay jN·j(14.4.1 ) 624 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).1. male 2. female .......... . .... . . .. . .. . .. . . 1. red # of red males N 11 # of red females N21# of green females N22# of green males N12# of males N1⋅ # of females N2⋅2. green # of red N⋅1# of green N⋅2total # N Figure 14.4.1. Example of a contingency table for two nominal variables, here sex and color. The row andcolumnmarginals(totals) areshown. Thevariables are “nominal,”i.e.,theorderinwhichtheir values are listed is arbitrary and does not affect the result of the contingency table analysis. If the ordering of values has some intrinsic meaning, then the variables are “ordinal”or“continuous, ”and correlation techniques ( §14.5- §14.6) can be utilized. N·jandNi·are sometimes called the row and column totals ormarginals , but we will use these terms cautiously since we can never keep straight which are the rows and which are the columns! Thenullhypothesisisthatthetwovariables xandyhavenoassociation. Inthis case, the probability of a particular value of xgiven a particular value of yshould be the same as the probability of that value of xregardless of y. Therefore, in the null hypothesis,the expectednumberfor any N ij, which we will denote nij, can be calculated from only the row and column totals, nij N·j=Ni· Nwhichimplies nij=Ni·N·j N(14.4.2 ) Notice that if a column or row total is zero, then the expected number for all the entries in that column or row is also zero; in that case, the never-occurringbin of x oryshould simply be removed from the analysis. Thechi-squarestatisticisnowgivenbyequation(14.3.1),which,inthepresent case, is summed over all entries in the table, χ2=/summationdisplay i,j(Nij−nij)2 nij(14.4.3 ) Thenumberofdegreesoffreedomis equaltothenumberofentriesinthetable (productof its row size and column size) minus the numberof constraints that have arisen from our use of the data themselves to determinethe nij. Each row total and columntotal is a constraint,exceptthat this overcountsbyone,since thetotal ofthe 14.4ContingencyTableAnalysisofTwoDistributions 625Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).column totals and the total of the row totals both equal N, the total number of data points. Therefore,if thetable is ofsize IbyJ, the numberofdegreesoffreedomis IJ−I−J+1. Equation (14.4.3), along with the chi-square probability function (§6.2),now givethe signi ficanceof an associationbetween the variables xandy. Suppose there is a signi ficant association. How do we quantify its strength, so that (e.g.) we can compare the strength of one association with another? The idea here is to find some reparametrization of χ2which maps it into some convenient interval,like0to1,wheretheresultis notdependentonthequantityofdatathatwe happento sample,butratherdependsonlyontheunderlyingpopulationfromwhich thedataweredrawn. Thereareseveraldifferentwaysofdoingthis. Twoofthemorecommon are called Cramer’s V and thecontingency coefficient C . The formula for Cramer ’sVis V=/radicalBigg χ2 Nmin (I−1,J−1)(14.4.4 ) where IandJare again the numbers of rows and columns, and Nis the total number of events. Cramer ’sVhas the pleasant property that it lies between zero and one inclusive, equals zero when there is no association, and equals one only when the association is perfect: All the events in anyrow lie in one uniquecolumn,and vice versa. (In chess parlance, no two rooks, placed on a nonzero table entry, can capture each other.) In the case of I=J=2, Cramer’sVis also referredto as the phistatistic. The contingency coef ficient Cis defined as C=/radicalBigg χ2 χ2+N(14.4.5 ) It also lies between zero and one, but (as is apparent from the formula) it can never achievethe upperlimit. While it can be used to comparethe strengthof associationof two tables with the same IandJ, its upper limit depends on IandJ. Therefore it can never be used to compare tables of different sizes. ThetroublewithbothCramer ’sVandthecontingencycoef ficient Cisthat,when they take on values in between their extremes, there is no very direct interpretation of what that value means. For example,you are in Las Vegas, and a friend tells youthat there is a small, but signi ficant, association between the color of a croupier ’s eyesandthe occurrenceofredandblackonhisroulettewheel. Cramer ’sVis about 0.028,yourfriendtellsyou. Youknowwhattheusualoddsagainstyouare(becauseof the green zero and double zero on the wheel). Is this association suf ficient for you to make money? Don ’t ask us! SUBROUTINE cntab1(nn,ni,nj,chisq,df,prob,cramrv,ccc) INTEGER ni,nj,nn(ni,nj),MAXI,MAXJ REAL ccc,chisq,cramrv,df,prob,TINY PARAMETER (MAXI=100,MAXJ=100,TINY=1.e-30) Maximum table size, and a small num- ber. C USES gammq Givenatwo-dimensionalcontingency tableintheformofanintegerarray nn(1:ni,1:nj) , thisroutine returns the chi-square chisq,the number ofdegrees offreedom df,thesignif- icance level prob(small values indicating a significant association), and two measures of association, Cramer’s V(cramrv) and the contingency coefficient C(ccc). INTEGER i,j,nni,nnj 626 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).REAL expctd,sum,sumi(MAXI),sumj(MAXJ),gammq sum=0 Will be total number of events. nni=ni Number of rows nnj=nj and columns. do12i=1,ni Get the row totals. sumi(i)=0. do11j=1,nj sumi(i)=sumi(i)+nn(i,j)sum=sum+nn(i,j) enddo 11 if(sumi(i).eq.0.)nni=nni-1 Eliminate any zero rows by reducing the number. enddo 12 do14j=1,nj Get the column totals. sumj(j)=0. do13i=1,ni sumj(j)=sumj(j)+nn(i,j) enddo 13 if(sumj(j).eq.0.)nnj=nnj-1 Eliminate any zero columns. enddo 14 df=nni*nnj-nni-nnj+1 Correctednumberofdegreesoffreedom. chisq=0. do16i=1,ni Do the chi-square sum. do15j=1,nj expctd=sumj(j)*sumi(i)/sum chisq=chisq+(nn(i,j)-expctd)**2/(expctd+TINY) Here TINYguarantees that any eliminated row or column will notcontribute to the sum.enddo 15 enddo 16 prob=gammq(0.5*df,0.5*chisq) Chi-square probability function. cramrv=sqrt(chisq/(sum*min(nni-1,nnj-1)))ccc=sqrt(chisq/(chisq+sum)) return END Measures of Association Based onEntropy Consider the game of “twenty questions, ”where by repeated yes/no questions youtry to eliminate all exceptone correctpossibility for an unknownobject. Better yet, consider a generalization of the game, where you are allowed to ask multiple choice questions as well as binary (yes/no) ones. The categories in your multiple choicequestionsaresupposedtobemutuallyexclusiveandexhaustive(asare “yes” and“no”). The value to you of an answer increases with the number of possibilities that it eliminates. More speci fically, an answer that eliminates all except a fraction pof the remaining possibilities can be assigned a value −lnp(a positive number, since p< 1). The purposeof the logarithmis to make the value additive,since (e.g.) one question that eliminates all but 1/6 of the possibilities is considered as good as twoquestions that, in sequence, reduce the number by factors 1/2 and 1/3. So that is the value of an answer; but what is the value of a question? If there areIpossibleanswerstothequestion (i=1,...,I )andthefractionofpossibilities consistent with the ith answer is p i(with the sum of the pi’s equal to one), then the valueofthe questionis theexpectationvalueofthevalueoftheanswer,denoted H, H=−I/summationdisplay i=1pilnpi (14.4.6 ) 14.4ContingencyTableAnalysisofTwoDistributions 627Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).In evaluating (14.4.6), note that lim p→0plnp=0 ( 14.4.7 ) The value Hlies between 0 and lnI. It is zero only when one of the pi’s is one, all theotherszero: Inthiscase,thequestionisvalueless,sinceitsanswerispreordained. Htakesonitsmaximumvaluewhenallthe pi’sareequal,inwhichcasethequestion is sure to eliminate all but a fraction 1/Iof the remaining possibilities. The value His conventionallytermed the entropyof the distribution given by thepi’s, a terminology borrowed from statistical physics. Sofarwehavesaid nothingabouttheassociationoftwovariables;butsuppose we are deciding what question to ask next in the game and have to choose between two candidates, or possibly want to ask both in one order or another. Suppose that one question, x, has Ipossible answers, labeled by i, and that the other question, y,a sJpossible answers, labeled by j. Then the possible outcomes of asking both questionsformacontingencytablewhoseentries N ij,whennormalizedbydividing by the total number of remaining possibilities N, give all the informationabout the p’s. In particular,we can make contact with the notation(14.4.1)by identifying pij=Nij N pi·=Ni· N(outcomesofquestion xalone) p·j=N·j N(outcomesofquestion yalone)(14.4.8 ) The entropies of the questions xandyare, respectively, H(x)=−/summationdisplay ipi·lnpi· H(y)=−/summationdisplay jp·jlnp·j (14.4.9 ) The entropy of the two questions together is H(x, y )=−/summationdisplay i,jpijlnpij (14.4.10 ) Now what is the entropy of the question ygiven x(that is, if xis askedfirst)? It is the expectation value over the answers to xof the entropy of the restricted ydistribution that lies in a single column of the contingency table (corresponding to the xanswer): H(y|x)=−/summationdisplay ipi·/summationdisplay jpij pi·lnpij pi·=−/summationdisplay i,jpijlnpij pi·(14.4.11 ) Correspondingly, the entropy of xgiven yis H(x|y)=−/summationdisplay jp·j/summationdisplay ipij p·jlnpij p·j=−/summationdisplay i,jpijlnpij p·j(14.4.12 ) 628 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).We can readily prove that the entropy of ygiven xis never more than the entropy of yalone, i.e., that asking xfirst can only reduce the usefulness of asking y(in which case the two variables are associated !): H(y|x)−H(y)=−/summationdisplay i,jpijlnpij/p i· p·j =/summationdisplay i,jpijlnp·jpi· pij ≤/summationdisplay i,jpij/parenleftbiggp·jpi· pij−1/parenrightbigg =/summationdisplay i,jpi·p·j−/summationdisplay i,jpij =1−1=0(14.4.13 ) where the inequality follows from the fact lnw≤w−1( 14.4.14 ) We nowhaveeverythingweneedtode fineameasureofthe “dependency ”ofy onx, that is to say a measure of association. This measure is sometimes called the uncertainty coefficient ofy. We will denote it as U(y|x), U(y|x)≡H(y)−H(y|x) H(y)(14.4.15 ) This measure lies between zero and one, with the value 0 indicating that xandy have no association, the value 1 indicating that knowledge of xcompletely predicts y. For in-between values, U(y|x)gives the fraction of y’s entropy H(y)that is lost if xis already known (i.e., that is redundant with the information in x). In our game of“twenty questions, ”U(y|x)is the fractional loss in the utility of question yif question xis to be asked first. If we wish to view xas the dependentvariable, yas the independentone, then interchanging xandywe can of course de fine the dependencyof xony, U(x|y)≡H(x)−H(x|y) H(x)(14.4.16 ) If we want to treat xandysymmetrically, then the useful combination turns out to be U(x, y )≡2/bracketleftbiggH(y)+H(x)−H(x, y ) H(x)+H(y)/bracketrightbigg (14.4.17 ) If the two variables are completely independent, then H(x, y )=H(x)+H(y),s o (14.4.17) vanishes. If the two variables are completely dependent, then H(x)= H(y)=H(x, y ),so(14.4.16)equalsunity. Infact,youcanusetheidentities(easily proved from equations 14.4.9 –14.4.12) H(x, y )=H(x)+H(y|x)=H(y)+H(x|y)( 14.4.18 ) 14.4ContingencyTableAnalysisofTwoDistributions 629Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).to show that U(x, y )=H(x)U(x|y)+H(y)U(y|x) H(x)+H(y)(14.4.19 ) i.e.,thatthesymmetricalmeasureisjustaweightedaverageofthetwoasymmetrical measures(14.4.15)and(14.4.16),weightedbytheentropyofeachvariableseparately. Here is a program for computing all the quantities discussed, H(x),H(y), H(x|y),H(y|x),H(x, y ),U(x|y),U(y|x),andU(x, y ): SUBROUTINE cntab2(nn,ni,nj,h,hx,hy,hygx,hxgy,uygx,uxgy,uxy) INTEGER ni,nj,nn(ni,nj),MAXI,MAXJ REAL h,hx,hxgy,hy,hygx,uxgy,uxy,uygx,TINYPARAMETER (MAXI=100,MAXJ=100,TINY=1.e-30) Given atwo-dimensional contingency table inthe formof aninteger array nn(i,j),w h e re ilabelsthe xvariableandrangesfrom1to ni,jlabelsthe yvariableandrangesfrom1to nj,thisroutinereturnstheentropy hofthewholetable,theentropy hxofthe xdistribution, the entropy hyof the ydistribution, the entropy hygxofygiven x,t h ee n t r o p y hxgyof xgiven y, the dependency uygxofyon x(eq. 14.4.15), the dependency uxgyofxon y (eq. 14.4.16), and the symmetrical dependency uxy(eq. 14.4.17). Parameters: MAXIandMAXJdefine the maximum size of table; TINYis a small number. INTEGER i,j REAL p,sum,sumi(MAXI),sumj(MAXJ) sum=0do 12i=1,ni Get the row totals. sumi(i)=0.0 do11j=1,nj sumi(i)=sumi(i)+nn(i,j)sum=sum+nn(i,j) enddo 11 enddo 12 do14j=1,nj Get the column totals. sumj(j)=0. do13i=1,ni sumj(j)=sumj(j)+nn(i,j) enddo 13 enddo 14 hx=0. Entropy of the xdistribution, do15i=1,ni if(sumi(i).ne.0.)then p=sumi(i)/sum hx=hx-p*log(p) endif enddo 15 hy=0. and of the ydistribution. do16j=1,nj if(sumj(j).ne.0.)then p=sumj(j)/sum hy=hy-p*log(p) endif enddo 16 h=0.do 18i=1,ni Total entropy: loop over both x do17j=1,nj and y. if(nn(i,j).ne.0)then p=nn(i,j)/sumh=h-p*log(p) endif enddo 17 enddo 18 hygx=h-hx Uses equation (14.4.18), hxgy=h-hy as does this. 630 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).uygx=(hy-hygx)/(hy+TINY) Equation (14.4.15). uxgy=(hx-hxgy)/(hx+TINY) Equation (14.4.16). uxy=2.*(hx+hy-h)/(hx+hy+TINY) Equation (14.4.17). returnEND CITED REFERENCES AND FURTHER READING: Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New York: Wiley). Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS- X Advanced Statistics Guide (New York: McGraw-Hill). Fano, R.M. 1961, Transmission of Information (NewYork: Wileyand MIT Press), Chapter 2. 14.5 Linear Correlation We next turn to measures of association between variables that are ordinal or continuous, rather than nominal. Most widely used is the linear correlation coefficient . For pairs of quantities (xi,y i),i =1,...,N, the linear correlation coefficient r(also called the product-moment correlation coef ficient, orPearson’s r) is given by the formula r=/summationtext i(xi−x)(yi−y) /radicalbigg/summationtext i(xi−x)2/radicalbigg/summationtext i(yi−y)2(14.5.1 ) where, as usual, xis the mean of the xi’s,yis the mean of the yi’s. Thevalueof rliesbetween −1and 1,inclusive. Ittakesonavalueof 1,termed “complete positive correlation, ”when the data points lie on a perfect straight line withpositiveslope,with xandyincreasingtogether. Thevalue 1holdsindependent of the magnitude of the slope. If the data points lie on a perfect straight line withnegative slope, ydecreasing as xincreases, then rhas the value −1; this is called “complete negative correlation. ”A value of rnear zero indicates that the variables xandyareuncorrelated . When a correlation is known to be signi ficant, ris one conventional way of summarizing its strength. In fact, the value of rcan be translated into a statement aboutwhat residuals(rootmeansquaredeviations)areto beexpectedifthe dataare fitted to a straight line by the least-squares method (see §15.2, especially equations 15.2.13–15.2.14). Unfortunately, ris a rather poor statistic for deciding whether an observed correlation is statistically signi ficant, and/or whether one observed correlation is signi ficantly stronger than another. The reason is that ris ignorant of the individual distributions of xandy, so there is no universal way to compute its distribution in the case of the null hypothesis. Abouttheonlygeneralstatementthatcanbemadeisthis: Ifthenullhypothesis is that xandyare uncorrelated, and if the distributions for xandyeach have enough convergent moments ( “tails”die off suf ficiently rapidly), and if Nis large