Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Scheid and numerical / Numerical Recipes in Fortran

f14-5

PDF · 4 pages · 57.2 KB
Open PDF file

Excerpt from the published textbook Numerical Recipes in Fortran 77 (Cambridge University Press, 1986-1992), pp. 630-633. Section 14.5 covers Pearson's r, its significance via erfc and Student's t, Fisher's z-transformation, and tests comparing correlations, with the Fortran routine pearsn. It ends at the opening of section 14.6 on nonparametric rank correlation. This is reference material by others, not Phil's own work.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
630 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).uygx=(hy-hygx)/(hy+TINY) Equation (14.4.15). uxgy=(hx-hxgy)/(hx+TINY) Equation (14.4.16). uxy=2.*(hx+hy-h)/(hx+hy+TINY) Equation (14.4.17). returnEND CITED REFERENCES AND FURTHER READING: Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New York: Wiley). Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS- X Advanced Statistics Guide (New York: McGraw-Hill). Fano, R.M. 1961, Transmission of Information (NewYork: Wileyand MIT Press), Chapter 2. 14.5 Linear Correlation We next turn to measures of association between variables that are ordinal or continuous, rather than nominal. Most widely used is the linear correlation coefficient . For pairs of quantities (xi,y i),i =1,...,N, the linear correlation coefficient r(also called the product-moment correlation coefficient, or Pearson’s r) is given by the formula r=/summationtext i(xi−x)(yi−y) /radicalbigg/summationtext i(xi−x)2/radicalbigg/summationtext i(yi−y)2(14.5.1 ) where, as usual, xis the mean of the xi’s,yis the mean of the yi’s. Thevalueof rliesbetween −1and 1,inclusive. Ittakesonavalueof 1,termed “complete positive correlation,” when the data points lie on a perfect straight line withpositiveslope,with xandyincreasingtogether. Thevalue 1holdsindependent of the magnitude of the slope. If the data points lie on a perfect straight line withnegative slope, ydecreasing as xincreases, then rhas the value −1; this is called “complete negative correlation.” A value of rnear zero indicates that the variables xandyareuncorrelated . When a correlation is known to be significant, ris one conventional way of summarizing its strength. In fact, the value of rcan be translated into a statement aboutwhat residuals(rootmeansquaredeviations)areto beexpectedifthe dataare fitted to a straight line by the least-squares method (see §15.2, especially equations 15.2.13 – 15.2.14). Unfortunately, ris a rather poor statistic for deciding whether an observed correlation is statistically significant, and/or whether one observed correlation is significantly stronger than another. The reason is that ris ignorant of the individual distributions of xandy, so there is no universal way to compute its distribution in the case of the null hypothesis. Abouttheonlygeneralstatementthatcanbemadeisthis: Ifthenullhypothesis is that xandyare uncorrelated, and if the distributions for xandyeach have enough convergent moments (“tails” die off sufficiently rapidly), and if Nis large 14.5 LinearCorrelation 631Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).(typically >500),then ris distributedapproximatelynormally,withameanofzero and a standard deviation of 1/√ N. In that case, the (double-sided) significance of the correlation, that is, the probability that |r|should be larger than its observed value in the null hypothesis, is erfc/parenleftBigg |r|√ N√ 2/parenrightBigg (14.5.2 ) where erfc (x)is the complementary error function, equation (6.2.8), computed by the routines erfcorerfccof§6.2. A small value of (14.5.2) indicates that the two distributions are significantly correlated. (See expression 14.5.9 below for a more accurate test.) Most statistics books try to go beyond (14.5.2) and give additional statistical tests that can be made using r. In almost all cases, however, these tests are valid only for a very special class of hypotheses, namely that the distributions of xandy jointlyforma binormal ortwo-dimensionalGaussian distributionaroundtheirmean values, with joint probability density p(x, y )dxdy =const. ×exp/bracketleftbigg −1 2(a11x2−2a12xy+a22y2)/bracketrightbigg dxdy (14.5.3 ) where a11,a12,anda22arearbitraryconstants. Forthis distribution rhas thevalue r=−a12√a11a22(14.5.4 ) There are occasions when (14.5.3) may be known to be a good model of the data. There may be other occasions when we are willing to take (14.5.3)as at least a rough and ready guess, since many two-dimensional distributions do resemble a binormaldistribution,atleastnottoofaroutontheirtails. Ineithersituation,wecanuse (14.5.3) to go beyond (14.5.2) in any of several directions: First, we can allow for the possibility that the number Nof data points is not large. Here, it turns out that the statistic t=r/radicalbigg N−2 1−r2(14.5.5 ) is distributed in the null case (of no correlation) like Student’s t-distribution with ν=N−2degrees of freedom, whose two-sided significance level is given by 1−A(t|ν)(equation 6.4.7). As Nbecomes large, this significance and (14.5.2) become asymptotically the same, so that one never does worse by using (14.5.5), even if the binormal assumption is not well substantiated. Second, when Nis only moderately large ( ≥10), we can compare whether the difference of two significantly nonzero r’s, e.g., from different experiments, is itself significant. Inotherwords, we can quantifywhethera changein somecontrol variablesignificantlyaltersanexistingcorrelationbetweentwoothervariables. Thisis done by using Fisher’s z-transformation to associate each measured rwith a corresponding z, z=1 2ln/parenleftbigg1+r 1−r/parenrightbigg (14.5.6 ) 632 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).Then, each zis approximately normally distributed with a mean value z=1 2/bracketleftbigg ln/parenleftbigg1+rtrue 1−rtrue/parenrightbigg +rtrue N−1/bracketrightbigg (14.5.7 ) where rtrueis the actual or population value of the correlation coefficient, and with a standard deviation σ(z)≈1√ N−3(14.5.8 ) Equations (14.5.7) and (14.5.8), when they are valid, give several useful statistical tests. For example, the significance level at which a measured value of r differs from some hypothesized value rtrueis given by erfc/parenleftbigg|z−z|√ N−3√ 2/parenrightbigg (14.5.9 ) where zandzare given by (14.5.6) and (14.5.7), with small values of (14.5.9) indicating a significant difference. (Setting z=0makes expression 14.5.9 a more accurate replacement for expression 14.5.2 above.) Similarly, the significance of a difference between two measured correlation coefficients r1andr2is erfc |z1−z2| √ 2/radicalBig 1 N1−3+1 N2−3  (14.5.10 ) where z1andz2are obtained from r1andr2using (14.5.6), and where N1andN2 are, respectively, the numberof data points in the measurementof r1andr2. All of the significances above are two-sided. If you wish to disprove the null hypothesisin favorofa one-sidedhypothesis,such as that r1>r 2(wherethe sense of the inequality was decided a priori), then (i) if your measured r1andr2have thewrongsense, you have failed to demonstrate your one-sided hypothesis, but (ii) if they have the right ordering, you can multiply the significances given above by 0.5, which makes them more significant. But keep in mind: These interpretations of the rstatistic can be completely meaningless if the joint probability distribution of your variables xandyis too different from a binormal distribution. SUBROUTINE pearsn(x,y,n,r,prob,z) INTEGER nREAL prob,r,z,x(n),y(n),TINY PARAMETER (TINY=1.e-20) Willregularize the unusual case of com- plete correlation. C USES betai Given two arrays x(1:n)andy(1:n), this routine computes their correlation coefficient r(returned as r), the significance level at which the null hypothesis of zero correlation is disproved ( probwhose small value indicates a significant correlation), and Fisher’s z (returned as z), whose valuecanbe used infurther statistical tests as described above. INTEGER j REAL ax,ay,df,sxx,sxy,syy,t,xt,yt,betai ax=0.ay=0. do 11j=1,n Find the means. 14.6 NonparametricorRankCorrelation 633Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).ax=ax+x(j) ay=ay+y(j) enddo 11 ax=ax/n ay=ay/n sxx=0. syy=0.sxy=0.do 12j=1,n Compute the correlation coefficient. xt=x(j)-ax yt=y(j)-aysxx=sxx+xt**2syy=syy+yt**2 sxy=sxy+xt*yt enddo 12 r=sxy/(sqrt(sxx*syy)+TINY) z=0.5*log(((1.+r)+TINY)/((1.-r)+TINY)) Fisher’s ztransformation. df=n-2t=r*sqrt(df/(((1.-r)+TINY)*((1.+r)+TINY))) Equation (14.5.5). prob=betai(0.5*df,0.5,df/(df+t**2)) Student’s tprobability. C prob=erfcc(abs(z*sqrt(n-1.))/1.4142136) For large n, this easier computation of prob,usingtheshortroutine erfcc, would give approximately the same value.return END CITED REFERENCES AND FURTHER READING: Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New York: Wiley). Hoel, P.G. 1971, Introduction to Mathematical Statistics , 4th ed. (New York: Wiley),Chapter 7. von Mises, R. 1964, Mathematical Theory of Probability and Statistics (New York: Academic Press), Chapters IX(A) and IX(B). Korn, G.A., andKorn, T.M. 1968, Mathematical Handbookfor Scientists andEngineers , 2nded. (New York: McGraw-Hill),§19.7. Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS- X Advanced Statistics Guide (New York: McGraw-Hill). 14.6 Nonparametric or Rank Correlation It is precisely the uncertainty in interpreting the significance of the linear correlationcoefficient rthat leads us to the importantconceptsof nonparametric or rankcorrelation . Asbefore,wearegiven Npairsofmeasurements (xi,y i). Before, difficulties arose because we did not necessarily know the probability distribution function from which the xi’s or yi’s were drawn. The key concept of nonparametric correlation is this: If we replace the value of each xiby the value of its rankamong all the other xi’s in the sample, that is,1,2,3,...,N, then the resulting list of numbers will be drawn from a perfectly known distribution function, namely uniformlyfrom the integers between 1andN, inclusive. Better than uniformly, in fact, since if the xi’s are all distinct, then each integer will occur precisely once. If some of the xi’s have identical values, it is conventionalto assign to all these “ties” the meanof the ranksthat theywould have had if their values had been slightly different. This midrankwill sometimes be an