f14-5
PDF · 4 pages · 57.2 KB
Open PDF file
Excerpt from the published textbook Numerical Recipes in Fortran 77 (Cambridge University Press, 1986-1992), pp. 630-633. Section 14.5 covers Pearson's r, its significance via erfc and Student's t, Fisher's z-transformation, and tests comparing correlations, with the Fortran routine pearsn. It ends at the opening of section 14.6 on nonparametric rank correlation. This is reference material by others, not Phil's own work.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
630 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).uygx=(hy-hygx)/(hy+TINY) Equation (14.4.15).
uxgy=(hx-hxgy)/(hx+TINY) Equation (14.4.16).
uxy=2.*(hx+hy-h)/(hx+hy+TINY) Equation (14.4.17).
returnEND
CITED REFERENCES AND FURTHER READING:
Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New
York: Wiley).
Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS-
X Advanced Statistics Guide (New York: McGraw-Hill).
Fano, R.M. 1961, Transmission of Information (NewYork: Wileyand MIT Press), Chapter 2.
14.5 Linear Correlation
We next turn to measures of association between variables that are ordinal
or continuous, rather than nominal. Most widely used is the linear correlation
coefficient . For pairs of quantities (xi,y i),i =1,...,N, the linear correlation
coefficient r(also called the product-moment correlation coefficient, or Pearson’s
r) is given by the formula
r=/summationtext
i(xi−x)(yi−y)
/radicalbigg/summationtext
i(xi−x)2/radicalbigg/summationtext
i(yi−y)2(14.5.1 )
where, as usual, xis the mean of the xi’s,yis the mean of the yi’s.
Thevalueof rliesbetween −1and 1,inclusive. Ittakesonavalueof 1,termed
“complete positive correlation,” when the data points lie on a perfect straight line
withpositiveslope,with xandyincreasingtogether. Thevalue 1holdsindependent
of the magnitude of the slope. If the data points lie on a perfect straight line withnegative slope, ydecreasing as xincreases, then rhas the value −1; this is called
“complete negative correlation.” A value of rnear zero indicates that the variables
xandyareuncorrelated .
When a correlation is known to be significant, ris one conventional way of
summarizing its strength. In fact, the value of rcan be translated into a statement
aboutwhat residuals(rootmeansquaredeviations)areto beexpectedifthe dataare
fitted to a straight line by the least-squares method (see §15.2, especially equations
15.2.13 – 15.2.14). Unfortunately, ris a rather poor statistic for deciding whether
an observed correlation is statistically significant, and/or whether one observed
correlation is significantly stronger than another. The reason is that ris ignorant of
the individual distributions of xandy, so there is no universal way to compute its
distribution in the case of the null hypothesis.
Abouttheonlygeneralstatementthatcanbemadeisthis: Ifthenullhypothesis
is that xandyare uncorrelated, and if the distributions for xandyeach have
enough convergent moments (“tails” die off sufficiently rapidly), and if Nis large
14.5 LinearCorrelation 631Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).(typically >500),then ris distributedapproximatelynormally,withameanofzero
and a standard deviation of 1/√
N. In that case, the (double-sided) significance of
the correlation, that is, the probability that |r|should be larger than its observed
value in the null hypothesis, is
erfc/parenleftBigg
|r|√
N√
2/parenrightBigg
(14.5.2 )
where erfc (x)is the complementary error function, equation (6.2.8), computed by
the routines erfcorerfccof§6.2. A small value of (14.5.2) indicates that the
two distributions are significantly correlated. (See expression 14.5.9 below for a
more accurate test.)
Most statistics books try to go beyond (14.5.2) and give additional statistical
tests that can be made using r. In almost all cases, however, these tests are valid
only for a very special class of hypotheses, namely that the distributions of xandy
jointlyforma binormal ortwo-dimensionalGaussian distributionaroundtheirmean
values, with joint probability density
p(x, y )dxdy =const. ×exp/bracketleftbigg
−1
2(a11x2−2a12xy+a22y2)/bracketrightbigg
dxdy (14.5.3 )
where a11,a12,anda22arearbitraryconstants. Forthis distribution rhas thevalue
r=−a12√a11a22(14.5.4 )
There are occasions when (14.5.3) may be known to be a good model of the
data. There may be other occasions when we are willing to take (14.5.3)as at least
a rough and ready guess, since many two-dimensional distributions do resemble a
binormaldistribution,atleastnottoofaroutontheirtails. Ineithersituation,wecanuse (14.5.3) to go beyond (14.5.2) in any of several directions:
First, we can allow for the possibility that the number Nof data points is not
large. Here, it turns out that the statistic
t=r/radicalbigg
N−2
1−r2(14.5.5 )
is distributed in the null case (of no correlation) like Student’s t-distribution with
ν=N−2degrees of freedom, whose two-sided significance level is given by
1−A(t|ν)(equation 6.4.7). As Nbecomes large, this significance and (14.5.2)
become asymptotically the same, so that one never does worse by using (14.5.5),
even if the binormal assumption is not well substantiated.
Second, when Nis only moderately large ( ≥10), we can compare whether
the difference of two significantly nonzero r’s, e.g., from different experiments, is
itself significant. Inotherwords, we can quantifywhethera changein somecontrol
variablesignificantlyaltersanexistingcorrelationbetweentwoothervariables. Thisis done by using Fisher’s z-transformation to associate each measured rwith a
corresponding z,
z=1
2ln/parenleftbigg1+r
1−r/parenrightbigg
(14.5.6 )
632 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).Then, each zis approximately normally distributed with a mean value
z=1
2/bracketleftbigg
ln/parenleftbigg1+rtrue
1−rtrue/parenrightbigg
+rtrue
N−1/bracketrightbigg
(14.5.7 )
where rtrueis the actual or population value of the correlation coefficient, and with
a standard deviation
σ(z)≈1√
N−3(14.5.8 )
Equations (14.5.7) and (14.5.8), when they are valid, give several useful
statistical tests. For example, the significance level at which a measured value of r
differs from some hypothesized value rtrueis given by
erfc/parenleftbigg|z−z|√
N−3√
2/parenrightbigg
(14.5.9 )
where zandzare given by (14.5.6) and (14.5.7), with small values of (14.5.9)
indicating a significant difference. (Setting z=0makes expression 14.5.9 a more
accurate replacement for expression 14.5.2 above.) Similarly, the significance of a
difference between two measured correlation coefficients r1andr2is
erfc
|z1−z2|
√
2/radicalBig
1
N1−3+1
N2−3
(14.5.10 )
where z1andz2are obtained from r1andr2using (14.5.6), and where N1andN2
are, respectively, the numberof data points in the measurementof r1andr2.
All of the significances above are two-sided. If you wish to disprove the null
hypothesisin favorofa one-sidedhypothesis,such as that r1>r 2(wherethe sense
of the inequality was decided a priori), then (i) if your measured r1andr2have
thewrongsense, you have failed to demonstrate your one-sided hypothesis, but (ii)
if they have the right ordering, you can multiply the significances given above by
0.5, which makes them more significant.
But keep in mind: These interpretations of the rstatistic can be completely
meaningless if the joint probability distribution of your variables xandyis too
different from a binormal distribution.
SUBROUTINE pearsn(x,y,n,r,prob,z)
INTEGER nREAL prob,r,z,x(n),y(n),TINY
PARAMETER (TINY=1.e-20) Willregularize the unusual case of com-
plete correlation.
C USES betai
Given two arrays x(1:n)andy(1:n), this routine computes their correlation coefficient
r(returned as r), the significance level at which the null hypothesis of zero correlation
is disproved ( probwhose small value indicates a significant correlation), and Fisher’s z
(returned as z), whose valuecanbe used infurther statistical tests as described above.
INTEGER j
REAL ax,ay,df,sxx,sxy,syy,t,xt,yt,betai
ax=0.ay=0.
do
11j=1,n Find the means.
14.6 NonparametricorRankCorrelation 633Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).ax=ax+x(j)
ay=ay+y(j)
enddo 11
ax=ax/n
ay=ay/n
sxx=0.
syy=0.sxy=0.do
12j=1,n Compute the correlation coefficient.
xt=x(j)-ax
yt=y(j)-aysxx=sxx+xt**2syy=syy+yt**2
sxy=sxy+xt*yt
enddo
12
r=sxy/(sqrt(sxx*syy)+TINY)
z=0.5*log(((1.+r)+TINY)/((1.-r)+TINY)) Fisher’s ztransformation.
df=n-2t=r*sqrt(df/(((1.-r)+TINY)*((1.+r)+TINY))) Equation (14.5.5).
prob=betai(0.5*df,0.5,df/(df+t**2)) Student’s tprobability.
C prob=erfcc(abs(z*sqrt(n-1.))/1.4142136) For large n, this easier computation of
prob,usingtheshortroutine erfcc,
would give approximately the same
value.return
END
CITED REFERENCES AND FURTHER READING:
Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New
York: Wiley).
Hoel, P.G. 1971, Introduction to Mathematical Statistics , 4th ed. (New York: Wiley),Chapter 7.
von Mises, R. 1964, Mathematical Theory of Probability and Statistics (New York: Academic
Press), Chapters IX(A) and IX(B).
Korn, G.A., andKorn, T.M. 1968, Mathematical Handbookfor Scientists andEngineers , 2nded.
(New York: McGraw-Hill),§19.7.
Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS-
X Advanced Statistics Guide (New York: McGraw-Hill).
14.6 Nonparametric or Rank Correlation
It is precisely the uncertainty in interpreting the significance of the linear
correlationcoefficient rthat leads us to the importantconceptsof nonparametric or
rankcorrelation . Asbefore,wearegiven Npairsofmeasurements (xi,y i). Before,
difficulties arose because we did not necessarily know the probability distribution
function from which the xi’s or yi’s were drawn.
The key concept of nonparametric correlation is this: If we replace the value
of each xiby the value of its rankamong all the other xi’s in the sample, that
is,1,2,3,...,N, then the resulting list of numbers will be drawn from a perfectly
known distribution function, namely uniformlyfrom the integers between 1andN,
inclusive. Better than uniformly, in fact, since if the xi’s are all distinct, then each
integer will occur precisely once. If some of the xi’s have identical values, it is
conventionalto assign to all these “ties” the meanof the ranksthat theywould have
had if their values had been slightly different. This midrankwill sometimes be an