Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Scheid and numerical / Numerical Recipes in Fortran

f14-6

PDF · 8 pages · 94.0 KB
Open PDF file

Sample pages from section 14.6 of Numerical Recipes in Fortran 77 (Cambridge University Press, 1986-1992), Chapter 14 on statistical description of data. It ends the linear correlation routine, then covers rank correlation, midranks for ties, the Spearman coefficient, its t-test, the sum-squared rank difference D with mean and variance, and the Fortran routines spear and crank. Kendall's tau is introduced; the text shown stops partway.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
14.6 NonparametricorRankCorrelation 633Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).ax=ax+x(j) ay=ay+y(j) enddo 11 ax=ax/n ay=ay/n sxx=0. syy=0.sxy=0.do 12j=1,n Compute the correlation coefficient. xt=x(j)-ax yt=y(j)-aysxx=sxx+xt**2syy=syy+yt**2 sxy=sxy+xt*yt enddo 12 r=sxy/(sqrt(sxx*syy)+TINY) z=0.5*log(((1.+r)+TINY)/((1.-r)+TINY)) Fisher’s ztransformation. df=n-2t=r*sqrt(df/(((1.-r)+TINY)*((1.+r)+TINY))) Equation (14.5.5). prob=betai(0.5*df,0.5,df/(df+t**2)) Student’s tprobability. C prob=erfcc(abs(z*sqrt(n-1.))/1.4142136) For large n, this easier computation of prob, usingthe short routine erfcc, would give approximately the same value.return END CITED REFERENCES AND FURTHER READING: Dunn, O.J., andClark, V.A. 1974, AppliedStatistics: Analysis of Variance andRegression (New York: Wiley). Hoel, P.G. 1971, Introduction to Mathematical Statistics , 4th ed. (New York: Wiley),Chapter 7. von Mises, R. 1964, Mathematical Theory of Probability and Statistics (New York: Academic Press), Chapters IX(A) and IX(B). Korn, G.A., andKorn, T.M. 1968, Mathematical Handbookfor Scientists andEngineers , 2nded. (New York: McGraw-Hill),§19.7. Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS- X Advanced Statistics Guide (New York: McGraw-Hill). 14.6 Nonparametric or Rank Correlation It is precisely the uncertainty in interpreting the significance of the linear correlationcoefficient rthat leads us to the importantconceptsof nonparametric or rankcorrelation . Asbefore,wearegiven Npairsofmeasurements (xi,yi). Before, difficulties arose because we did not necessarily know the probability distribution function from which the xi’s or yi’s were drawn. The key concept of nonparametric correlation is this: If we replace the value of each xiby the value of its rankamong all the other xi’s in the sample, that is,1,2,3,...,N, then the resulting list of numbers will be drawn from a perfectly known distribution function, namely uniformlyfrom the integers between 1andN, inclusive. Better than uniformly, in fact, since if the xi’s are all distinct, then each integer will occur precisely once. If some of the xi’s have identical values, it is conventionalto assign to all these “ties” the meanof the ranksthat theywould have had if their values had been slightly different. This midrankwill sometimes be an 634 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).integer, sometimes a half-integer. In all cases the sum of all assigned ranks will be the same as the sum of the integers from 1toN, namely1 2N(N+1 ). Of course we do exactly the same procedure for the yi’s, replacing each value by its rank among the other yi’s in the sample. Now we are free to invent statistics for detecting correlation between uniform setsofintegersbetween 1andN,keepinginmindthepossibilityoftiesintheranks. There is, of course, some loss of information in replacing the original numbers by ranks. We couldconstructsome ratherartificial exampleswherea correlationcould bedetectedparametrically(e.g.,inthelinearcorrelationcoefficient r),butcouldnot be detected nonparametrically. Such examples are very rare in real life, however,and the slight loss of informationin rankingis a small price to pay for a very major advantage: When a correlation is demonstrated to be present nonparametrically, then it is really there! (That is, to a certainty level that depends on the significancechosen.) Nonparametric correlation is more robust than linear correlation, more resistant to unplanned defects in the data, in the same sort of sense that the median is more robustthan the mean. For more onthe conceptof robustness,see §15.7. As always in statistics, some particular choices of a statistic have already been invented for us and consecrated, if not beatified, by popular use. We will discuss two, theSpearmanrank-ordercorrelationcoefficient (r s), andKendall’stau (τ). Spearman Rank-Order CorrelationCoefficient LetRibe the rank of xiamong the other x’s,Sibe the rank of yiamong the other y’s, ties being assigned the appropriatemidrankas described above. Then the rank-order correlation coefficient is defined to be the linear correlation coefficient of the ranks, namely, rs=/summationtext i(Ri−R)(Si−S)/radicalBig/summationtext i(Ri−R)2/radicalBig/summationtext i(Si−S)2(14.6.1 ) The significance of a nonzero value of rsis tested by computing t=rs/radicalBigg N−2 1−r2s(14.6.2 ) which is distributed approximately as Student’s distribution with N−2degrees of freedom. A key point is that this approximation does not depend on the original distribution of the x’s and y’s; it is always the same approximation, and always pretty good. It turns out that rsis closely related to another conventional measure of nonparametriccorrelation,theso-called sumsquareddifferenceofranks ,definedas D=N/summationdisplay i=1(Ri−Si)2(14.6.3 ) (This Dis sometimes denoted D**, where the asterisks are used to indicate that ties are treated by midranking.) 14.6 NonparametricorRankCorrelation 635Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).Whenthereare noties in thedata,thenthe exactrelationbetween Dandrsis rs=1−6D N3−N(14.6.4 ) When there are ties, then the exact relation is slightly more complicated: Let fkbe thenumberoftiesinthe kthgroupoftiesamongthe Ri’s,andlet gmbethenumber of ties in the mth group of ties among the Si’s. Then it turns out that rs=1−6 N3−N/bracketleftbig D+1 12/summationtext k(f3 k−fk)+1 12/summationtext m(g3 m−gm)/bracketrightbig /bracketleftBigg 1−/summationtext k(f3 k−fk) N3−N/bracketrightBigg1/2/bracketleftBigg 1−/summationtext m(g3 m−gm) N3−N/bracketrightBigg1/2(14.6.5 ) holds exactly. Notice that if all the fk’s and all the gm’s are equal to one, meaning that there are no ties, then equation (14.6.5) reduces to equation (14.6.4). In (14.6.2)we gave a t-statistic that tests the significance of a nonzero rs.I ti s also possible to test the significance of Ddirectly. The expectation value of Din the null hypothesis of uncorrelated data sets is D=1 6(N3−N)−1 12/summationdisplay k(f3 k−fk)−1 12/summationdisplay m(g3 m−gm)( 14.6.6 ) its variance is Var(D)=(N−1)N2(N+1 )2 36 ×/bracketleftbigg 1−/summationtext k(f3 k−fk) N3−N/bracketrightbigg/bracketleftbigg 1−/summationtext m(g3 m−gm) N3−N/bracketrightbigg (14.6.7 ) and it is approximately normally distributed, so that the significance level is a complementaryerrorfunction(cf.equation14.5.2). Ofcourse,(14.6.2)and(14.6.7)are not independent tests, but simply variants of the same test. In the program that follows, we return both the significance level obtained by using (14.6.2) and the significancelevelobtainedbyusing(14.6.7);theirdiscrepancywillgiveyouanideaof how goodthe approximationsare. You will also notice that we breakoff the task of assigning ranks (includingtied midranks) into a separate routine, crank. SUBROUTINE spear(data1,data2,n,wksp1,wksp2,d,zd,probd,rs,probrs) INTEGER nREAL d,probd,probrs,rs,zd,data1(n),data2(n),wksp1(n),wksp2(n) C USES betai,crank,erfcc,sort2 Given two data arrays, data1(1:n) anddata2(1:n) ,e a c ho fl e n gt h n,a n dgi v e nt w o workspaces of the same length, this routine returns their sum-squared difference of ranks asD, the number of standard deviations by which Ddeviates from its null-hypothesis expected value as zd, the two-sided significance level of this deviation as probd,S p e a r m a n ’ sr a n k correlation rsasrs, and the two-sided significance level of its deviation from zero as probrs. The workspaces can be identical to the data arrays, but in that case the data arrays are destroyed. The external routines crank(below) and sort2(§8.2) are used. A 636 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).small value of either probdorprobrsindicates a significant correlation ( rspositive) or anticorrelation ( rsnegative). INTEGER j REAL aved,df,en,en3n,fac,sf,sg,t,vard,betai,erfccdo 11j=1,n wksp1(j)=data1(j) wksp2(j)=data2(j) enddo 11 call sort2(n,wksp1,wksp2) Sort each of the data arrays, and convert the entries to ranks. The values sfand sgreturn the sums/summationtext(f3 k−fk) and/summationtext(g3 m−gm), respectively.call crank(n,wksp1,sf) call sort2(n,wksp2,wksp1)call crank(n,wksp2,sg)d=0. do 12j=1,n Sum the squared difference of ranks. d=d+(wksp1(j)-wksp2(j))**2 enddo 12 en=nen3n=en**3-enaved=en3n/6.-(sf+sg)/12. Expectation value of D, fac=(1.-sf/en3n)*(1.-sg/en3n) vard=((en-1.)*en**2*(en+1.)**2/36.)*fac and variance of Dgive zd=(d-aved)/sqrt(vard) number of standard deviations, probd=erfcc(abs(zd)/1.4142136) and significance. rs=(1.-(6./en3n)*(d+(sf+sg)/12.))/sqrt(fac) Rank correlation coefficient, fac=(1.+rs)*(1.-rs)if(fac.gt.0.)then t=rs*sqrt((en-2.)/fac) and its tvalue, df=en-2. probrs=betai(0.5*df,0.5,df/(df+t**2)) give its significance. else probrs=0. endifreturnEND SUBROUTINE crank(n,w,s) INTEGER n REAL s,w(n) Given a sorted array w(1:n), replaces the elements by their rank, includingmidrankingof ties, and returns as sthe sum of f3−f,w h e r e fis the number of elements in each tie. INTEGER j,ji,jt REAL rank,ts=0.j=1 The next rank to be assigned. 1 if(j.lt.n)then “do while ” structure. if(w(j+1).ne.w(j))then Not a tie. w(j)=j j=j+1 else At i e : do 11jt=j+1,n How far does it go? if(w(jt).ne.w(j))goto 2 enddo 11 jt=n+1 If here, it goes all the way to the last element. 2 rank=0.5*(j+jt-1) This is the mean rank of the tie, do12ji=j,jt-1 so enter it into all the tied entries, w(ji)=rank enddo 12 t=jt-j s=s+t**3-t and update s. j=jt endif goto 1 14.6 NonparametricorRankCorrelation 637Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).endif if(j.eq.n)w(n)=n I ft h el a s te l e m e n tw a sn o tt i e d ,t h i si si t sr a n k . return END Kendall’sTau Kendall’s τis even more nonparametric than Spearman’s rsorD. Instead of using the numerical difference of ranks, it uses only the relative ordering of ranks:higher in rank, lower in rank, or the same in rank. But in that case we don’t even have to rank the data! Ranks will be higher, lower, or the same if and only if the values are larger, smaller, or equal, respectively. On balance, we prefer r sas being the more straightforwardnonparametrictest, but both statistics are in general use. In fact, τandrsare very strongly correlated and, in most applications, are effectively the same test. To define τ, we start with the Ndata points (xi,yi). Now consider all 1 2N(N−1)pairsof data points, where a data point cannot be paired with itself, and where the points in either order count as one pair. We call a pair concordant if the relative ordering of the ranks of the two x’s (or for that matter the two x’s themselves) is the same as the relative ordering of the ranks of the two y’s (or for thatmatterthetwo y’sthemselves). Wecallapair discordant iftherelativeordering of the ranks of the two x’s is opposite from the relative ordering of the ranks of the twoy’s. If there is a tie in either the ranks of the two x’s or the ranks of the two y’s, then we don’t call the pair either concordant or discordant. If the tie is in the x’s, we will call the pair an “extra ypair.” If the tie is in the y’s, we will call the pair an “extra xpair.” If the tie is in both the x’s and the y’s, we don’t call the pair anything at all. Are you still with us? Kendall’s τis nowthe followingsimplecombinationofthese variouscounts: τ=concordant −discordant√concordant +discordant +extra- y√ concordant +discordant +extra- x (14.6.8 ) You can easily convince yourself that this must lie between 1and−1, and that it takes on the extreme values only for complete rank agreement or complete rank reversal, respectively. More important, Kendall has worked out, from the combinatorics, the approx- imate distribution of τin the null hypothesis of no association between xandy. In this case τis approximately normally distributed, with zero expectation value and a variance of Var(τ)=4N+1 0 9N(N−1)(14.6.9 ) The following program proceeds according to the above description, and therefore loops over all pairs of data points. Beware: This is an O(N2)algorithm, unlikethealgorithmfor rs,whosedominantsortoperationsareoforder NlogN.I f you are routinely computing Kendall’s τfor data sets of more than a few thousand points, you may be in for some serious computing. If, however, you are willing to bin your data into a moderate number of bins, then read on. 638 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).SUBROUTINE kendl1(data1,data2,n,tau,z,prob) INTEGER n REAL prob,tau,z,data1(n),data2(n) C USES erfcc Given data arrays data1(1:n) anddata2(1:n) , this program returns Kendall’s τastau, itsnumber of standard deviationsfromzeroas z, anditstwo-sidedsignificancelevelas prob. Small values of probindicate a significant correlation ( taupositive) or anticorrelation ( tau negative). INTEGER is,j,k,n1,n2 REAL a1,a2,aa,var,erfcc n1=0 This will be the argument of one square root in (14.6.8), n2=0 and this the other. is=0 This will be the numerator in (14.6.8). do12j=1,n-1 Loop over first member of pair, do11k=j+1,n and second member. a1=data1(j)-data1(k) a2=data2(j)-data2(k) aa=a1*a2if(aa.ne.0.)then Neither array has a tie. n1=n1+1 n2=n2+1 if(aa.gt.0.)then is=is+1 else is=is-1 endif else One or both arrays have ties. if(a1.ne.0.)n1=n1+1 An “extra x”e v e n t . if(a2.ne.0.)n2=n2+1 An “extra y”e v e n t . endif enddo 11 enddo 12 tau=float(is)/sqrt(float(n1)*float(n2)) Equation (14.6.8). var=(4.*n+10.)/(9.*n*(n-1.)) Equation (14.6.9). z=tau/sqrt(var) prob=erfcc(abs(z)/1.4142136) Significance. returnEND Sometimes it happens that there are only a few possible values each for xand y. Inthatcase,thedatacanberecordedasacontingencytable(see §14.4)thatgives the number of data points for each contingency of xandy. Spearman’s rank-order correlation coefficient is not a very natural statistic underthesecircumstances,sinceitassignstoeach xandybinanot-very-meaningful midrankvalueandthentotalsupvastnumbersofidenticalrankdifferences. Kendall’stau,ontheotherhand,withits simplecounting,remainsquitenatural. Furthermore, itsO(N 2)algorithmis nolongera problem,sincewecanarrangeforittoloopover pairsofcontingencytableentries(eachcontainingmanydatapoints)insteadofover pairs of data points. This is implemented in the program that follows. Note that Kendall’s tau can be applied only to contingency tables where both variables are ordinal, i.e., well-ordered, and that it looks specifically for monotonic correlations,notforarbitraryassociations. Thesetwopropertiesmakeitlessgeneral than the methods of §14.4, which applied to nominal, i.e., unordered, variables and arbitrary associations. Comparing kendl1above with kendl2below, you will see that we have “floated” a number of variables. This is because the number of events in a contingency table might be sufficiently large as to cause overflows in some of the 14.6 NonparametricorRankCorrelation 639Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).integer arithmetic, while the number of individual data points in a list could not possibly be that large [for an O(N2)routine!]. SUBROUTINE kendl2(tab,i,j,ip,jp,tau,z,prob) INTEGER i,ip,j,jp REAL prob,tau,z,tab(ip,jp) C USES erfcc Given a two-dimensional table tabof physical dimension (ip,jp) and logical dimension (i,j), such that tab(k,l) contains the number of events fallingin bin kof one variable and bin lof another, this program returns Kendall’s τastau, its number of standard deviations from zero as z, and its two-sided significance level as prob. Small values of prob indicate a significant correlation ( taupositive) or anticorrelation ( taunegative) between the two variables. Although tabis a real array, it will normally contain integral values. INTEGER k,ki,kj,l,li,lj,m1,m2,mm,nnREAL en1,en2,pairs,points,s,var,erfcc en1=0. Seekendl1above. en2=0.s=0. nn=i*j Total number of entries in contingency table. points=tab(i,j)do 12k=0,nn-2 Loop over entries in table, ki=k/j decodinga row index, kj=k-j*ki and a column index. points=points+tab(ki+1,kj+1) Increment the total count of events. do11l=k+1,nn-1 Loop over other member of the pair, li=l/j decodingits row lj=l-j*li and column. m1=li-kim2=lj-kj mm=m1*m2 pairs=tab(ki+1,kj+1)*tab(li+1,lj+1)if(mm.ne.0)then Not a tie. en1=en1+pairs en2=en2+pairs if(mm.gt.0)then Concordant, or s=s+pairs else discordant. s=s-pairs endif else if(m1.ne.0)en1=en1+pairs if(m2.ne.0)en2=en2+pairs endif enddo 11 enddo 12 tau=s/sqrt(en1*en2) var=(4.*points+10.)/(9.*points*(points-1.)) z=tau/sqrt(var) prob=erfcc(abs(z)/1.4142136)returnEND CITED REFERENCES AND FURTHER READING: Lehmann, E.L. 1975, Nonparametrics: Statistical Methods Based on Ranks (San Francisco: Holden-Day). Downie, N.M., and Heath, R.W. 1965, Basic Statistical Methods , 2nd ed. (New York: Harper & Row), pp. 206–209. Norusis,M.J.1982, SPSSIntroductoryGuide:BasicStatisticsandOperations ;and1985, SPSS- X Advanced Statistics Guide (New York: McGraw-Hill). 640 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X) Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine- readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).14.7 Do Two-DimensionalDistributions Differ? We here discuss a useful generalization of the K–S test ( §14.3) totwo-dimensional distributions. ThisgeneralizationisduetoFasanoandFranceschini [1],avariant onanearlier idea due to Peacock [2]. In a two-dimensional distribution, each data point is characterized by an (x, y )pair of values. An example near to our hearts is that each of the 19 neutrinos that were detectedfrom Supernova 1987A is characterized by a time t iand by an energy Ei(see[3]). We might wish to know whether these measured pairs (ti,E i),i=1 ...19are consistent with a theoretical model that predicts neutrino flux as a function of both time and energy — that is,a two-dimensional probability distribution in the (x, y )[here, (t, E )] plane. That would be a one-sample test. Or, given two sets of neutrino detections, from two comparable detectors, we might want to know whether they are compatible with each other, a two-sample test. In the spirit of the tried-and-true, one-dimensional K–S test, we want to range over the (x, y )plane in search of some kind of maximum cumulative difference between two two-dimensional distributions. Unfortunately, cumulative probability distribution is notwell-defined in more than one dimension! Peacock’s insight was that a good surrogate istheintegrated probability in each of four natural quadrants around a given point (x i,y i), namely the total probabilities (or fraction of data) in (x>x i,y > y i),(x<x i,y > y i), (x<x i,y < y i),(x>x i,y < y i). The two-dimensional K–S statistic Dis now taken to be the maximum difference (ranging both over data points and over quadrants) of thecorresponding integrated probabilities. When comparing two data sets, the value of Dmay depend on which data set is ranged over. In that case, define an effective Das the average of the two values obtained. If you are confused at this point about the exact definition of D, don’t fret; the accompanying computer routines amount to a precise algorithmic definition. Figure14.7.1gives afeelingforwhatisgoing on. The65trianglesand 35squares seem to have somewhat different distributions in the plane. The dotted lines are centered on thetriangle that maximizes the Dstatistic; the maximum occurs in the upper-left quadrant. That quadrant contains only 0.12 of all the triangles, but it contains 0.56 of all the squares. Thevalue of Dis thus 0.44. Is this statistically significant? Evenforfixedsamplesizes,itisunfortunately notrigorously truethatthedistributionof Dinthenullhypothesisisindependentoftheshapeofthetwo-dimensionaldistribution. Inthis respectthetwo-dimensionalK–Stestisnotasnaturalasitsone-dimensionalparent. However,extensive Monte Carlo integrations have shown that the distribution of the two-dimensionalDisvery nearly identical for even quite different distributions, as long as they have the same coefficient of correlation r, defined in the usual way by equation (14.5.1). In their paper, FasanoandFranceschinitabulateMonteCarloresultsfor(whatamountsto)thedistributionofDas a function of (of course) D, sample size N, and coefficient of correlation r. Analyzing their results, one finds that the significance levels for the two-dimensional K–S test can besummarized by the simple, though approximate, formulas, Probability (D>observed )= Q KS/parenleftbigg √ ND 1+√ 1−r2(0.25−0.75/√ N)/parenrightbigg (14.7.1 ) for the one-sample case, and the same for the two-sample case, but with N=N1N2 N1+N2. (14.7.2 ) The above formulas are accurate enough when N>∼20, and when the indicated probability (significance level) is less than (more significant than) 0.20or so. When the indicated probability is >0.20, its value may not be accurate, but the implication that the data and model (or two data sets) are not significantly different is certainly correct. Noticethat in the limit of r→ 1(perfect correlation), equations (14.7.1) and (14.7.2) reduce to equations (14.3.9) and (14.3.10): The two-dimensional data lie on a perfect straight line, andthe two-dimensional K–S test becomes a one-dimensional K–S test.