f14-3
PDF · 9 pages · 104.9 KB
Open PDF file
Photocopied or downloaded sample pages (pp. 614 onward) from Numerical Recipes in Fortran 77, Chapter 14, Statistical Description of Data, by Press et al. (Cambridge University Press, 1986-1992). It covers the null hypothesis, the chi-square statistic for binned data, degrees of freedom and constraints, and the Fortran routines chsone and chstwo. It then moves on to unequal sample sizes and the Kolmogorov-Smirnov test for continuous data.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
614 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).14.3 Are Two Distributions Different?
Given two sets of data, we can generalize the questions asked in the previous
sectionandaskthesinglequestion: Arethetwosetsdrawnfromthesamedistribution
function, or from different distribution functions? Equivalently,in proper statistical
language, “Can we disprove, to a certain required level of significance, the null
hypothesis that two data sets are drawn from the same population distribution
function?” Disprovingthenullhypothesisineffectprovesthatthedatasetsarefromdifferent distributions. Failing to disprove the null hypothesis, on the other hand,
only shows that the data sets can be consistent with a single distribution function.
One can never provethat two data sets come from a single distribution, since (e.g.)
no practical amount of data can distinguish between two distributions which differ
only by one part in 10
10.
Provingthattwodistributionsaredifferent,orshowingthattheyareconsistent,
is a task that comes up all the time in many areas of research: Are the visible stars
distributed uniformly in the sky? (That is, is the distribution of stars as a functionof declination — position in the sky — the same as the distribution of sky area as
a function of declination?) Are educational patterns the same in Brooklyn as in the
Bronx? (That is, are the distributions of people as a function of last-grade-attendedthe same?) Do two brands of fluorescent lights have the same distribution of
burn-outtimes? Istheincidenceofchickenpoxthesameforfirst-born,second-born,
third-born children, etc.?
Thesefourexamplesillustratethefourcombinationsarisingfromtwodifferent
dichotomies: (1) The data are either continuous or binned. (2) Either we wish tocompare one data set to a known distribution, or we wish to compare two equally
unknown data sets. The data sets on fluorescent lights and on stars are continuous,
since we can be given lists of individual burnout times or of stellar positions. Thedata sets on chicken pox and educational level are binned, since we are given
tables of numbers of events in discrete categories: first-born, second-born, etc.; or
6th Grade, 7th Grade, etc. Stars and chicken pox, on the other hand, share the
property that the null hypothesis is a known distribution (distribution of area in the
sky, or incidence of chicken pox in the general population). Fluorescent lights andeducationallevel involvethe comparisonoftwo equallyunknowndatasets (the two
brands, or Brooklyn and the Bronx).
One can always turn continuous data into binned data, by grouping the events
into specified ranges of the continuous variable(s): declinations between 0 and 10
degrees,10and20,20and30,etc. Binninginvolvesa loss ofinformation,however.
Also, there is often considerable arbitrariness as to how the bins should be chosen.
Alongwithmanyotherinvestigators,weprefertoavoidunnecessarybinningofdata.
Theacceptedtestfordifferencesbetweenbinneddistributionsisthe chi-square
test. For continuous data as a function of a single variable, the most generally
accepted test is the Kolmogorov-Smirnovtest . We consider each in turn.
Chi-Square Test
Suppose that Niis the number of events observed in the ith bin, and that niis
the number expected according to some known distribution. Note that the Ni’s are
14.3AreTwoDistributionsDifferent? 615Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).integers, while the ni’s may not be. Then the chi-square statistic is
χ2=/summationdisplay
i(Ni−ni)2
ni(14.3.1 )
wherethesumis overall bins. Alargevalueof χ2indicatesthatthenullhypothesis
(thatthe Ni’saredrawnfromthepopulationrepresentedbythe ni’s)isratherunlikely.
Any term jin (14.3.1)with 0=nj=Njshould be omitted from the sum. A
term with nj=0,Nj/negationslash=0gives an infinite χ2, as it should, since in this case the
Ni’s cannot possibly be drawn from the ni’s!
Thechi-squareprobabilityfunction Q(χ2|ν)isanincompletegammafunction,
and was already discussed in §6.2 (see equation 6.2.18). Strictly speaking Q(χ2|ν)
is the probability that the sum of the squares of νrandomnormalvariables of unit
variance (and zero mean) will be greater than χ2. The terms in the sum (14.3.1)
are not individually normal. However, if either the number of bins is large ( /greatermuch1),
or the number of events in each bin is large ( /greatermuch1), then the chi-square probability
functionisagoodapproximationtothedistributionof(14.3.1)inthecaseofthenull
hypothesis. Its use to estimate the significance of the chi-squaretest is standard.
The appropriate value of ν, the number of degrees of freedom, bears some
additional discussion. If the data are collected with the model ni’s fixed — that
is, not later renormalized to fit the total observed number of events ΣNi— then ν
equals the number of bins NB. (Note that this is notthe total number of events!)
Muchmorecommonly,the ni’sarenormalizedafterthefactsothattheirsumequals
the sum of the Ni’s. In this case the correct value for νisNB−1, and the model
is said to have one constraint ( knstrn=1 in the program below). If the model that
givesthe ni’shasadditionalfreeparametersthatwereadjustedafterthefacttoagree
with the data, then each of these additional “fitted” parameters decreases ν(and
increases knstrn) by one additional unit.
We have, then, the following program:
SUBROUTINE chsone(bins,ebins,nbins,knstrn,df,chsq,prob)
INTEGER knstrn,nbins
REAL chsq,df,prob,bins(nbins),ebins(nbins)
C USES gammq
Giventhe array bins(1:nbins) containing theobserved numbers ofevents, andanarray
ebins(1:nbins) containing the expected numbers of events, and given the number of
constraints knstrn(normallyone),thisroutinereturns(trivially)thenumberofdegreesof
freedom df,and(nontrivially)thechi-square chsqandthesignificance prob. Asmallvalue
ofprobindicates asignificantdifference between thedistributions binsandebins.N o t e
that binsandebinsare both real arrays, although binswill normally contain integer
values.
INTEGER jREAL gammq
df=nbins-knstrn
chsq=0.do
11j=1,nbins
if(ebins(j).le.0.)pause ’bad expected number in chsone’
chsq=chsq+(bins(j)-ebins(j))**2/ebins(j)
enddo 11
prob=gammq(0.5*df,0.5*chsq) Chi-square probability function. See §6.2.
return
END
616 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).Next we consider the case of comparing twobinned data sets. Let Ribe the
number of events in bin ifor the first data set, Sithe number of events in the same
binifor the second data set. Then the chi-square statistic is
χ2=/summationdisplay
i(Ri−Si)2
Ri+Si(14.3.2 )
Comparing (14.3.2)to (14.3.1),you should note that the denominatorof (14.3.2)is
notjust the average of RiandSi(which would be an estimator of niin 14.3.1).
Rather,it is twice the average,the sum. Thereasonis that each termin a chi-square
sum is supposed to approximate the square of a normally distributed quantity withunit variance. The variance of the difference of two normal quantities is the sum
of their individual variances, not the average.
If the data were collected in such a way that the sum of the R
i’s is necessarily
equal to the sum of Si’s, then the number of degrees of freedom is equal to one
less than the number of bins, NB−1(that is, knstrn =1), the usual case. If
this requirementwere absent,then the numberofdegreesof freedomwouldbe NB.
Example: A birdwatcher wants to know whether the distribution of sighted birds
as a function of species is the same this year as last. Each bin corresponds to onespecies. If the birdwatcher takes his data to be the first 1000 birds that he saw in
each year, then the numberof degreesof freedomis N
B−1. If he takes his data to
beallthebirdshesawonarandomsampleofdays,thesamedaysineachyear,thenthe numberof degreesof freedomis N
B(knstrn =0). Inthis latter case, notethat
he is also testing whether the birds were more numerous overall in one year or the
other: That is the extra degree of freedom. Of course, any additionalconstraints on
the data set lower the number of degrees of freedom (i.e., increase knstrntomore
positivevalues) in accordance with their number.
The program is
SUBROUTINE chstwo(bins1,bins2,nbins,knstrn,df,chsq,prob)
INTEGER knstrn,nbinsREAL chsq,df,prob,bins1(nbins),bins2(nbins)
C USES gammq
Giventhe arrays bins1(1:nbins) andbins2(1:nbins) , containing two sets of binned
data, and given the number of constraints knstrn(normally 1or0), this routine returns
the number of degrees of freedom df, the chi-square chsq, and the significance prob.
A small value of probindicates a significant difference between the distributions bins1
andbins2.N o t et h a t bins1andbins2areboth real arrays,although they willnormally
contain integer values.
INTEGER j
REAL gammqdf=nbins-knstrn
chsq=0.
do
11j=1,nbins
if(bins1(j).eq.0..and.bins2(j).eq.0.)then
df=df-1. No data means one less degree of freedom.
else
chsq=chsq+(bins1(j)-bins2(j))**2/(bins1(j)+bins2(j))
endif
enddo 11
prob=gammq(0.5*df,0.5*chsq) Chi-square probability function. See §6.2.
returnEND
14.3AreTwoDistributionsDifferent? 617Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).Equation(14.3.2)andtheroutine chstwobothapplytothecasewherethetotal
number of data points is the same in the two binned sets. For unequal numbers ofdata points, the formula analogous to (14.3.2) is
χ
2=/summationdisplay
i(/radicalbig
S/RR i−/radicalbig
R/SS i)2
Ri+Si(14.3.3 )
where
R≡/summationdisplay
iRi S≡/summationdisplay
iSi (14.3.4 )
are the respective numbers of data points. It is straightforward to make the
corresponding change in chstwo.
Kolmogorov-SmirnovTest
The Kolmogorov-Smirnov(or K–S) test is applicableto unbinneddistributions
that are functions of a single independent variable, that is, to data sets where each
data point can be associated with a single number (lifetime of each lightbulb when
it burns out, or declination of each star). In such cases, the list of data points can
be easily converted to an unbiased estimator SN(x)of thecumulative distribution
functionoftheprobabilitydistributionfromwhichitwasdrawn: Ifthe Neventsare
located at values xi,i=1,...,N, then SN(x)is the function giving the fraction
of data points to the left of a given value x. This function is obviously constant
between consecutive(i.e., sorted into ascending order) xi’s, and jumps by the same
constant 1/Nat each xi. (See Figure 14.3.1.)
Different distribution functions, or sets of data, give different cumulative
distribution function estimates by the above procedure. However, all cumulative
distribution functions agree at the smallest allowable value of x(where they are
zero), and at the largest allowable value of x(where they are unity). (The smallest
andlargestvaluesmightofcoursebe ±∞.) Soitis thebehaviorbetweenthelargest
and smallest values that distinguishes distributions.
One can think of any number of statistics to measure the overall difference
betweentwocumulativedistributionfunctions: theabsolutevalueoftheareabetween
them, for example. Or their integrated mean square difference. The Kolmogorov-
Smirnov Dis a particularly simple measure: It is defined as the maximum value
of the absolute difference between two cumulative distribution functions. Thus,
for comparing one data set’s SN(x)to a known cumulative distribution function
P(x), the K–S statistic is
D=m a x
−∞<x<∞|SN(x)−P(x)| (14.3.5 )
while for comparing two different cumulative distribution functions SN1(x)and
SN2(x), the K–S statistic is
D=m a x
−∞<x<∞|SN1(x)−SN2(x)| (14.3.6 )
618 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).x
xD
P(x)SN(x)cumulative probability distribution
Figure 14.3.1. Kolmogorov-Smirnov statistic D. A measured distribution of values in x(shown
asNdots on the lower abscissa) is to be compared with a theoretical distribution whose cumulative
probability distribution is plotted as P(x). A step-function cumulative probability distribution SN(x)is
constructed, one that rises an equal amount at each measured point. Dis the greatest distance between
the two cumulative distributions.
WhatmakestheK –Sstatisticusefulisthat itsdistributioninthecaseofthenull
hypothesis(datasets drawnfromthesamedistribution)canbecalculated,atleast to
usefulapproximation,thusgivingthe signi ficanceof anyobservednonzerovalueof
D. A central feature of the K –S test is that it is invariant under reparametrization
ofx; in other words, you can locally slide or stretch the xaxis in Figure 14.3.1,
and the maximum distance Dremains unchanged. For example, you will get the
same signi ficance using xas using logx.
The function that enters into the calculation of the signi ficance can be written
as the following sum:
QKS(λ)=2∞/summationdisplay
j=1(−1)j−1e−2j2λ2(14.3.7 )
which is a monotonic function with the limiting values
QKS(0) = 1 QKS(∞)=0 ( 14.3.8 )
In terms of this function, the signi ficance level of an observed value of D(as
a disproof of the null hypothesis that the distributions are the same) is given
approximately [1]by the formula
Probability (D>observed )=QKS/parenleftBig/bracketleftBig/radicalbig
Ne+0.12 + 0 .11//radicalbig
Ne/bracketrightBig
D/parenrightBig
(14.3.9 )
14.3Are TwoDistributionsDifferent? 619Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).where Neis the effective number of data points, Ne=Nfor the case (14.3.5)
of one distribution, and
Ne=N1N2
N1+N2(14.3.10 )
for the case (14.3.6) of two distributions, where N1is the number of data points in
thefirst distribution, N2the number in the second.
The natureof the approximationinvolvedin (14.3.9)is that it becomesasymp-
totically accurate as the Nebecomes large,but is alreadyquite good for Ne≥4,a s
small a number as one might ever actually use. (See [1].)
So, we havethe followingroutinesfor thecases ofoneandtwo distributions:
SUBROUTINE ksone(data,n,func,d,prob)
INTEGER nREAL d,data(n),func,prob
EXTERNAL func
C USES probks,sort
G i v e na na r r a y data(1:n), and given a user-supplied function of a single variable func
which is a cumulative distribution function ranging from 0 (for smallest values of its argu-
ment)to1(forlargestvaluesofitsargument), thisroutinereturnstheK–Sstatistic d,and
the significance level prob. Small values of probshowthat the cumulative distribution
function of dataissignificantly different from func. Thearray dataismodifiedbybeing
sorted into ascending order.
INTEGER jREAL dt,en,ff,fn,fo,probks
call sort(n,data) If the data are already sorted into ascending or-
der, then this call can be omitted. en=n
d=0.fo=0. Data’s c.d.f. before the next step.
do
11j=1,n Loop over the sorted data points.
fn=j/en Data’s c.d.f. after this step.
ff=func(data(j)) Compare to the user-supplied function.
dt=max(abs(fo-ff),abs(fn-ff)) Maximum distance.
if(dt.gt.d)d=dt
fo=fn
enddo 11
en=sqrt(en)prob=probks((en+0.12+0.11/en)*d) Compute significance.
returnEND
SUBROUTINE kstwo(data1,n1,data2,n2,d,prob)
INTEGER n1,n2
REAL d,prob,data1(n1),data2(n2)
C USES probks,sort
Given an array data1(1:n1) , and an array data2(1:n2) , this routine returns the K–
S statistic d, and the significance level probfor the null hypothesis that the data sets
are drawn from the same distribution. Small values of probshowthat the cumulative
distribution function of data1is significantly different from that of data2. The arrays
data1anddata2are modified by being sorted into ascending order.
INTEGER j1,j2
REAL d1,d2,dt,en1,en2,en,fn1,fn2,probkscall sort(n1,data1)call sort(n2,data2)
en1=n1
en2=n2j1=1 Next value of data1to be processed.
j2=1 Ditto, data2.
620 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).fn1=0.
fn2=0.
d=0.
1 if(j1.le.n1.and.j2.le.n2)then If we are not done...
d1=data1(j1)
d2=data2(j2)
if(d1.le.d2)then Next step is in data1.
fn1=j1/en1j1=j1+1
endif
if(d2.le.d1)then Next step is in data2.
fn2=j2/en2j2=j2+1
endif
dt=abs(fn2-fn1)if(dt.gt.d)d=dt
goto 1
endifen=sqrt(en1*en2/(en1+en2))prob=probks((en+0.12+0.11/en)*d) Compute significance.
return
END
Bothoftheaboveroutinesusethefollowingroutineforcalculatingthefunction
QKS:
FUNCTION probks(alam)
REAL probks,alam,EPS1,EPS2PARAMETER (EPS1=0.001, EPS2=1.e-8)
Kolmogorov-Smirnov probability function.
INTEGER j
REAL a2,fac,term,termbfa2=-2.*alam**2
fac=2.
probks=0.termbf=0. Previous term in sum.
do
11j=1,100
term=fac*exp(a2*j**2)
probks=probks+termif(abs(term).le.EPS1*termbf.or.abs(term).le.EPS2*probks)returnfac=-fac Alternating signs in sum.
termbf=abs(term)
enddo
11
probks=1. Get here only by failing to converge.
return
END
Variants onthe K–S Test
The sensitivity of the K –S test to deviations from a cumulative distribution function
P(x)is not independent of x. In fact, the K –S test tends to be most sensitive around the
median value, where P(x)=0 .5, and less sensitive at the extreme ends of the distribution,
where P(x)is near 0or1. The reason is that the difference |SN(x)−P(x)|does not, in the
nullhypothesis,haveaprobability distributionthatisindependent of x. Rather,itsvariance is
proportional to P(x)[1−P(x)],which islargestat P=0.5. SincetheK –Sstatistic(14.3.5)
isthemaximumdifferenceoverall xoftwocumulativedistributionfunctions,adeviationthat
might be statistically signi ficant atits ownvalue of xgets compared to the expected chance
deviation at P=0.5, and is thus discounted. A result is that, while the K –S test is good at
14.3AreTwoDistributionsDifferent? 621Sample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).findingshiftsin a probability distribution, especially changes in the median value, it is not
always so good at findingspreads, which more affect the tails of the probability distribution,
and which may leave the median unchanged.
One way of increasing the power of the K –S statistic out on the tails is to replace
D(equation 14.3.5) by a so-called stabilized orweighted statistic[2-4], for example the
Anderson-Darling statistic ,
D*= m a x
−∞<x<∞|SN(x)−P(x)|/radicalbig
P(x)[1−P(x)](14.3.11 )
Unfortunately,thereisnosimpleformulaanalogous toequations (14.3.7)and(14.3.9)forthis
statistic,althoughNo ´e[5]givesacomputationalmethodusingarecursionrelationandprovides
a graph of numerical results. There are many other possible similar statistics, for example
D** =/integraldisplay1
P=0[SN(x)−P(x)]2
P(x)[1−P(x)]dP(x)( 14.3.12 )
which is also discussed by Anderson and Darling (see [3]).
Another approach, which we prefer as simpler and more direct, is due to Kuiper [6,7].
We already mentioned that the standard K –S test is invariant under reparametrizations of the
variable x. An evenmore generalsymmetry, which guarantees equalsensitivitiesatallvalues
ofx, is to wrap the xaxis around into a circle (identifying the points at ±∞),and to look for
astatisticthatisnowinvariant underallshiftsandparametrizationsonthecircle. Thisallows,for example, a probability distribution to be “cut”at some central value of x, and the left and
right halves to be interchanged, without altering the statistic or its signi ficance.
Kuiper’s statistic ,d efined as
V=D
++D−=m a x
−∞<x<∞[SN(x)−P(x)] + max
−∞<x<∞[P(x)−SN(x)] (14.3.13 )
is the sum of the maximum distance of SN(x)above and below P(x). You should be able
to convince yourself that this statistic has the desired invariance on the circle: Sketch theindefinite integral of two probability distributions de fined on the circle as a function of angle
around the circle, as the angle goes through several times 360
◦. If you change the starting
point of the integration, D+andD−change individually, but their sum is constant.
Furthermore, there is a simple formula for the asymptotic distribution of the statistic V,
directly analogous to equations (14.3.7) –(14.3.10). Let
QKP(λ)=2∞/summationdisplay
j=1(4j2λ2−1)e−2j2λ2(14.3.14 )
which is monotonic and satis fies
QKP(0) = 1 QKP(∞)=0 ( 14.3.15 )
In terms of this function the signi ficance level is [1]
Probability (V>observed )=QKP/parenleftbigg/bracketleftBig√
Ne+0.155 + 0 .24/√
Ne/bracketrightBig
V/parenrightbigg
(14.3.16 )
Here NeisNin the one-sample case, or is given by equation (14.3.10) in the case of
two samples.
Of course, Kuiper ’s test is ideal for any problem originally de fined on a circle, for
example, to test whether the distribution in longitude of something agrees with some theory,or whether two somethings have different distributions in longitude. (See also
[8].)
We will leave to you the coding of routines analogous to ksone,kstwo, and probks,
above. (For λ< 0.4, don’ttrytodo thesum 14.3.14. Itsvalue is1,to7 figures,buttheseries
can require many terms to converge, and loses accuracy to roundoff.)
Twofinal cautionary notes: First, we should mention that all varieties of K –S test lack
the ability to discriminate some kinds of distributions. A simple example is a probabilitydistribution with a narrow “notch”within which the probability falls to zero. Such a
distribution is of course ruled out by the existence of even one data point within the notch,
622 Chapter14. StatisticalDescriptionofDataSample page from NUMERICAL RECIPES IN FORTRAN 77: THE ART OF SCIENTIFIC COMPUTING (ISBN 0-521-43064-X)
Copyright (C) 1986-1992 by Cambridge University Press.Programs Copyright (C) 1986-1992 by Numerical Recipes Software. Permission is granted for internet users to make one paper copy for their own personal use. Further reproduction, or any copyin g of machine-
readable files (including this one) to any servercomputer, is strictly prohibited. To order Numerical Recipes booksor CDROMs, v isit website
http://www.nr.com or call 1-800-872-7423 (North America only),or send email to [email protected] (outside North Amer ica).but, because of its cumulative nature, a K –S test would require many data points in the notch
before signaling a discrepancy.
Second,weshouldnotethat,ifyouestimateanyparametersfromadataset(e.g.,amean
andvariance),thenthedistributionoftheK –Sstatistic Dforacumulativedistributionfunction
P(x)thatuses the estimated parameters is no longer given by equation (14.3.9). In general,
you will have to determine the new distribution yourself, e.g., by Monte Carlo methods.
CITED REFERENCES AND FURTHER READING:
von Mises, R. 1964, Mathematical Theory of Probability and Statistics (New York: Academic
Press), Chapters IX(C) and IX(E).
Stephens, M.A. 1970, Journalof the RoyalStatistical Society , ser. B, vol. 32, pp. 115–122.[1]
Anderson,T.W.,andDarling,D.A.1952, AnnalsofMathematicalStatistics , vol.23,pp.193–212.
[2]
Darling, D.A. 1957, Annals of Mathematical Statistics , vol. 28, pp. 823–838. [3]
Michael, J.R. 1983, Biometrika , vol. 70, no. 1, pp. 11–17. [4]
No´e, M. 1972, Annals of Mathematical Statistics , vol. 43, pp. 58–64. [5]
Kuiper,N.H. 1962, Proceedingsof theKoninklijkeNederlandseAkademie vanWetenschappen ,
ser. A., vol. 63, pp. 38–47. [6]
Stephens, M.A. 1965, Biometrika , vol. 52, pp. 309–321. [7]
Fisher, N.I., Lewis, T., and Embleton, B.J.J. 1987, Statistical Analysis of Spherical Data (New
York: Cambridge University Press). [8]
14.4 Contingency Table Analysis of Two
Distributions
Inthissection,andthenexttwosections,wedealwith measuresofassociation
fortwodistributions. Thesituationisthis: Eachdatapointhastwoormoredifferent
quantitiesassociatedwithit,andwewanttoknowwhetherknowledgeofonequantity
gives us any demonstrableadvantage in predicting the value of another quantity. In
manycases,onevariablewillbean “independent ”or“control”variable,andanother
will be a “dependent ”or“measured ”variable. Then, we want to know if the latter
variableisin fact dependent on or associated with the former variable. If it is, we
wanttohavesomequantitativemeasureofthestrengthoftheassociation. Oneoftenhears this loosely stated as the question of whether two variables are correlated or
uncorrelated , but we will reserve those terms for a particular kind of association
(linear, or at least monotonic), as discussed in §14.5 and §14.6.
Notice that, as in previous sections, the different concepts of signi ficance and
strength appear: The association between two distributions may be very signi ficant
even if that association is weak —if the quantity of data is large enough.
Itisusefultodistinguishamongsomedifferentkindsofvariables,withdifferent
categories forming a loose hierarchy.
•Avariableiscalled nominalifitsvaluesarethemembersofsomeunordered
set. For example, “state of residence ”is a nominal variable that (in the
U.S.) takes on one of 50 values; in astrophysics, “type of galaxy ”is a
nominalvariablewiththethreevalues “spiral,”“elliptical,”and“irregular.”