Lee Elisa T., Wang John Wenyu - Statistical Methods for Survival Data Analysis (3rd Edition, Wiley)
PDF · 536 pages · 6.9 MB
Open PDF file
Published textbook by Elisa T. Lee and John Wenyu Wang, not Phil's own work, aimed at biomedical researchers and statisticians. It covers survival functions, censored data, product-limit and life-table estimates, and nonparametric and parametric comparisons. Further topics are exponential, Weibull, lognormal, gamma and log-logistic models, goodness of fit, Cox proportional hazards and logistic regression, with statistical tables.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
FinancialEbooks.NET
||ry
Visit FinancialEbooks.NET for more financial information
@)WILEY
Statistical Methods for
Survival Data Analysis
Third Edition
Elisa T.Lee
ftp:/, JohnWenyuWang
WILEY SERIES INPROBABILITY AND STATISTICS
StatisticalMethodsfor
SurvivalDataAnalysis
StatisticalMethodsfor
SurvivalDataAnalysis
ThirdEdition
ELISAT.LEE
JOHNWENYUWANG
DepartmentofBiostatisticsandEpidemiologyand
CenterforAmericanIndianHealthResearchCollegeofPublicHealthUniversityofOklahomaHealthSciencesCenterOklahomaCity,Oklahoma
A JOHN WILEY & SONS, INC., PUBLICATION
Copyright /p72003byJohnWiley&Sons,Inc.Allrightsreserved.
PublishedbyJohnWiley&Sons,Inc.,Hoboken,NewJersey.
PublishedsimultaneouslyinCanada.
Nopartofthispublicationmaybereproduced,storedinaretrievalsystem,ortransmittedinany
formorbyanymeans,electronic,mechanical,photocopying,recording,scanning,orotherwise,exceptaspermittedunderSection107or108ofthe1976UnitedStatesCopyrightAct,
withouteitherthepriorwrittenpermissionofthePublisher,orauthorizationthroughpaymentof
theappropriateper-copyfeetotheCopyrightClearanceCenter,Inc.,222RosewoodDrive,Danvers,MA01923,978-750-8400,fax978-750-4470,oronthewebatwww.copyright.com.
RequeststothePublisherforpermissionshouldbeaddressedtothePermissionsDepartment,
JohnWiley&Sons,Inc.,111RiverStreet,Hoboken,NJ07030, (201)748-6011,fax (201)748-6008,
e-mail:permreq /p28wiley.com.
LimitofLiability/DisclaimerofWarranty:Whilethepublisherandauthorhaveusedtheirbest
effortsinpreparingthisbook,theymakenorepresentationsorwarrantieswithrespecttothe
accuracyorcompletenessofthecontentsofthisbookandspecificallydisclaimanyimplied
warrantiesofmerchantabilityorfitnessforaparticularpurpose.Nowarrantymaybecreatedorextendedbysalesrepresentativesorwrittensalesmaterials.Theadviceandstrategiescontained
hereinmaynotbesuitableforyoursituationYoushouldconsultwithaprofessionalwhere
appropriate.Neitherthepublishernorauthorshallbeliableforanylossofprofitoranyothercommercialdamages,includingbutnotlimitedtospecial,incidental,consequential,orother
damages.
ForgeneralinformationonourotherproductsandservicespleasecontactourCustomerCare
DepartmentwithintheU.S.at877-762-2974,outsidetheU.S.at317-572-3993orfax317-572-4002.
Wileyalsopublishesitsbooksinavarietyofelectronicformats.Somecontentthatappearsin
print,however,maynotbeavailableinelectronicformat.
Library of Congress Cataloging-in-Publication Data:
Lee,ElisaT.
Statisticalmethodsforsurvivaldataanalysis.--3rded./ElisaT.LeeandJohnWenyuWang.
p.cm.-- (Wileyseriesinprobabilityandstatistics )
Includesbibliographicalreferencesandindex.ISBN0-471-36997-7 (cloth:alk.paper )
1.Medicine--Research--Statisticalmethods.2.Failuretimedataanalysis.3.
Prognosis--Statisticalmethods.I.Wang,JohnWenyu.II.Title.III.Series.
R853.S7L432003
610/p30.72--dc21 2002027025
PrintedintheUnitedStatesofAmerica.
10987654321
Tothememoryofourparents
Mr.Chi-LanTanandMrs.Hwei-ChiLeeTan
(E.T.L. )
Mr.BeijunZhangandMrs.XiangyiWang
(J.W.W. )
Contents
Preface xi
1 Introduction
1.1 Preliminaries,1
1.2 CensoredData,1
1.3 ScopeoftheBook,5
BibliographicalRemarks,7
2 Functions of Survival Time 8
2.1 Definitions,8
2.2 RelationshipsoftheSurvivalFunctions,15
BibliographicalRemarks,17Exercises,17
3 Examples of Survival Data Analysis 19
3.1 Example3.1:ComparisonofTwoTreatmentsandThree
Diets,19
3.2 Example3.2:ComparisonofTwoSurvivalPatterns
UsingLifeTables,26
3.3 Example3.3:FittingSurvivalDistributionstoRemission
Data,29
3.4 Example3.4:RelativeMortalityandIdentificationof
PrognosticFactors,32
3.5 Example3.5:IdentificationofRiskFactors,40
BibliographicalRemarks,47
Exercises,47
vii
4 Nonparametric Methods of Estimating Survival Functions 64
4.1 Product-LimitEstimatesofSurvivorshipFunction,65
4.2 Life-TableAnalysis,77
4.3 Relative,Five-Year,andCorrectedSurvivalRates,94
4.4 StandardizedRatesandRatios,97
BibliographicalRemarks,102
Exercises,102
5 Nonparametric Methods for Comparing Survival Distributions 106
5.1 ComparisonofTwoSurvivalDistributions,106
5.2 Mantel —HaenszelTest,121
5.3 Comparisonof K(K/p572)Samples,125
BibliographicalRemarks,131
Exercises,131
6 Some Well-Known Parametric Survival Distributions
and Their Applications 134
6.1 ExponentialDistribution,134
6.2 WeibullDistribution,1386.3 LognormalDistribution,143
6.4 GammaandGeneralizedGammaDistributions,148
6.5 Log-LogisticDistribution,1546.6 OtherSurvivalDistributions,155
BibliographicalRemarks,160
Exercises,160
7 Estimation Procedures for Parametric Survival Distributions
without Covariates 162
7.1 GeneralMaximumLikelihoodEstimationProcedure,162
7.2 ExponentialDistribution,166
7.3 WeibullDistribution,1787.4 LognormalDistribution,180
7.5 StandardandGeneralizedGammaDistributions,188
7.6 Log-LogisticDistribution,1957.7 OtherParametricSurvivalDistributions,196
BibliographicalRemarks,196
Exercises,197viii
8 Graphical Methods for Survival Distribution Fitting 198
8.1 Introduction,198
8.2 ProbabilityPlotting,200
8.3 HazardPlotting,209
8.4 Cox —SnellResidualMethod,215
BibliographicalRemarks,219
Exercises,219
9 Tests of Goodness of Fit and Distribution Selection 221
9.1 Goodness-of-FitTestStatisticsBasedonAsymptotic
LikelihoodInferences,222
9.2 TestsforAppropriatenessofaFamilyofDistributions,225
9.3 SelectionofaDistributionUsingBIC
orAICProcedures,230
9.4 TestsforaSpecificDistributionwith
KnownParameters,233
9.5 HollanderandProschan’sTestforAppropriateness
ofaGivenDistributionwithKnownParameters,236
BibliographicalRemarks,238
Exercises,240
10 Parametric Methods for Comparing Two Survival Distributions 243
10.1 LikelihoodRatioTestforComparingTwoSurvival
Distributions,243
10.2 ComparisonofTwoExponentialDistributions,246
10.3 ComparisonofTwoWeibullDistributions,25110.4 ComparisonofTwoGammaDistributions,252
BibliographicalRemarks,254
Exercises,254
11 Parametric Methods for Regression Model Fitting and
Identification of Prognostic Factors 256
11.1 PreliminaryExaminationofData,257
11.2 GeneralStructureofParametricRegressionModels
andTheirAsymptoticLikelihoodInference,259
11.3 ExponentialRegressionModel,26311.4 WeibullRegressionModel,269
11.5 LognormalRegressionModel,274
11.6 ExtendedGeneralizedGammaRegressionModel,277 ix
11.7 Log-LogisticRegressionModel,280
11.8 OtherParametricRegressionModels,283
11.9 ModelSelectionMethods,286
BibliographicalRemarks,295Exercises,295
12 Identification of Prognostic Factors Related to Survival Time:
Cox Proportional Hazards Model 298
12.1 PartialLikelihoodFunctionforSurvivalTimes,298
12.2 IdentificationofSignificantCovariates,31412.3 EstimationoftheSurvivorshipFunctionwithCovariates,319
12.4 AdequacyAssessmentoftheProportionalHazardsModel,326
BibliographicalRemarks,336Exercises,337
13 Identification of Prognostic Factors Related to Survival Time:
Nonproportional Hazards Models 339
13.1 ModelswithTime-DependentCovariates,339
13.2 StratifiedProportionalHazardsModels,34813.3 CompetingRisksModel,352
13.4 RecurrentEventsModels,356
13.5 ModelsforRelatedObservations,374
BibliographicalRemarks,376
Exercises,376
14 Identification of Risk Factors Related to Dichotomous
and Polychotomous Outcomes 377
14.1 UnivariateAnalysis,378
14.2 LogisticandConditionalLogisticRegressionModels
forDichotomousResponses,385
14.3 ModelsforPolychotomousOutcomes,413
BibliographicalRemarks,425Exercises,425
Appendix A Newton--Raphson Method 428
Appendix B Statistical Tables 433References 488Index 511x
Preface
Statisticalmethodsforsurvivaldataanalysishavecontinuedtoflourishinthe
lasttwodecades.Applicationsofthemethodshavebeenwidenedfromtheirhistorical use in cancer and reliability research to business, criminology,epidemiology,andsocialandbehavioralsciences.Thethirdeditionof Statisti-
cal Methods for Survival Data Analysis isintendedtoprovideacomprehensive
introductionofthemostcommonlyusedmethodsforanalyzingsurvivaldata.Itbeginswithbasicdefinitionsandinterpretationsofsurvivalfunctions.Fromthere,the readerisguidedthroughmethods,parametricandnonparametric,forestimatingandcomparingthesefunctionsandthesearchforatheoreticaldistribution (or model )to fit the data. Parametric and nonparametric ap-
proachestotheidentificationofprognosticfactorsthatarerelatedtosurvivalare then discussed. Finally, regression methods, primarily linear logistic re-gressionmodels,toidentifyriskfactorsfordichotomousandpolychotomousoutcomesareintroduced.
The third edition continues to be application-oriented, with a minimum
levelofmathematics.Inafewchapters,someknowledgeofcalculusandmatrixalgebrais needed.Thefewsections thatintroducethe generalmathematicalstructureforthemethodscanbeskippedwithoutlossofcontinuity.Alargenumberofpracticalexamplesaregiventoassistthereaderinunderstandingthemethodsandapplicationsandininterpretingtheresults.Readerswithonlycollegealgebrashouldfindthebookreadableandunderstandable.
Therearemanyexcellentbooksonclinicaltrials.Wethereforehavedeleted
thetwochaptersonthesubjectthatwereinthesecondedition.Instead,wehaveincludeddiscussionsofmorestatisticalmethodsforsurvivaldataanalysis.A brief summary of the improvements made for the third edition is givenbelow.
1. Twoadditionaldistributions,thelog-logisticdistributionandageneral-
izedgammadistribution,havebeenaddedtotheapplicationofparamet-ric models that can be used in model fitting and prognostic factoridentification (Chapters6,7,and11 ).
xi
2. Inseveralsections (Sections7.1,9.1,10.1,11.2,and12.1 ),discussionsof
the asymptotic likelihood inference of the methods covered in thechaptersaregiven.Thesesectionsareintendedtoprovideamoregeneralmathematicalstructureforstatisticians.
3. The Cox —Snell residual method has been added to the chapter on
graphicalmethodsforsurvivaldistributionfitting (Chapter8 ).Inaddi-
tion,thesectionsonprobabilityandhazardplottinghavebeenrevised
sothatnospecialgraphicalpapersarerequiredtomaketheplots.
4. More tests of goodness of fit are given, including the BIC and AIC
procedures (Chapters9and11 ).
5. For Cox’s proportional hazards model (Chapter 12 ), we have now
includedmethodstoassessitsadequencyandprocedurestoestimatethesurvivorshipfunctionwithcovariates.
6. Theconceptofnonproportionalhazardsmodelsisintroduced (Chapter
13), which includes models with time-dependent covariates, stratified
models,competingrisksmodels,recurrenteventmodels,andmodelsforrelatedobservations.
7. Thechapteronlinearlogisticregression (Chapter14 )hasbeenexpanded
to cover regression models for polychotomous outcomes. In addition,methods for a general m:nmatching design have been added to the
sectiononconditionallogisticregressionforcase —controlstudies.
8. ComputerprogrammingcodesforsoftwarepackagesBMDP,SAS,and
SPSSareprovidedformostexamplesinthetext.
Wewouldliketothankthemanyresearchers,teachers,andstudentswho
haveusedthesecondeditionofthebook.Thesuggestionsforimprovementthatmanyofthemhaveprovidedareinvaluable.SpecialthanksgotoXingWang, Linda Hutton, Tracy Mankin, and Imran Ahmed for typing themanuscript. Steve Quigley of John Wiley convinced us to work on a thirdedition.Wethankhimforhisenthusiasm.
Finally,wearemostgratefultoourfamilies,Sam,Vivian,Benedict,Jennifer,
andAnnelisa (E.T.L. ),andAliceandXing (J.W.W. ),fortheconstantjoy,love,
andsupporttheyhavegivenus.
ET.L
JWW
Oklahoma City, OK
April 18, 2001xii
CHAPTER 1
Introduction
1.1 PRELIMINARIES
This book is for biomedical researchers, epidemiologists, consulting statisti-
cians, students taking a first course on survival data analysis, and othersinterested in survival time study. It deals with statistical methods for analyzingsurvival data derived from laboratory studies of animals, clinical and epi-demiologic studies of humans, and other appropriate applications.
Survival time can be defined broadly as the time to the occurrence of a given
event. This event can be the development of a disease, response to a treatment,relapse,ordeath.Therefore,survivaltime canbetumor-freetime,thetimefromthe start of treatment to response, length of remission, and time to death.Survival data can include survival time, response to a given treatment, andpatient characteristics related to response, survival, and the development of adisease. The study of survival data has focused on predicting the probability ofresponse, survival, or mean lifetime, comparing the survival distributions ofexperimentalanimalsor of humanpatientsand the identificationof risk and/orprognostic factors related to response, survival, and the development of adisease.In this book,specialconsiderationis givento thestudy ofsurvival datain biomedical sciences, although all the methods are suitable for applicationsin industrial reliability, social sciences, and business. Examples of survival datain these fields are the lifetime of electronic devices, components, or systems(reliability engineering ); felons’ time to parole (criminology ); duration of first
marriage (sociology ); length of newspaper or magazine subscription (market-
ing); and worker’s compensation claims (insurance )and their various influenc-
ing risk or prognostic factors.
1.2 CENSORED DATA
Many researchers consider survival data analysis to be merely the application
of twoconventionalstatisticalmethodsto a specialtypeof problem: parametric
if the distribution of survival times is known to be normal and nonparametric
1
if the distribution is unknown. This assumption would be true if the survival
times of all the subjects were exact and known; however, some survival timesare not. Further, the survival distribution is often skewed, or far from beingnormal. Thus there is a need for new statistical techniques. One of the mostimportant developments is due to a special feature of survival data in the lifesciences that occurs when some subjects in the study have not experienced theevent of interest at the end of the study or time of analysis. For example, somepatients may still be alive or disease-free at the end of the study period. The
exact survival times of these subjects are unknown. These are called censored
observations orcensored times and can also occur when people are lost to
follow-up after a period of study. When these are not censored observations,the set of survival times is complete. There are three types of censoring.
Type I Censoring
Animal studies usually start with a fixed number of animals, to which thetreatment or treatments is given. Because of time and/or cost limitations, theresearcher often cannot wait for the death of all the animals. One option is toobserve for a fixed period of time, say six months, after which the survivinganimals are sacrificed. Survival times recorded for the animals that died duringthe study period are the times from the start of the experiment to their death.These are called exactoruncensored observations . The survival times of the
sacrificedanimals are not known exactly but are recorded as at least the lengthof the study period. Theseare called censored observations. Someanimalscould
be lost or die accidentally. Their survival times, from the start of experimentto loss or death, are also censored observations. In type I censoring , if there are
no accidental losses, all censored observations equal the length of the studyperiod.
For example, suppose that six rats have been exposed to carcinogens by
injecting tumor cells into their foot pads. The times to develop a tumor of agiven size are observed. The investigator decides to terminate the experimentafter 30 weeks. Figure 1.1 is a plot of the development times of the tumors.Rats A, B, and D developed tumors after 10, 15, and 25 weeks, respectively.Rats C and E did not develop tumors by the end of the study; their tumor-freetimes are thus 30-plus weeks. Rat F died accidentally without tumors after 19weeks of observation. The survival data (tumor-free times )are 10, 15, 30 /p59, 25,
30/p59, and 19/p59weeks. (The plus indicates a censored observation. )
Type II Censoring
Another option in animal studies is to wait until a fixed portion of the animalshave died, say 80 of 100, after which the surviving animals are sacrificed. Inthis case, type II censoring , if there are no accidental losses, the censored
observations equal the largest uncensored observation. For example, in anexperimentof six rats (Figure 1.2 ), the investigatormay decide to terminate the
study after four of the six rats have developed tumors. The survival ortumor-free times are then 10, 15, 35 /p59, 25, 35, and 19 /p59weeks.2
Figure 1.1 Example of type I censored data.
Figure 1.2 Example of type II censored data.
Type III Censoring
In most clinical and epidemiologic studies the period of study is fixed andpatients enter the study at different times during that period. Some may diebefore the end of the study; their exact survival times are known. Others maywithdraw before the end of the study and are lost to follow-up. Still others maybe alive at the end of the study. For ‘‘lost’’ patients, survival times are at leastfrom their entrance to the last contact. For patients still alive, survival timesare at least from entry to the end of the study. The latter two kinds ofobservations are censored observations. Since the entry times are not simulta-neous, the censored times are also different. This is type III censoring . For
example, suppose that six patients with acute leukemia enter a clinical study 3
Figure 1.3 Example of type III censored data.
during a total study period of one year. Suppose also that all six respond to
treatment and achieve remission. The remission times are plotted in Figure 1.3.Patients A, C, and E achieve remission at the beginning of the second, fourth,and ninth months, and relapse after four, six, and three months, respectively.Patient B achieves remission at the beginning of the third month but is lost tofollow-up four months later; the remission duration is thus at least fourmonths. Patients D and F achieve remission at the beginning of the fifth andtenth months, respectively, and are still in remission at the end of the study;their remission times are thus at least eight and three months. The respectiveremission times of the six patients are 4, 4 /p59,6 ,8/p59, 3, and 3 /p59months.
Type I and type II censored observations are also called singly censored
data, and type III, progressively censored data , by Cohen (1965 ). Another
commonly used name for type III censoring is random censoring . All of these
types of censoring are right censoring orcensoring to the right . There are also
left censoring and interval censoring cases. L eft censoring occurs when it is
knownthattheevent ofinterestoccurredpriorto acertaintime t, but theexact
timeof occurrenceis unknown.For example,anepidemiologistwishes toknowthe age at diagnosis in a follow-up study of diabetic retinopathy. At the time ofthe examination, a 50-year-old participant was found to have already develop-ed retinopathy,but there is no recordof the exacttime at whichinitial evidencewas found. Thus the age at examination (i.e., 50 )is a left-censored observation.
It means that the age of diagnosis for this patient is at most50 years.
Interval censoring occurs when the event of interest is known to have
occurred between times aand b. For example, if medical records indicate that
at age 45, the patient in the example above did not have retinopathy, his ageat diagnosis is between 45 and 50 years.
We will study descriptive and analytic methods for complete, singly cen-
sored, and progressively censored survival data using numerical and graphical4
techniques.Analytic methods discussed include parametricand nonparametric.
Parametric approaches are used either when a suitable model or distributionis fitted to the data or when a distribution can be assumed for the populationfrom which the sampleis drawn. Commonly used survival distributions are theexponential,Weibull,lognormal,and gamma.If a survivaldistributionis foundto fit the data properly, the survival pattern can then be described by theparameters in a compact way. Statistical inference can be based on thedistribution chosen. If the search for an appropriate model or distribution is
too time consuming or not economicalor no theoreticaldistribution adequate-ly fits the data, nonparametric methods, which are generally easy to apply,should be considered.
1.3 SCOPE OF THE BOOK
This book is divided into four parts.
Part I (Chapters 1, 2, and 3 )defines survival functions and gives examples
of survival data analysis. Survival distribution is most commonly described bythree functions: the survivorship function (also called the cumulative survival
rate or survival function ), the probability density function, and the hazard
function (hazard rate or age-specific rate ). In Chapter 2 we define these three
functions and their equivalence relationships. Chapter 3 illustrates survivaldata analysis with five examples taken from actual research situations. Clinicaland laboratory data are systematically analyzed in progressive steps and theresults are interpreted. Section and chapter numbers are given for quickreference. The actual calculations are given as examples or left as exercises inthe chapters where the methods are discussed. Four sets of data are providedin the exercise section for the reader to analyze. These data are referred to inthe various chapters.
In Part II (Chapters 4 and 5 )we introduce some of the most widely used
nonparametric methods for estimating and comparing survival distributions.Chapter 4 deals with the nonparametric methods for estimating the threesurvival functions: the Kaplan and Meier product-limit (PL)estimate and the
life-table technique (population life tables and clinical life tables ). Also covered
is standardization of rates by direct and indirect methods, including thestandardized mortality ratio. Chapter 5 is devoted to nonparametric tech-niques for comparing survival distributions. A common practice is to comparethe survival experiences of two or more groups differing in their treatment orin a given characteristic. Several nonparametric tests are described.
Part III (Chapters 6 to 10 )introduces the parametric approach to survival
data analysis. Although nonparametric methods play an important role insurvival studies, parametric techniques cannot be ignored. In Chapter 6 weintroduce and discuss the exponential, Weibull, lognormal, gamma, andlog-logistic survival distributions. Practical applications of these distributionstaken from the literature are included. 5
An important part of survival data analysis is model or distribution fitting.
Once an appropriate statistical model for survival time has been constructedand its parameters estimated, its information can help predict survival, developoptimal treatment regimens, plan future clinical or laboratory studies, and soon. The graphical technique is a simple informal way to select a statisticalmodel and estimate its parameters. When a statistical distribution is found tofit the data well, the parameters can be estimated by analytical methods. InChapter 7 we discuss analytical estimation procedures for survival distribu-
tions. Most of the estimationprocedures are based on the maximum likelihoodmethod. Mathematical derivations are omitted; only formulas for the estimatesand examples are given. In Chapter 8 we introduce three kinds of graphicalmethods: probability plotting, hazard plotting, and the Cox —Snell residual
method for survival distribution fitting. In Chapter 9 we discuss several testsof goodness of fit and distribution selection. In Chapter 10 we describe severalparametric methods for comparing survival distributions.
A topic that has received increasing attention is the identification of
prognostic factors related to survival time. For example, who is likely tosurvivelongest after mastectomy,and what are the most important factorsthatinfluence that survival? Another subject important to both biomedical re-searchers and epidemiologists is identification of the risk factors related to thedevelopment of a given disease and the response to a given treatment. Whatare the factorsmost closely related to the developmentof a given disease? Whois more likely to develop lung cancer, diabetes, or coronary disease? In manydiseases, such as cancer, patients who respond to treatment have a betterprognosis than patients who do not. The question, then, relates to what thefactors are that influence response. Who is more likely to respond to treatmentand thus perhaps survive longer?
Part IV (Chapters 11 to 14 )deals with prognostic/risk factors and survival
times. In Chapter 11 we introduce parametric methods for identifying impor-tant prognostic factors. Chapters 12 and 13 cover, respectively, the Coxproportional hazards model and several nonproportional hazards models forthe identification of prognostic factors. In the final chapter, Chapter 14, weintroduce the linear logistic regression model for binary outcome variables andits extension to handle polychotomous outcomes.
In Appendix A we describe a numerical procedure for solving nonlinear
equations, the Newton —Raphson method. This method is suggested in Chap-
ters 7, 11, 12, and 13. Appendix B comprises a number of statistical tables.
Most nonparametric techniques discussed here are easy to understand and
simple to apply. Parametric methods require an understanding of survivaldistributions. Unfortunately, most of survival distributions are not simple.Readers without calculus may find it difficult to apply them on their own.However, if the main purpose is not model fitting, most parametric techniquescan be substituted for by their nonparametric competitors. In fact, a largepercentage of survival studies in clinical or epidemiological journals areanalyzed by nonparametric methods. Researchers not interested in survival6
modelfitting shouldreadthe chaptersand sectionsonnonparametricmethods.
Computerprograms for survivaldata analysis are available in several commer-cially available software packages: for example, BMDP, SAS, and SPSS. Thesecomputer programs are referred to in various chapters when applicable.Computer programming codes are given for many of the examples.
Bibliographical Remarks
Cross and Clark (1975 )was the first book to discuss parametric models and
nonparametric and graphical techniques for both complete and censoredsurvival data. Since then, several other books have been published in additionto the first edition of this book (Lee, 1980, 1992 ). Elandt-Johnsonand Johnson
(1980 )discuss extensively the construction of life tables, model fitting, compet-
ing risk, and mathematical models of biological processes of disease pro-gression and aging. Kalbfleisch and Prentice (1980 )focus on regression
problems with survival data, particularly Cox’s proportional hazards model.Miller (1981 )covers a number of parametric and nonparametric methods for
survival analysis. Cox and Oakes (1984 )also cover the topic concisely with an
emphasis on the examination of explanatory variables.
Nelson (1982 )providesa gooddiscussionofparametric,nonparametric,and
graphical methods. The book is more suited for industrial reliability engineersthan for biomedical researchers, as are Hahn and Shapiro (1967 )and Mann et
al.(1974 ). In addition, Lawless (1982 )gives a broad coverage of the area with
applications in engineering and biomedical sciences.
More recent publications include Marubini and Valsecchi (1994 ), Klein-
baum (1995 ), Klein and Moeschberger (1997 ), and Hosmer and Lemeshow
(1999 ). Most of these books take a more rigorous mathematical approach and
require knowledge of mathematical statistics. 7
CHAPTER 2
Functions of Survival Time
Survival time data measure the time to a certain event, such as failure, death,
response, relapse, the development of a given disease, parole, or divorce. Thesetimes are subject to random variations, and like any random variables, form adistribution. The distribution of survival times is usually described or charac-terized by three functions: (1)the survivorship function, (2)the probability
density function, and (3)the hazard function. These three functions are
mathematically equivalent — if one of them is given, the other two can bederived.
In practice, the three functions can be used to illustrate different aspects of
the data. A basic problem in survival data analysis is to estimate from thesampled data one or more of these three functions and to draw inferencesabout the survival pattern in the population. In Section 2.1 we define the threefunctions and in Section 2.2, discuss the equivalence relationship among thethree functions.
2.1 DEFINITIONS
LetTdenote the survival time. The distribution of Tcan be characterized by
three equivalent functions.
Survivorship Function (or Survival Function)
This function, denoted by S(t), is defined as the probability that an individual
survives longer than t:
S(t)/p58P(an individual survives longer than t)
/p58P(T/p57t) (2.1.1 )
From the definition of the cumulative distribution function F(t)o fT,
S(t)/p581-P(an individual fails before t)
/p581/p57F(t)( 2 .1.2)
8
Figure 2.1 Two examples of survival curves.HereS(t) is a nonincreasing function of time twith the properties
S(t)/p58/p71 for t/p580
0 for t/p58/p45
That is, the probability of surviving at least at the time zero is 1 and that of
surviving an infinite time is zero.
The function S(t) is also known as the cumulativesurvivalrate. To depict the
course of survival, Berkson (1942 )recommended a graphic presentation of S(t).
The graph of S(t) is called the survival curve. A steep survival curve, such as
the one shown in Figure 2.1 a, represents low survival rate or short survival
time. A gradual or flat survival curve such as in Figure 2.1 brepresents high
survival rate or longer survival.
The survivorship function or the survival curve is used to find the 50th
percentile (the median )and other percentiles (e.g., 25th and 75th )of survival
time and to compare survival distributions of two or more groups. The mediansurvival times in Figure 2.1 aandbare approximately 5 and 36 units of time,
respectively. The mean is generally used to describe the central tendency of adistribution, but in survival distributions the median is often better because asmall number of individuals with exceptionally long or short lifetimes willcause the mean survival time to be disproportionately large or small.
In practice, if there are no censored observations, the survivorship function
is estimated as the proportion of patients surviving longer than t:
S/p19(t)/p58number of patients surviving longer than t
total number of patients
(2.1.3 )
where the circumflex denotes an estimate of the function. When censored
observations are present, the numerator of (2.1.3 )cannot always be determined.
For example, consider the following set of survival data: 4, 6, 6 /p59,1 0/p59, 15, 20. 9
Figure 2.2 Two examples of density curves.Using (2.1.3 ), we can compute S/p19(5)/p585/6/p580.833. However, we cannot obtain
S/p19(11) since the exact number of patients surviving longer than 11 is unknown.
Either the third or the fourth patient (6/p59and 10 /p59)could survive longer than
or less than 11. Thus, when censored observations are present, (2.1.3 )is no
longer appropriate for estimating S(t). Nonparametric methods of estimating
S(t) for censored data are discussed in Chapter 4.
Probability Density Function (or Density Function)
Like any other continuous random variable, the survival time Thas a
probability density function defined as the limit of the probability that anindividual fails in the short interval ttot/p59/afii9773tper unit width /afii9773t, or simply the
probability of failure in a small interval per unit time. It can be expressed as
f(t)/p58lim/p9/p82/p29/p15P[an individual dying in the interval (t,t/p59/afii9773t)]
/afii9773t(2.1.4)
The graph of f(t) is called the density curve. Figure 2.2 aandbgive two
examples of the density curve. The density function has the following twoproperties:
1.f(t) is a nonnegative function:
f(t)/p460 for all t/p460
/p580 fort/p580
2. The area between the density curve and the taxis is equal to 1.
In practice, if there are no censored observations, the probability density
functionf(t) is estimated as the proportion of patients dying in an interval per10
unit width:
f/p19(t)/p58number of patients dying in the interval beginning at time t
(total number of patients )/p59(interval width )(2.1.5 )
Similar to the estimation of S(t), when censored observations are present,
(2.1.5 )is not applicable. We discuss an appropriate method in Chapter 4.
The proportion of individuals that fail in any time interval and the peaks of
high frequency of failure can be found from the density function. The densitycurve in Figure 2.2 agives a pattern of high failure rate at the beginning of the
study and decreasing failure rate as time increases. In Figure 2.2 b, the peak of
high failure frequency occurs at approximately 1.7 units of time. The propor-tion of individuals that fail between 1 and 2 units of time is equal to the shadedarea between the density curve and the axis. The density function is also knownas theunconditional failure rate.
Hazard Function
The hazard function h(t) of survival time Tgives the conditional failure rate.
This is defined as the probability of failure during a very small time interval,assuming that the individual has survived to the beginning of the interval, oras the limit of the probability that an individual fails in a very short interval,t/p59/afii9773t, given that the individual has survived to time t:
h(t)/p58lim/p9/p82/p29/p15P
/p3an individual fails in the time interval (t,t/p59/afii9773t)
given the individual has survived to t /p4
/afii9773t(2.1.6)
The hazard function can also be defined in terms of the cumulative
distribution function F(t) and the probability density function f(t):
h(t)/p58f(t)
1/p57F(t)(2.1.7)
The hazard function is also known as the instantaneous failure rate ,force of
mortality ,conditional mortality rate , andage-specific failure rate. Iftin(2.1.6 )
is age, it is a measure of the proneness to failure as a function of the age of theindividual in the sense that the quantity /afii9773th(t) is the expected proportion of
agetindividuals who will fail in the short time interval t/p59/afii9773t. The hazard
function thus gives the risk of failure per unit time during the aging process. Itplays an important role in survival data analysis.
In practice, when there are no censored observations the hazard function is
estimated as the proportion of patients dying in an interval per unit time, given 11
Figure 2.3 Examples of the hazard function.that they have survived to the beginning of the interval:
h/p19(t)/p58number of patients dying in the interval beginning at time t
(number of patients surviving at t)/p59(interval width )
/p58number of patients dying per unit time in the interval
number of patients surviving at t(2.1.8 )
Actuaries usually use the average hazard rate of the interval in which the
number of patients dying per unit time in the interval is divided by the averagenumber of survivors at the midpoint of the interval:
h/p19(t)/p58
number of patients dying per unit time in the interval
(number of patients surviving at t)/p57(number of deaths in the interval )/2
(2.1.9 )
The actuarial estimate in (2.1.9 )gives a higher hazard rate than (2.1.8 )and thus
a more conservative estimate.
The hazard function may increase, decrease, remain constant, or indicate a
more complicated process. Figure 2.3 is a plot of several kinds of hazardfunction. For example, patients with acute leukemia who do not respond totreatment have an increasing hazard rate, h/p16(t),h/p17(t) is a decreasing hazard
function that, for example, indicates the risk of soldiers wounded by bulletswho undergo surgery. The main danger is the operation itself and this dangerdecreases if the surgery is successful. An example of a constant hazard function,h/p18(t), is the risk of healthy persons between 18 and 40 years of age whose main
risks of death are accidents. The bathtub curve ,h/p19(t), describes the process of12
Table 2.1 Survival Data and Estimated Survival Functions of40 Myeloma Patients
Number of Patients
Surviving at Number of Patients
Survival Time Beginning of Dying int(months ) Interval Interval S/p19(t)f/p19(t)h/p19(t)
0—5 40 5 1.000 0.025 0.027
5—10 35 7 0.875 0.035 0.044
10—15 28 6 0.700 0.030 0.048
15—20 22 4 0.550 0.020 0.040
20—25 18 5 0.450 0.025 0.065
25—30 13 4 0.325 0.020 0.072
30—35 9 4 0.225 0.020 0.114
35—40 5 0 0.125 0.000 0.000
40—45 5 2 0.125 0.010 0.100
45—50 3 1 0.075 0.005 0.080
/p4650 2 2 0.050 — —human life. During an initial period, the risk is high (high infant mortality ).
Subsequently, h(t) stays approximately constant until a certain time, after
which it increases because of wear-out failures. Finally, patients with tubercu-losis have risks that increase initially, then decrease after treatment. Such anincreasing, then decreasing hazard function is described by h/p20(t).
Thecumulative hazard function is defined as
H(t)/p58
/p16/p82
/p15h(x)dx (2.1.10)
It will be shown in Section 2.2 that
H(t)/p58/p57 logS(t) (2.1.11 )
Thus, at t/p580,S(t)/p581,H(t)/p580, and at t/p58/p45,S(t)/p580,H(t)/p58/p45. The
cumulative hazard function can be any value between zero and infinity. All logfunctions in this book are natural logs (basee)unless otherwise indicated.
The following example illustrates how these functions can be estimated from
a complete sample of grouped survival times without censored observations.
Example 2.1 The first three columns of Table 2.1 give the survival data of
40 patients with myeloma. The survival times are grouped into intervals of fivemonths. The estimated survivorship function, density function, and hazardfunction are also given, with the corresponding graphs plotted in Figure2.4a—c. 13
Figure 2.4 Estimated survival functions of myeloma patients.14
Figure 2.4 (Continued).
The estimated survivorship function ,S/p19(t), is calculated following (2.1.3 )at the
beginning or the end of each interval. For example, at the beginning of the firstinterval, all 40 patients are alive, S/p19(0)/p581, and at the beginning of the second
interval, 35 of the 40 patients are still alive, S/p19(5)/p5835/40/p580.875. Similarly,
S/p19(10)/p5828/40/p580.700. The estimated density function f/p19(t) is computed follow-
ing(2.1.5 ). For example, the density function of the first interval (0—5)is
5/(40/p595)/p580.025, and that of the second interval (5—10)is 7/(40/p595)/p580.035.
The estimated density function is plotted at the midpoint of each interval(Figure 2.4 b). The estimated hazard function, h/p19(t), is computed following the
actuarial method given in (2.1.9 ). For example, the hazard function of the first
interval 5/[5 (40/p575/2)]/p580.027 and that of the second interval is 7/[5 (35/p577/
2)]/p580.044. The estimated hazard function is also plotted at the midpoint of
each interval (Figure 2.4 c).
From Table 2.1 or Figure 2.4 a, the median survival time of myeloma
patients is approximately 17.5 months, and the peak of high frequency of deathoccurs in 5 to 10 months. In addition, the hazard function shows an increasingtrend and reaches its peak at approximately 32.5 months and then fluctuates.
2.2 RELATIONSHIPS OF THE SURVIVAL FUNCTIONS
The three functions defined in Section 2.1 are mathematically equivalent. Given
any one of them, the other two can be derived. Readers not interested in themathematical relationship among the three survival functions can skip this 15
section without loss of continuity.
1. From (2.1.2 )and (2.1.7 ),
h(t)/p58f(t)
S(t)(2.2.1)
This relationship can also be derived from (2.1.6 )using basic definitions of
conditional probabilities.
2. Since the probability density function is the derivative of the cumulative
distribution function,
f(t)/p58d
dt[1/p57S(t)]/p58/p57S/p30(t)( 2 .2.2)
3. Substituting (2.2.2 )into (2.2.1 )yields
h(t)/p58/p57S/p30(t)
S(t)/p58/p57d
dtlogS(t)( 2 .2.3)
4.Integrating (2.2.3 )from zero to tand using S(0)/p581, we have
/p57/p16/p82
/p15h(x)dx/p58logS(t)
or
H(t)/p58/p57 logS(t)
or
S(t)/p58exp[/p57H(t)]/p58exp/p3/p57/p16/p82
/p15h(x)dx/p4(2.2.4 )
5. From (2.2.1 )and (2.2.4 )we obtain
f(t)/p58h(t) exp[ /p57H(t)] (2 .2.5)
Hence, iff(t) is known, the survivorship function can be obtained from the
basic relationship between f(t),F(t), and (2.1.2 ). The hazard function can then
be determined from (2.2.1 ).I fS(t) is known, f(t) andh(t) can be determined
from (2.2.2 )and(2.2.1 ), respectively, or h(t) can be derived first from (2.2.3 )and
thenf(t) from (2.2.1 ).I fh(t) is given, S(t) andf(t) can be obtained, respectively,
from (2.2.4 )and (2.2.5 ). Thus, given any one of the three survival functions, the
other two can easily be derived. The following example illustrates theseequivalence relationships.16
Example 2.2 Suppose that the survival time of a population has the
following density function:
f(t)/p58e/p92/p82t/p460
Using the definition of the cumulative distribution function,
F(t)/p58/p16/p82
/p15f(x)dx/p58/p16/p82
/p15e/p92/p86dx/p58/p57e/p92/p86/p11/p82
/p15/p581/p57e/p92/p82
From (2.1.2 )we obtain the survivorship function
S(t)/p58e/p92/p82
The hazard function can then be obtained from (2.2.1 ):
h(t)/p58e/p92/p82
e/p92/p82/p581
A complete treatment of this distribution is given in Section 6.1.
Bibliographical Remarks
The three survival functions and their equivalents are discussed in every text
cited in the Bibliographical Remarks in Chapter 1.
EXERCISES
2.1 Consider the survival data given in Exercise Table 2.1. Compute and plot
the estimated survivorship function, the probability density function, andthe hazard function.
Exercise Table 2.1
Year of Number Alive at Number Dying in
Follow-up Beginning of Interval Interval
0—1 1100 240
1—2 860 180
2—3 680 184
3—4 496 138
4—5 358 118
5—6 240 60
6—7 180 52
7—8 128 44
8—98 4 3 2
/p4695 2 2 8 17
2.2 Exercise Table 2.2 is a life table for the total population (of 100,000 live
births )in the United States, 1959 —1961. Compute and plot the estimated
survivorship function, the probability density function, and the hazardfunction.
Exercise Table 2.2
Age Number Living at Number Dying in
Interval Beginning of Age Interval Age Interval
0—1 100,000 2,593
1—5 97,407 409
5—10 96,998 233
10—15 96,765 214
15—20 96,551 440
20—25 96,111 594
25—30 95,517 612
30—35 94,905 761
35—40 94,144 1,080
40—45 93,064 1.686
45—50 91,378 2,622
50—55 88,756 4,045
55—60 84,711 5,644
60—65 79,067 7,920
65—70 71,147 10,290
70—75 60,857 12,687
75—80 48,170 14,594
80—85 33,576 15,034
85 and over 18,542 18,542
Source: U.S. National Center for Health Statistics, Life Tables 1959 —1961,
Vol. 1, No. 1, ‘‘United States Life Tables 1959 —61,’’ December 1964, pp. 8 —9.
2.3 Derive (2.2.1 )using (2.1.6 )and basic definitions of conditional probabil-
ity.
2.4 Given the hazard function
h(t)/p58c
derive the survivorship function and the probability density function.
2.5 Given the survivorship function
S(t)/p58exp(/p57t/p65)
derive the probability density function and the hazard function.18
CHAPTER 3
Examples of Survival Data Analysis
The investigator who has assembled a large amount of data must decide what
to do with it and what it indicates. In this chapter we take several sets ofsurvival data from actual research situations and analyze them. In Example 3.1we analyze two sets of data obtained, respectively, from two and threetreatment groups to compare the treatment’s abilities to prolong life. Example3.2 is an example of the life-table technique for large samples. Example 3.3 givesremission data from two treatments; the investigator seeks a well-knowndistribution for the remission patterns to compare the two groups. In Example3.4 we study survival data and several other patient characteristics to identifyimportant prognostic factors; the patient characteristics are analyzed individ-ually and simultaneously for their prognosticvalues. In Example 3.5 weintroduce a case in which the interest is to identify risk factors in thedevelopment of a given disease. Four sets of real data are presented in theexercises so that the reader can plan analysis.
3.1 EXAMPLE 3.1: COMPARISON OF TWO TREATMENTS
AND THREE DIETS
3.1.1 Comparison of Two Treatments
Thirty melanoma patients (stages 2 to 4 )were studied to compare the
immunotherapies BCG (Bacillus Calmette-Guerin )andCorynebacterium par-
vumfor their abilities to prolong remission duration and survival time. The age,
gender, disease stage, treatment received, remission duration, and survival timeare given in Table 3.1. All the patients were resected before treatment beganand thus had no evidence of melanoma at the time of first treatment.
The usual objective with this type of data is to determine the length of
remission and survival and to compare the distributions of remission andsurvival time in each group. Before comparing the remission and survival
19
Table 3.1 Data for 30 Resected Melanoma Patients
Initial Treatment Remission Survival
Patient Age Gender /p63Stage Received /p64Duration /p65 Time /p65
1 59 2 3B 1 33.7 /p59 33.7/p59
2 50 2 3B 1 3.8 3.9
3 76 1 3B 1 6.3 10.54 66 2 3B 1 2.3 5.4
5 33 1 3B 1 6.4 19.5
6 23 2 3B 1 23.8 /p59 23.8/p59
7 40 2 3B 1 1.8 7.9
8 34 1 3B 1 5.5 16.9 /p59
9 34 1 3B 1 16.6 /p59 16.6/p59
10 38 2 2 1 33.7 /p59 33.7/p59
11 54 2 2 1 17.1 /p59 17.1/p59
12 49 1 3B 2 4.3 8.0
13 35 1 3B 2 26.9 /p59 26.9/p59
14 22 1 3B 2 21.4 /p59 21.4/p59
f15 30 1 3B 2 18.1 /p59 18.1/p59
16 26 2 3B 2 5.8 16.0 /p59
17 27 1 3B 2 3.0 6.9
18 45 2 3B 2 11.0 /p59 11.0/p59
19 76 2 3A 2 22.1 24.8 /p59
20 48 1 3A 2 23.0 /p59 23.0/p59
21 91 1 4A 2 6.8 8.3
22 82 2 4A 2 10.8 /p59 10.8/p59
23 50 2 4A 2 2.8 12.2 /p59
24 40 1 4A 2 9.2 12.5 /p59
25 34 1 3A 2 15.9 24.4
26 38 1 4A 2 4.5 7.7
27 50 1 2 2 9.2 14.8 /p59
28 53 2 2 2 8.2 /p59 8.2/p59
29 48 2 2 2 8.2 /p59 8.2/p59
30 40 2 2 2 7.8 /p59 7.8/p59
Source: Data courtesy of Richard Ishmael.
/p631, male; 2, female.
/p641, BCG; 2, C. parvum.
/p65Remission and survival times are in months.
distributions, we attempt to determine if the two treatment groups are
comparable with respect to prognostic factors. Let us use the survival time toillustrate the steps. (The remission time could be analyzed similarly. )
1.Estimate and plot the survival function of the two treatment groups . The
resulting curves are called survival curves. Points on the curve estimate the
proportion of patients who will survive at least a given period of time. For suchsmall samples with progressively censored observations, the Kaplan —Meier
product-limit (PL)method is appropriate for estimating the survival function.20
EXAMPLES OF SURVIVAL DATA ANALYSIS
Table 3.2 Kaplan--Meier Product-Limit Estimate of Survival
Function S(t)
BCG Patients
Death time ( t) 3.9 5.4 7.9 10.5 19.5
S/p19(t) 0.909 0.818 0.727 0.636 0.477
C. parvum Patients
Death time ( t) 6.9 7.7 8.0
S/p19(t) 0.947 0.895 0.839
It does not require any assumptions about the form of the function that is
being estimated. We discuss this method in detail in Section 4.1. Computerprograms for the method can be found in BMDP (Dixon et al. 1990 ), SPSS
Version 10.1 (2000 ), and SAS Version 8.1 (2000 ). Examples for computer codes
will be given in Section 4.1.
Table 3.2 gives the PL estimate of the survival function S/p19(t) for the two
treatment groups. Note that S/p19(t) is estimated only at death times; however, the
censored observations were used to estimate S(t). Themedian survival time can
be estimated by linear interpolation. For BCG patients the median survivaltime was about 18.2 months. The median survival time for the C.parvum group
cannot be calculated since 15 of the 19 patients were still alive. Most computerprograms give not only S/p19(t) but also the standard error of S/p19(t), and the 75-,
50-, and 25-percentile points.
Figure 3.1 plots the estimated survival function S/p19(t) for patients receiving
the two treatments: The median survival time (50-percentile point )for the BCG
group can also be determined graphically. The survival curves clearly showthatC. parvum patients had slightly better survival experience than BCG
patients. For example, 50%of the BCG patients survived at least 18.2 months,whereas about 61% of the C. parvum patients survived that long.
2.Examine the prognostic homogeneity of the two groups . The next question
to ask is whether the difference in survival between the two treatment groupsis statistically significant. Is the difference shown by the data significant orsimply random variation in the sample? A statistical test of significance isneeded. However, a statistical test without considering patient characteristicsmakes sense only if the two groups of patients are homogeneous with respectto prognosticfac tors. It has been assumed thus far that the patients in the twogroups are comparable and that the only difference between them is treatment.Thus, before performing a statistical test it is necessary to examine thehomogeneity between the two groups.
Although prognosticfac tors for melanoma patients are not well established,
it has been reported that women and the young have a better survivalEXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 21
Figure3.1 Survival curves of patients receiving BCG and C. parvum.
experience than men and the elderly. Also, the disease stage plays an important
role in survival. Let us check the homogeneity of the two treatment groupswith respect to age, gender, and disease stage.
The age distributions are estimated and plotted in Figure 3.2. The median
age is 39 for the BCG group and 43 for the C. parvum patients. To test the
significance of the difference between the two age distributions, the two-samplet-test (Armitage, 1971; Daniel, 1987 )or nonparametrictests suc h as the
Mann—Whitney U-test or the Kolmogorov —Smirnov test (Marascuilo and
McSweeney, 1977 )are appropriate. However, the generalized Wilcoxon tests
given in Section 5.1 can also be used, since they reduce to the Mann —Whitney
U-test. Using Geham’s generalized Wilcoxon test, the difference between the
two age distributions is not found to be statistically significant. More about thetest will be given in Section 5.1.
The number of male and female patients in the two treatment groups is
given in Table 3.3. Sixty-four percent of the BCG patients and 42% of the C.
parvum patients are women. A chi-square test can be used to compare the two
proportions (see Section 14.1 ). It can be used only for r/p59ctables in which the
entries are frequencies, not for tables in which the entries are mean values ormedians of a certain variable. For a 2 /p592 table, the chi-square value can be
computed by hand. Computer programs for the test can be found in manycomputer program packages, such as BMDP (Dixon et al., 1990 ), SPSS
Version 10.1 (2000 ), and SAS Version 8.1 (SAS Institute, 2000 ).
The chi-square value for treatment by gender in Table 3.3 is 1.29 with 1
degree of freedom, which is not significant at the 0.05 or 0.10 level. Therefore,the difference between the two proportions is not statistically significant. Thenumber of stage 2 patients and the number of patients with more advanced22
EXAMPLES OF SURVIVAL DATA ANALYSIS
Figure3.2 Age distribution of two treatment groups.
Table 3.3 Treatment by Gender and Disease Stage
BCG C. parvum BCG C. Parvum
Disease
Gender Number % Number % Total Stage Number % Number % Total
Male 4 36 11 58 15 2 2 18 4 21 6
Female 7 64 8 42 15 3 and 4 9 82 15 79 24—— — —— —11 19 30 11 19 30disease in the two treatment groups are also given in Table 3.3. Eighteen
percent of the BCG patients are at stage 2 against 21% of the C. parvum
patients. However, a chi-square test result shows that the difference is notsignificant.
Thus we can say that the data do not show heterogeneity between the two
treatment groups. If heterogeneity is found, the groups can be divided intosubgroups of members who are similar in their prognoses.EXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 23
3.Compare the two survival distributions . There are several parametricand
nonparametric tests to compare two survival distributions. They are describedin Chapters 5 and 10. Since we have no information of the survival distributionthat the data follow, we would continue to use nonparametric methods tocompare the two survival distributions. The four tests described in Sections5.1.1 to 5.1.4 are suitable. The performance of these tests is discussed at the endof Section 5.1. We chose Gehan’s generalized Wilcoxon test here to demon-strate the analysis procedure only because of its simplicity of calculation.
In testing the significance of the difference between two survival distribu-
tions, the hypothesis is that the survival distribution of the BCG patients is thesame as that of the C. parvum patients. Let S/p16(t) andS/p17(t) be the survival
function of the BCG and C.parvum groups, respectively. The null hypothesis is
H/p15:S/p16(t)/p58S/p17(t)
The alternative hypothesis chosen is two-sided:
H/p16:S/p16(t)/p34S/p17(t)
since we have no prior information concerning the superiority of either of the
two treatments. The slight difference between the two estimated survival curvescould be due to random variation. The one-sided alternative H/p16:S/p16(t)/p58S/p17(t)
should be considered inappropriate.
Using Gehan’s generalized Wilcoxon test, the difference in survival distribu-
tion of the two treatment groups is found to be insignificant (p/p580.33).
Therefore, we do not reject the null hypothesis that the two survival distribu-tions are equal. Although our conclusion is that the data do not provideenough evidence to reject the hypothesis, ‘‘not to reject the null hypothesis’’does not automatically mean ‘‘to accept the null hypothesis.’’ The differencebetween the two statements is that the error probability of the latter statementis usually much larger than that of the former.
3.1.2 Comparison of Three Diets
A laboratory investigator interested in the relationship between diet and the
development of tumors divided 90 rats into three groups and fed them low-fat,saturated fat, and unsaturated fat diets, respectively (King et al., 1979 ). The rats
were of the same age and species and were in similar physical condition. Anidentical amount of tumor cells were injected into a foot pad of each rat. Therats were observed for 200 days. Many developed a recognizable tumor earlyin the study period. Some were tumor-free at the end of the 200 days. Rat 16in the low-fat group and rat 24 in the saturated group died accidentally after140 days and 170 days, respectively, with no evidence of tumor. Table 3.4 givesthe tumor-free time, the time from injection to the time that a tumor developsor to the end of the study. Fifteen of the 30 rats on the low-fat diet developeda tumor before the experiment was terminated. The rat that died had atumor-free time of at least 140 days. The other 14 rats did not develop any24
EXAMPLES OF SURVIVAL DATA ANALYSIS
Table 3.4 Tumor-Free Time (Days) of 90 Rats on Three Different Diets
Rat Low-Fat Rat Saturated Fat Rat Unsaturated Fat
1 140 1 124 1 112
2 177 2 58 2 683 50 3 56 3 844 65 4 68 4 109
5 86 5 79 5 153
6 153 6 89 6 143
7 181 7 107 7 608 191 8 86 8 709 77 9 142 9 98
10 84 10 110 10 16411 87 11 96 11 63
12 56 12 142 12 63
13 66 13 86 13 7714 73 14 75 14 9115 119 15 117 15 9116 140 /p59 16 98 16 66
17 200 /p59 17 105 17 70
18 200 /p59 18 126 18 77
19 200 /p59 19 43 19 63
20 200 /p59 20 46 20 66
21 200 /p59 21 81 21 66
22 200 /p59 22 133 22 94
23 200 /p59 23 165 23 101
24 200 /p59 24 170 /p59 24 105
25 200 /p59 25 200 /p59 25 108
26 200 /p59 26 200 /p59 26 112
27 200 /p59 27 200 /p59 27 115
28 200 /p59 28 200 /p59 28 126
29 200 /p59 29 200 /p59 29 161
30 200 /p59 30 200 /p59 30 178
Source: King et al. (1979 ). Data are used by permission of the author.
tumor by the end of the experiment; their tumor-free times were at least 200
days. Among the 30 rats in the saturated fat diet group, 23 developed a tumor,one died tumor-free after 170 days, and six were tumor-free at the end of theexperiment. All 30 rats in the unsaturated fat diet group developed tumorswithin 200 days. The two early deaths can be considered losses to follow-up.The data are singly censored if the two early deaths are excluded.
The investigator’s main interest here is to compare the three diets’ abilities
to keep the rats tumor-free. To obtain information about the distribution ofthe tumor-free time, we can first estimate the survival (tumor-free )function of
the three diet groups. The three survival functions were estimated using theEXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 25
Figure3.3 Survival curves of rats in three diet groups.
Kaplan—Meier PL method and plotted in Figure 3.3. The median tumor-free
times for the low-fat, saturated fat, and unsaturated fat groups were 188, 107,and 91 days, respectively. Since the three groups are homogeneous, we can skipthe step that checks for homogeneity and compare the three distributions oftumor-free time.
TheK-sample test described in Section 5.3.3 can be used to test the
significance of the differences among the three diet groups. Using this test, the
investigator finds that the differences among the three groups are highlysignificant (p/p580.002 ). Note that the K-sample test can tell the investigator
only that the differences among the groups are statistically significant. It cannottell which two groups contribute the most to the differences—whether thelow-fat diet produces a significantly different tumor-free time than thesaturated fat diet or whether the saturated fat diet is significantly different fromthe unsaturated fat diet. All one can conclude is that the data show a significantdifference among the tumor-free times produced by the three diets.
3.2 EXAMPLE 3.2: COMPARISON OF TWO SURVIVAL PATTERNS
USING LIFE TABLES
When the sample of patients is so large that their groupings are meaning-
ful, the life-table technique can be used to estimate the survival distribution.A method developed by Mantel and Haenszel (1959 )and applied to life26
EXAMPLES OF SURVIVAL DATA ANALYSIS
Table 3.5 Life Table for Male Patients with Localized Cancer of Rectum Diagnosed in
Connecticut, 1935--1944 and 1945--1954 /p63
1935—1944 1945 —1954
Interval
(t/p71) n/p30/p71d/p71w/p71/p59l/p71n/p71S/p19(t/p71)n/p30/p71d/p71w/p71/p59l/p71n/p71S/p19(t/p71)
1 388 167 2 387.0 0.5685 749 185 10 744.0 0.7513
2 219 45 1 218.5 0.4514 554 88 10 549.0 0.6309
3 173 45 1 172.5 0.3336 456 55 10 451.0 0.55394 127 19 0 127.0 0.2837 391 43 10 386.0 0.4922
5 108 17 0 108.0 0.2390 338 32 14 331.0 0.4446
6 91 11 1 90.5 0.2100 292 31 52 266.0 0.39287 79 8 0 79.0 0.1887 209 20 38 190.0 0.3514
8 71 5 0 71.0 0.1754 151 7 24 139.0 0.3337
9 66 6 1 65.5 0.1593 120 6 25 107.5 0.3151
10 59 7 0 59.0 0.1404 89 6 24 77.0 0.2905
Source: Myers (1969 ).
/p63Symbols: n/p30/p71, number of patients alive at beginning of interval t/p71;d/p71, number of patients dying
during interval t/p71;w/p71/p59l/p71, number of patients withdrawn alive or lost to follow-up during interval
t/p71;n/p71/p58n/p30/p71/p57/p16/p17(w/p71/p59l/p71);S/p19(t/p71), cumulative proportion surviving from beginning of study to end of
intervalt/p71.tables by Mantel (1966 )can be used to compare two survival patterns in the
life-table analysis.
Consider the data of male patients with localized cancer of the rectum
diagnosed in Connecticut from 1935 to 1954 (Myers, 1969 ). A total of 388
patients were diagnosed between 1935 and 1944, and 749 patients werediagnosed between 1945 and 1954. For such large sample sizes the data can begrouped and tabulated as shown in Table 3.5. The 10 intervals indicate thenumber of years after diagnosis. For the tabulated life tables the survival
functionS(t/p71) can be estimated for each interval t/p71. In Section 4.2 we discuss
the estimation procedures of S(t/p71) and density and hazard functions. The
survival, density, and hazard functions are the three most important functionsthat characterize a survival distribution.
TheS/p19(t/p71) column in Table 3.5 gives the estimated survival function for the
two time periods; these are plotted in Figure 3.4. Patients diagnosed in the1945—1954 period had considerably longer survival times (median 3.87 years )
than did patients diagnosed in the 1935 —1944 period (median 1.58 years ). The
five-year survival rate is frequently used by cancer researchers and can easilybe determined from a life table. Patients diagnosed in 1935 —1944 had a
five-year survival rate of 0.2390, or 23.9%. The patients diagnosed in 1945 —
1954 had a rate of 0.4446, or 44.5%. In comparing two sets of survival data,one can compare the proportions of patients surviving some stated period,such as five years, or the five-year survival rates. However, one cannotanticipate that two survival patterns will always stand in a superior —inferiorEXAMPLE 3.2: COMPARISON OF TWO SURVIVAL PATTERNS 27
Figure3.4 Survival curves for male patients with localized cancer of the rectum,
diagnosed in Connecticut, 1935 —1944 versus 1945 —1954.
relationship. It is more desirable to make a whole-pattern comparison (see
Sections 4.3 and 5.2 ).
The Mantel —Haenszel method described in Section 5.2 is a whole-pattern
comparison and can be used to compare two survival patterns in life tables.Application of this method to the data in Table 3.5 results in a chi-square valueof 51.996 with 1 degree of freedom. We can conclude that the difference
between the two survival patterns is highly significant (p/p580.001 ).
Estimates of the survival function or survival rate depend on the life-table
interval used. If each interval is very short, resulting in a large number ofintervals, the computation becomes very tedious and the life-table advantageis not fully taken. One assumption underlying the life table is that thepopulation has the same survival probability in each interval. If the intervallength is long, this assumption may be violated and the estimates inaccurate;this should be avoided except for rough calculations. Although the length ofeach interval and the total number of intervals are important, they will notcause trouble in most clinical studies since the study periods normally cover ashort period of time, such as one, two, or three years. Life tables with about10 to 20 intervals of several months to one year each are reasonable. Theinvestigator should also consider the disease under study. If the variation insurvival is large in a short period of time, the interval length should be short.However, in some demographicor other studies it is often of interest to c overa life span from birth to age 85 or more. The number of intervals would be28
EXAMPLES OF SURVIVAL DATA ANALYSIS
very large if short intervals were used. In this case five-year intervals are
sufficient to take into account the important variations in survival rateestimates (Shryock et al., 1971 ).
3.3 EXAMPLE 3.3: FITTING SURVIVAL DISTRIBUTIONS TO
REMISSION DATA
The remission times of 42 patients with acute leukemia were reported by
Freireich et al. (1963 )in a clinical trial undertaken to assess the ability of
6-mercaptopurine (6-MP )to maintain remission. /p16Each patient was ran-
domized to receive 6-MP or a placebo. The study was terminated after oneyear. The following remission times, in weeks, were recorded:
6-MP (21 patients ): 6, 6, 6, 7, 10, 13, 16, 22, 23, 6 /p59,9/p59,1 0/p59,1 1/p59,1 7/p59,
19/p59,2 0/p59,2 5/p59,3 2/p59,3 2/p59,3 4/p59,3 5/p59
Placebo (21 patients ): 1, 1, 2, 2, 3, 4, 4, 5, 5, 8, 8, 8, 8, 11, 11, 12, 12, 15, 17,
22, 23
Suppose that we are interested in a distribution to describe the remission times
of these patients but that no information is available as to which distributionwill fit. We need to find a distribution that fits the data well. If we can find one,the remission experience can then be described by the properties of thedistribution, and the remission time of new patients can be predicted. Paramet-ric tests can be used to compare the effectiveness of the two treatments, butsince there are a large number of well-known functions and distributions tochoose from, the search becomes an art as much as a scientific task.
The simplest and most efficient tool is the graph. Probability plotting can
be done for complete data; for data that include censored observations, hazardplotting and the Cox —Snell method are more appropriate. It is not difficult to
use the computer to generate these plots. Detailed discussions of probabilityplotting and hazard plotting are presented in Chapter 8. In both probabilityand hazard plotting, a linear configuration indicates that the distribution fitswell and its parameters can be estimated from the graph.
Let us begin by trying to fit a distribution to the remission duration of 6-MP
patients. Since the data consist of both censored and uncensored observations,we use the technique of hazard plotting. In this example we limit ourselves tothree distributions: the exponential, Weibull, and lognormal. In practice, moredistributions may need to be considered. Figures 3.5, 3.6, and 3.7 give thehazard plots for the exponential, Weibull, and lognormal distributions, respec-tively. A straight line is fitted to the points by eye in each of the plots. Amongthese graphs, the Weibull distribution appears to provide the best fit to theremission data. The straight line fits the points fairly closely. The estimates of
/p16Data are used by permission of the publisher.EXAMPLE 3.3: FITTING SURVIVAL DISTRIBUTIONS 29
Figure3.5 Exponential hazard plot of the remission times of 21 leukemia patients who
received 6-MP.
Figure3.6 Weibull hazard plot of the remission times of 21 leukemia patients who
received 6-MP.30 EXAMPLES OF SURVIVAL DATA ANALYSIS
/afii9818/p92/p16/p431/p57exp[H(t)]/p44
Figure3.7 Log-normal hazard plot of the remission times of 21 leukemia patients who
received 6-MP.
the parameters /afii9838and/afii9828of the Weibull distribution obtained from the line are
equal to 0.033 and 1.143, respectively (methods discussed in Chapter 8 ). After
knowing that the Weibull distribution provides a good fit, we can use ananalytical method, the maximum likelihood method, to obtain a more accurateestimate of the parameters. Following the procedures discussed in Chapter 7,
the maximum likelihood estimates of /afii9838and/afii9828are/afii9838/p19/p580.03 and /afii9828/p24/p581.354, which
are quite close to the graphical estimates.
After an appropriate distribution has been identified and parameters es-
timated, we can estimate the probability of having a given duration ofremission and other probabilities. For example, the probability of having aremission time longer than 10 weeks can be predicted as
P(T/p5710)/p58e/p92/p7/p16/p15
/afii9838/p19/p8/p65/p19/p58e/p57(10*0.03)/p16/p13/p18/p20/p19/p580.822
For the placebo group, we can use the probability plotting technique since the
data are complete. Figures 3.8, 3.9, and 3.10 give the exponential, Weibull, andlognormal probability plots. Comparing the three graphs, again, the straightline in the Weibull plot appears to give the best fit. From the Weibull plot,estimates of /afii9838and/afii9828are found to be 0.111 and 1.250, respectively. The
maximum likelihood estimates of /afii9838and/afii9828are, respectively, 0.105 and 1.371.
Again, the graphical estimates are very close to the maximum likelihoodestimates. Based on the maximum likelihood estimates of the parameters, weEXAMPLE 3.3: FITTING SURVIVAL DISTRIBUTIONS 31
log/p431/[1/p57F(t)]/p44
Figure3.8 Exponential probability plot of the remission times of 21 leukemia patients
who received placebo.
can estimate the probability of having a remission time longer than 10 weeks.
Using the same formula as given above, the probability for a patient receiving
placebo to have a remission duration longer than 10 weeks is found to be 0.34,which is smaller than that of a patient receiving 6-MP.
These graphical methods are subjective. The judgment as to whether the
assumed distribution fits the data is based on a visual examination rather thanon an objective statistical test. However, the methods are very simple and doprovide a great deal of information. Even in a case where none of thedistributions discussed in this book fit well, graphs can help find the reasonsand thus help modify the model. Therefore, graphical methods are usuallyrecommended as the first thing to try.
3.4 EXAMPLE 3.4: RELATIVE MORTALITY AND IDENTIFICATION
OF PROGNOSTIC FACTORS
One thousand and twelve Oklahoma Indians (379 men and 633 women )
with non-insulin-dependent diabetes mellitus (NIDDM )were examined in32
EXAMPLES OF SURVIVAL DATA ANALYSIS
log log /p431/[1/p57F(t)]/p44
Figure3.9 Weibull probability plot of the remission times of 21 leukemia patients who
received placebo.
/afii9818/p92/p16(F(t))
Figure3.10 Log-normal probability plot of the remission times of 21 leukemia patients
who received placebo.EXAMPLE 3.4: RELATIVE MORTALITY 33
1972—1980 and a mortality follow-up study was conducted in 1986 —1989 (Lee
et al., 1993 ). The mean [standard deviation (SD)] age and duration of diabetes
at baseline examination were 52 (11)and 7 (6)years. The average duration of
follow-up was 10 (SD 4 )years. As of December 31, 1989, 548 patients were
alive, 452 (187 men and 265 women )were dead, and 12 could not be traced.
Table 3.6 gives the survival time in years (T)of the first 40 male patients along
with 12 potential prognosticfac tors: age, duration of diabetes (DUR )in years,
family history of diabetes (FAM ), use of insulin within one year of diagnosis
(INS), use of diuretics (DIU ), hypertension (HBP ), retinopathy (EVD ), pro-
teinuria (PRO ), fasting plasma glucose (GLU )in milligrams per deciliter,
cholesterol (TC)in milligrams per deciliter, triglyceride (TG)in milligrams per
deciliter, and body mass index (BMI ), which is defined as weight in kilograms
divided by height in meters squared.
Among other things, the authors compared the mortality experience of the
diabeticpatients with that of the general population in Oklahoma over thefollow-up period. Taking changes in age distribution into consideration, thepatients were divided into five groups according to their age at baselineexamination: /p5835, 35—44, 45—54, 55—64, and /p4665. The expected survival rates
were calculated on a yearly basis following the methods described in Section4.3 and using the death rates given in the 1970 and 1980 Oklahoma populationlife tables. Death rates for the years between 1970 and 1980 and between 1980and 1989 were estimated based on the 1970 and 1980 statistics and theassumption that changes in death rates between 1970 and 1980 and after 1980follow a linear trend. The observed and expected survival curves for the groupswere plotted (Figure 3.11 ), and ratios of the observed and expected number of
deaths (O/E ratios )by age were tabulated (Table 3.7 ).
Figure 3.11 shows that the diabeticpatients had a muc h lower survivorship
than the general Oklahoma population for this age —gender distribution. At the
beginning of the fifteenth year after baseline examination, the relative survivalfor the diabeticOklahoma Indians was only 60%. The overall O/E ratios inTable 3.7 are 2.92 [or standardized mortality ratio (SMR )292] for men and
4.09 (or SMR 409 )for women, which indicates a significantly higher mortality
rate in the diabeticOklahoma Indians than in the general population.Although patients in every group experienced excessive mortality, the youngerpatients had the highest rate.
The relationship between the 12 potential prognosticvariables and the
survival time of men was examined using univariate and multivariate methods.The procedures are summarized below.
1.Examine the individual relationship of each variable to survival. One way
to analyze the data is first to determine which of the 12 variables could beconsidered of significant prognostic importance. In addition to correlationanalysis of these variables, the survival times in subcategories are compared(Table 3.8 ). Patients are grouped into subgroups in a meaningful way or in a
way that maximizes the observed difference in survival time between the34
EXAMPLES OF SURVIVAL DATA ANALYSIS
Figure3.11 Observed and expected survivorship from baseline examination for
dabeticOklahoma Indians.
subgroups (subject to the constraint that each subgroup contains at least 10%
of the total number of patients ).
The survivorship function for every subgroup of each variable was estimated
using the Kaplan —Meier method (discussed in Chapter 4 )and plotted. Figure
3.12 gives an example. Survival functions among the subgroups were comparedby the logrank test (one of the available tests discussed in Chapter 5 ). Table
3.8 shows that except cholesterol and triglyceride, every one of the 12 variablesis significant. The median survival time decreases as age and duration ofdiabetes increase. Patients with a family history of diabetes, elevated fastingplasma glucose, hypertension, or retinopathy have significantly shorter survivaldurations than those without these characteristics. Patients with baseline BMIvalues greater than or equal to 30 had much better survivorship than didpatients with a lower BMI value.EXAMPLE 3.4: RELATIVE MORTALITY 35
Table3.6 Data of First 40 MalePatie nts Enrolle d in Mortality Study of Oklahoma Diabe tic Indians
Patient Status /p63T Age DUR FAM /p64INS /p64DIU /p64HBP /p64EVD /p64PRO /p64GLU TC TG BMI
1 1 1 2 . 4 4 4 . 0 3 100001 2 4 2 3 9 2 5 3 8 3 4 . 2
2 1 1 4 . 1 4 3 . 5 1 0000009 4 1 5 8 9 4 4 2 . 2
3 0 1 4 . 4 4 7 . 6 4 101100 1 0 0 1 9 5 4 0 5 3 3 . 1
4 0 1 4 . 2 3 6 . 3 3 100000 1 7 1 2 1 2 2 1 8 3 8 . 55 0 1 4 . 4 5 4 . 4 7 000000 1 1 2 2 0 4 7 7 3 1 . 76 0 1 2 . 4 5 0 . 8 4 1010008 3 2 0 6 1 7 8 4 1 . 57 0 1 2 . 4 5 0 . 0 2 100000 1 0 4 1 7 8 1 0 0 3 9 . 58 17 . 0 6 6 . 9 8 000001 1 6 1 1 8 9 9 9 2 9 . 7
9 1 1 3 . 6 4 0 . 2 1 4 100110 2 6 2 2 4 1 3 0 1 2 7 . 5
1 0 0 1 4 . 4 5 4 . 1 4 101010 1 1 5 1 8 3 3 9 2 2 4 . 41 1 0 1 2 . 4 3 8 . 9 3 100101 1 0 8 2 3 7 4 9 3 2 . 41 2 19 . 8 5 1 . 2 4 100001 1 8 4 1 1 4 1 1 8 2 6 . 51 3 10 . 0 5 3 . 0 6 001001 1 2 6 2 0 6 4 8 0 3 4 . 51 4 1 1 2 . 1 4 5 . 0 5 001100 1 1 5 2 3 8 1 7 7 1 8 . 9
1 5 0 1 4 . 4 3 8 . 3 2 100000 1 1 0 2 0 4 1 8 0 3 1 . 2
1 6 0 1 2 . 4 4 0 . 0 5 101111 2 2 7 1 5 9 3 3 7 3 9 . 21 7 0 1 4 . 4 4 4 . 4 9 100000 1 8 2 1 9 3 3 3 2 3 2 . 71 8 0 1 4 . 2 4 8 . 2 5 101111 1 8 4 1 7 1 2 3 8 3 3 . 51 9 0 1 2 . 4 3 6 . 3 6 100000 2 3 8 1 9 6 1 6 2 2 4 . 22 0 0 1 3 . 7 4 1 . 9 3 100010 2 8 4 1 7 0 1 2 5 3 0 . 7 36
2 1 1 1 3 . 4 4 9 . 7 1 5 000000 2 4 8 2 1 7 2 3 4 2 8 . 0
2 2 1 1 2 . 6 5 1 . 5 9 1011109 6 1 7 4 1 4 4 2 4 . 22 3 0 1 4 . 0 4 4 . 1 1 100000 1 1 6 2 5 1 1 5 3 3 3 . 3
2 4 0 1 2 . 4 3 5 . 9 3 100000 1 2 0 1 2 9 7 6 3 0 . 1
2 5 0 1 2 . 9 5 0 . 4 1 000100 1 2 8 1 9 0 1 2 3 2 7 . 72 6 0 1 2 . 4 4 8 . 0 1 0011009 5 2 0 7 5 9 2 8 . 12 7 0 1 4 . 5 4 0 . 1 3 0000008 9 2 5 2 1 1 2 3 1 . 72 8 0 1 3 . 4 5 4 . 5 0 010100 1 0 4 4 9 0 5 4 0 3 0 . 82 9 1 1 3 . 9 6 9 . 3 6 101100 1 2 2 1 6 2 2 0 9 2 4 . 2
3 0 1 1 2 . 0 3 8 . 5 0 101101 2 2 5 1 8 3 3 4 3 4 3 . 1
3 1 1 3 . 6 6 3 . 7 1 8 000110 2 1 1 2 1 7 1 2 4 2 5 . 13 2 1 1 5 . 4 7 1 . 8 1 2 100000 1 5 0 2 2 7 1 3 7 2 6 . 03 3 0 1 0 . 3 5 9 . 5 3 101100 1 8 0 1 8 8 3 0 8 2 8 . 13 4 1 5 . 8 5 0 . 1 1 100000 4 0 0 2 0 0 1 6 6 2 6 . 13 5 1 2 . 5 7 5 . 4 1 8 100101 2 8 1 6 9 2 3 6 4 4 9 . 7
3 6 0 1 5 . 0 2 9 . 6 1 100001 1 6 7 1 5 4 1 5 7 6 0 . 2
3 7 1 5 . 5 6 0 . 2 1 8 000010 3 5 1 2 0 6 1 4 1 2 6 . 03 8 1 4 . 5 6 3 . 9 4 100101 1 2 7 1 8 0 1 3 1 2 1 . 8
3 9 1 6 . 8 5 7 . 4 1 7 100110 1 5 3 7 8 6 4 6 6 3 4 . 1
4 0 1 3 . 6 7 1 . 3 1 8 111111 1 7 9 4 8 8 3 6 4 2 5 . 6
/p631, dead; 0, alive.
/p641, yes; 0, no.
37
Figure3.12 Survival curves of diabetic patients by hypertension status at baseline.Table 3.7 Observed and Expected Number of Deaths and O /E Ratios During Follow-up
Period by Gender and Age at Baseline
Age at Males FemalesBaseline
(yr) Observed Expected O/E Observed Expected O/E
/p4544 34 6.48 5.25 31 5.51 5.63
45—54 79 25.64 3.08 85 18.40 4.62
55—64 32 11.67 2.74 69 15.03 4.59
65/p59 42 20.32 2.06 80 25.86 3.09
Total 187 64.11 2.92 265 64.80 4.09
2.Examine the simultaneous relationship of the variables to survival. Exam-
ination of each variable can give only a preliminary idea of which variablesmight be of prognostic importance. The simultaneous effect of the variablesmust be analyzed by an appropriate multivariate statistical method to deter-mine the relative importance of each. Cox’s (1972 )proportional hazards model38
EXAMPLES OF SURVIVAL DATA ANALYSIS
Table 3.8 Survival Time by Potential Prognostic Variable for Male Diabetic Patients
Number of Number of Median
Variable Patients Deaths Survival Time (yr)pValue
Age (yr)
/p5845 102 34 15.2
45—54 173 79 14.2/p580.00155—64 53 31 9.1
/p4665 46 43 5.1
Family history of diabetes
No 104 62 11.4/p580.01Yes 238 104 14.8
Duration of diabetes (yr)
/p587 207 84 15.2
7—13 91 49 12.2 /p580.001
/p4614 58 48 7.9
Use of diuretics
No 254 117 14.5/p580.001Yes 102 64 10.0
Use of insulin /p581 year
of diagnosis
No 317 157 13.9/p580.05Yes 39 24 11.9
Hypertension
No 211 84 15.3/p580.001Yes 151 99 9.8
Retinopathy
No 332 163 13.9/p580.001Yes 24 18 6.5
Proteinuria
Negative 250 112 14.5Slight 54 27 12.4 /p580.001
Heavy 57 44 8.0
Fasting plasma glucose
/p58200 235 106 11.9/p580.05/p46200 139 81 8.4
Cholesterol
/p58240 300 144 14.80.12/p46240 64 38 12.2
Triglyceride
/p58220 223 105 14.10.44/p46220 141 77 13.2
BMI
/p5830 189 114 11.8/p580.001/p4630 184 73 15.4EXAMPLE 3.4: RELATIVE MORTALITY 39
can be applied. This model, presented in Chapter 12, is a regression model
that relates patient characteristics directly to the risk of failure and thusindirectly to survival. The assumption of this model is that the hazards fordifferent strata of each independent (or prognostic )variable are proportional
over time. This assumption was verified by a graphical method (discussed in
Chapter 12 )using each of the variables. Figure 3.13 gives an example of the
graph of log[ /p57S(t)] versustfor the two hypertension groups. The two almost
parallel curves indicate that the hazards of dying are proportional. Therefore,
the assumption of the proportional hazards model is satisfied and the modelappropriate.
This model can be fitted by a stepwise procedure that results in a ranking
of the prognosticvariables. The first variable selec ted to enter the model is themost important single variable in predicting the risk of dying. The secondvariable is the second most important, and so on. A significance level can beobtained from a likelihood ratio test at each step, which indicates the level ofcontribution given by the additional variable.
Using the proportional hazards model and a stepwise procedure, seven of
the 12 variables were identified as significant at the 0.05 level based on thelikelihood ratio test at each step. These variables, the regression coefficients,and the significance levels based on the Ward test, which uses the regressioncoefficient and its standard error (S.E.)are given in Table 3.9. The sign of the
coefficient indicates whether the variable is positively or negatively related tothe hazard of dying. For example, age and duration of diabetes are bothpositively related to the risk of dying and therefore negatively related to thesurvival time. Table 3.9 also gives the ratio of risk (or hazard )for values of
each variable unfavorable to survival to values of that variable favorable tosurvival. For example, patients who were 60 years of age at baseline had a 3.05times higher risk of dying during the follow-up period (10—16 years, average
13 years )than did patients who were only 40 years old at baseline. For
dichotomous variables, the ratio of risk is equal to exp (coefficient ), which is
also interpreted as the relative risk of the variable adjusting for the othervariables. Consequently, the confidence interval for the relative risk can becalculated (not shown in Table 3.9 ). Based on this set of data, the authors
conclude that age, hypertension, duration of diabetes, fasting plasma glucose,BMI, proteinuria, and use of diuretics are significantly related to survival. Themultivariate method also showed that high values of BMI might be protective.
3.5 EXAMPLE 3.5: IDENTIFICATION OF RISK FACTORS
A study of the incidence of retinopathy in Oklahoma Indians with NIDDM
was conducted in 1987 —1990 as part of a prospective study of diabetic
complications (Lee et al., 1992 ). Among the 312 patients who were free of
retinopathy at initial examination in the 1970s, 228 were found to have40
EXAMPLES OF SURVIVAL DATA ANALYSIS
Figure3.13 Curves of log[ /p57logS(t)] for the two hypertension groups.
Table 3.9 Significant Variables (at 0.05 Level) Identified by Proportional Hazards
Model
Relative Risk /p64 Ratio
Regression pValue of
Variable /p63 Coefficient (Ward Test )Favorable Unfavorable Risk
Age 0.0558 /p580.001 9.32 28.45 3.05
Hypertension: 0.6360 /p580.001 1.00 1.89 1.89
1, yes, 0, no
Duration of 0.0559 /p580.001 1.32 2.19 1.66
diabetes
Fasting plasma 0.0023 /p580.010 1.35 1.58 1.17
glucose
BMI /p570.0330 0.035 0.32 0.44 0.72
Proteinuria: 0.3744 0.025 1.00 1.45 1.45
1, yes, 0, no
Use of diuretics: 0.4191 0.030 1.00 1.52 1.52
1, yes; 0, no
/p63Variables are listed in order of entry into model with a p-value limit for entry of 0.05.
/p64Favorable categories are 40 years of age, no hypertension, duration of diabetes 5 years, fasting
plasma glucose 130mg/dL, BMI 35, no proteinuria, and no diuretics use. Unfavorable categoriesare 60 years of age, hypertensive, duration of diabetes 14 years, fasting plasma glucose 200mg/dL,
BMI 25, having proteinuria, and diuretics use.EXAMPLE 3.5: IDENTIFICATION OF RISK FACTORS
41
developed the eye disease during the 10 to 16-year follow-up period (average
follow-up time 12.7 years ). Twelve potential factors (assessed at time of baseline
examination )were examined by univariate and multivariate methods for their
relationship to retinopathy (RET ): age, gender, duration of diabetes (DUR ),
fasting plasma glucose (GLU ), initial treatment (TRT ), systolic (SBP )and
diastolicblood pressure (DBP ), body mass index (BMI ), plasma cholesterol
(TC), plasma triglyceride (TG), and presence of macrovascular disease (LVD )
or renal disease (RD). Table 3.10 gives the data for the first 40 patients. Among
other things, the authors related these variables to the development ofretinopathy.
1.Examine the individual relationship of each variable to the development of
diabetic retinopathy. Table 3.11 gives some summary statistics of the eight
continuous variables for patients who have developed retinopathy and forthose who have not. Notice that patients who have developed the disease wereyounger at baseline and had much higher fasting plasma glucose, systolic anddiastolicblood pressure, and plasma triglyc eride than did patients who havenot. Table 3.12 summarizes the contingency table analysis of retinopathyincidence rates. The number of patients at risk of developing retinopathy andthe number of patients who developed the disease (and rate )are given by
subcategory of each potential risk factor. Using the chi-square test, it is foundthat there was a significant difference in the retinopathy rate among thesubcategories of several variables using a significance level of 0.05: duration ofdiabetes, fasting plasma glucose, systolic and diastolic blood pressure, andtreatment. It appears that patients with poor glucose control or high bloodpressure or treated with oral agents or insulin have a higher incidence ofretinopathy. In addition, patients with high triglyceride levels tend to havehigher incidence of retinopathy (p/p580.064). However, patients who had
developed macrovascular disease at the time of baseline examination had alower retinopathy incidence. The authors state that this may be due to the factthat 68% of the patients who had macrovascular disease either died (54% )
during the follow-up period or were lost to follow-up (14% ). Many of these
patients may have developed retinopathy, particularly the patients who havedied, but were not included. Therefore, the lower incidence of retinopathy inpatients who had macrovascular disease at baseline is probably the result of aselection bias. Similarly, the large number of death plus the losses to follow-upmay also contribute to the drop in retinopathy rate in patients who had haddiabetes for more than 12 years at baseline. Among the 80 patients in thisduration of diabetes category, 56% have died and 10% did not participate inthe follow-up examination. The large number of deaths may also be responsiblefor the finding that patients who survived long enough to develop retinopathywere younger at baseline. The deceased patients were significantly older (mean
57 years )than the survivors who participated in the follow-up examination
(mean 48 years ).42
EXAMPLES OF SURVIVAL DATA ANALYSIS
Table 3.10 First 40 Patients Involved in Study of Risk Factors in Development of Diabetic Retinopathy
Patient RET /p63Age Gender DUR TRT /p64GLU SBP DBP TC TG BMI LVD /p63RD /p63
1 1 47.6 M 4 2 100 156 98 195 405 33.1 0 0
2 0 54.4 M 7 1 112 112 78 204 77 31.7 0 03 0 50.8 M 4 2 83 134 80 206 178 41.5 1 04 1 49.4 F 5 2 276 102 70 190 222 24.1 0 05 1 50.0 M 2 2 104 142 86 178 100 39.5 0 06 1 50.7 F 7 2 242 142 78 217 268 31.6 0 1
7 1 35.3 F 2 2 130 134 80 390 564 47.0 0 0
8 0 50.2 F 5 2 130 100 70 174 128 29.8 0 09 1 45.0 M 5 2 115 134 100 238 177 18.9 1 0
10 0 38.3 M 2 2 110 132 80 204 180 31.2 1 011 0 45.8 F 5 3 130 118 68 185 316 26.6 1 012 1 51.6 F 3 2 141 112 78 152 77 31.2 0 1
13 1 36.3 M 6 2 238 142 94 194 162 24.2 0 0
14 1 44.7 F 16 2 190 152 90 132 161 32.0 0 015 1 37.2 F 1 2 126 136 90 133 211 32.7 0 016 1 52.6 F 9 2 159 140 76 151 132 26.7 0 017 1 44.1 M 1 1 116 126 76 251 153 33.3 0 018 1 35.9 M 3 1 120 132 80 129 76 30.1 0 0
19 1 50.4 M 1 2 128 144 90 190 123 27.7 0 0
20 1 48.0 M 1 2 95 128 74 207 59 28.1 1 021 0 47.5 F 1 2 85 124 82 161 190 31.6 1 022 1 50.1 F 6 2 138 106 72 181 135 30.5 0 1
(Continued overleaf )
43
Table3.10 Continued
Patient RET /p63Age Gender DUR TRT /p64GLU SBP DBP TC TG BMI LVD /p63RD /p63
23 0 43.3 F 1 1 104 128 86 204 198 26.1 0 0
24 1 54.5 M 0 3 104 142 84 490 540 30.8 1 025 1 52.2 F 3 3 304 132 84 192 119 36.9 1 026 1 53.3 F 3 2 249 128 72 120 85 35.5 0 027 1 64.3 F 6 2 297 138 80 145 64 30.1 1 028 1 44.6 F 4 1 139 112 80 156 111 36.1 0 1
29 1 47.1 F 2 2 169 130 84 198 99 39.0 0 1
30 1 46.5 F 7 1 159 128 78 238 157 34.5 0 031 0 51.5 F 3 1 147 128 78 185 182 32.0 0 032 1 59.5 M 3 2 180 132 78 188 308 28.1 1 033 1 52.0 F 6 2 183 142 84 175 68 37.6 0 034 1 45.7 F 6 2 180 138 80 179 189 44.1 1 0
35 1 48.2 F 18 2 267 158 100 195 112 25.2 1 0
36 1 57.4 M 0 1 159 172 108 219 294 33.7 0 137 1 42.0 F 4 2 158 106 68 224 157 33.6 0 038 1 50.7 F 1 2 211 142 84 390 645 37.1 0 039 1 53.8 F 0 1 177 154 80 175 208 35.3 0 140 0 56.9 M 1 1 98 116 70 146 97 26.5 0 1
/p631, yes; 0, no.
/p641, diet only; 2, oral agent; 3, insulin.
44
Table 3.11 Summary Statistics for Eight Variables by Retinopathy Status at Follow-up
Retinopathy Status
No Yes
Variable Mean S.D. Mean S.D. pValue
Age 50.0 9.0 47.2 7.4 0.01
Duration of diabetes 4.2 4.5 4.8 4.4 0.34
Fasting plasma glucose 141.8 65.6 196.3 76.6 /p580.0001
Systolicblood pressure 128.0 15.7 132.6 17.3 0.04Diastolicblood pressure 80.3 10.8 84.9 10.1 /p580.001
Body mass index 32.3 6.3 32.5 5.9 0.76Cholesterol 204.4 66.0 206.8 58.7 0.76
Triglyceride 180.5 111.1 234.4 273.3 0.01
2.Examine the simultaneous relationship of the variables to the development
of retinopathy. Univariate analysis of each variable using the contingency table
or the chi-square test gives a preliminary idea of which individual variablemight be of prognostic importance. The simultaneous effect of all the variablescan be analyzed by the linear logistic regression model (discussed in Section
14.2)to determine the relative importance of each.
The 12 variables were fitted to the linear logisticregression model using a
stepwise selection procedure. The variables most significantly related to thedevelopment of retinopathy were found to be initial treatment, fasting plasmaglucose, age, and diastolic blood pressure (p/p450.001). Table 3.13 gives the
regression coefficients of the four most significant variables (p/p450.05), the
standard errors, and adjusted odds ratios [exp (coefficient )]. Thepvalues used
here are the significance levels based on the likelihood ratio test or theimprovement in the maximum likelihood due to the addition of the variable inthe stepwise procedure. This method is more powerful than the Wald test,which is based on the standardized regression coefficients (Chapter 14 ). The
results are consistent with those in the univariate analysis.
On the basis of the regression coefficients, the probability of developing
retinopathy during a 10 to 16-year follow-up can be estimated by substitutingvalues of the risk factors into the regression equation,
logP
1/p57P
/p58/p572.373/p591.495(oral agent )/p590.882 (insulin )
/p590.014 (GLU )/p570.074 (age)/p590.048 (DBP )
For example, for a 50-year-old patient who is on oral agents and whose fasting
plasma glucose and diastolic blood pressure are 170 mg/dl and 95 mmHg,EXAMPLE 3.5: IDENTIFICATION OF RISK FACTORS 45
Table 3.12 Cumulative Incidence Rates of Retinopathy by Baseline Variables
Developed Retinopathy
Number of
Variable Persons at Risk Number Percent pvalue
Gender
Female 211 151 71.60.384Male 101 77 76.2
Age (yr)
/p5835 13 10 76.9
35—44 101 77 76.20.24245—54 155 115 74.2
/p4655 43 26 60.5
Duration of diabetes (yr)
/p584 153 105 68.6
4—7 113 86 76.10.0338—11 23 22 95.7
/p4612 23 15 65.2
Fasting plama glucose (mg/dl )
/p58140 117 62 53.0
140—199 90 74 82.2 /p580.001
/p46200 105 92 87.6
Systolicblood pressure (mmHg )
/p58130 145 95 65.5
130—159 149 115 78.8 0.016
/p46160 20 18 85.7
Diastolicblood pressure (mmHg )
/p5885 179 118 65.9
85—94 87 73 83.9 0.004
/p4695 46 37 80.4
Plasma cholesterol (mg/dl )
/p58240 267 193 72.30.442/p46240 45 35 77.8
Plasma triglyceride (mg/dl )
/p58250 237 167 70.50.064/p46250 75 61 81.3
Body mass index (kg/m /p17)
/p5828 73 49 67.1
28—33 121 94 77.7 0.261
/p4634 118 85 72.0
Renal disease
No 251 179 71.30.155Yes 61 49 80.3
Macrovascular disease
No 205 157 76.60.053Yes 107 71 66.4
Treatment (initial )
Diet alone 115 62 53.9
Oral agent 158 136 86.1 /p580.001
Insulin 37 29 78.446 EXAMPLES OF SURVIVAL DATA ANALYSIS
Table 3.13 Results of Logistic Regression Analysis
Standard
Variable Coefficient Error exp (coefficient )Coefficient/S.E.
Constant /p572.373 1.557
Initial treatment
Oral agent 1.495 0.330 4.459 4.53
Insulin 0.882 0.488 2.416 1.81
Fasting plasma
0.014 0003 1.014 4.67 glucose
Age /p570.074 0019 0.929 /p573.89
Diastolicblood
0.048 0.015 1.049 3.20 pressure
respectively, the chance of developing retinopathy in the next 10 to 16 years is
91%.
The linear logisticregression model is useful in identifying important risk
factors. However, complete measurements of all the variables are needed;missing data are a problem. In this example, complete data are available onmost of the patients. This may not always be the case. Although there aremethods of coping with missing data (discussed in Section 11.1 ), none is
perfect. Thus it is extremely important for investigators to make every effort toobtain complete data on every subject.
Bibliographical Remarks
It is impossible to cite all the published examples of survival data analysis
similar to those in this chapter. Other similar studies can be found in theliterature: for example, Biometrics ,Biometrika ,Cancer,Journal of Chronic
Disease,Journal of the National Cancer Institute ,American Journal of Epi-
demiology ,Journal of the American Medical Association , andNew England
Journal of Medicine . An easy way to find examples is to use the National
Library of Medicine’s Web site and search the file PubMed with appropriatekeywords.
EXERCISES
The four sets of data below are taken from actual research situations. Although
the data can be used for various analyses throughout the book, the reader isasked here only to describe in detail how the data can be analyzed. The dataappear in examples and other exercises in subsequent chapters.
3.1Thirty-three patients with hypernephroma were treated with combined
chemotherapy,immunotherapy,andhormonaltherapy.ExerciseTable3.1EXERCISES 47
Exercise Table 3.1 Data for 33 Patients with Hypernephroma
Date of
Date Death or Skin Test Results /p65
Treatment Last
Patient Age Gender Started Response /p63Follow-up Status /p64Monilia Mumps PPD PHA SK-SD
1 53 F 3/31/77 1 10/1/77 0 7 /p5972 3 /p5923 0 /p5902 5 /p5925 0 /p590
2 61 M 6/18/76 0 8/21/76 1 10 /p5910 15 /p5920 0 /p5901 3 /p5913 9 /p599
3 53 F 2/1/77 3 10/1/77 0 0 /p5907 /p5970 /p5902 5 /p5925 0 /p590
4 48 M 12/19/74 2 1/15/76 1 0 /p5900 /p5900 /p5900 /p5900 /p590
5 55 M 11/10/75 0 1/15/76 1 12 /p5912 ND 10 /p5910 8 /p5985 /p595
6 62 F 10/7/74 2 4/5/75 1 10 /p5910 5 /p5950 /p5907 /p5975 /p595
7 57 M 10/28/74 0 1/6/75 1 15 /p5915 15 /p5915 0 /p5900 /p5901 0 /p5910
8 53 M 10/6/75 2 6/18/77 1 0 /p590N D 0 /p5901 2 /p5912 0 /p590
9 45 M 4/11/77 0 10/1/77 0 6 /p5944 /p5940 /p5900 /p5900 /p590
10 58 M 8/4/76 3 2/11/77 1 13 /p5913 13 /p5913 22 /p5922 23 /p5923 0 /p590
11 61 F 1/1/77 3 10/1/77 0 0 /p5908 /p5981 7 /p5917 11 /p5911 0 /p590
12 61 M 7/25/76 1 10/1/77 0 9 /p5991 2 /p5912 0 /p5902 0 /p5920 0 /p590
13 77 M 5/8/75 0 9/26/75 1 0 /p5900 /p5900 /p5900 /p5900 /p590
14 55 M 4/27/77 2 10/1/77 0 0 /p5900 /p5901 5 /p5915 10 /p5910 0 /p590
15 50 M 4/20/77 3 10/1/77 0 0 /p5901 4 /p5914 5 /p5953 2 /p5932 21 /p5921
16 42 M 8/24/76 0 10/1/77 0 11 /p5911 7 /p5970 /p5901 2 /p5912 0 /p590
48
17 50 F 1/8/75 0 6/30/75 1 0 /p5900 /p5900 /p5900 /p5900 /p590
18 66 F 9/8/76 3 10/1/77 0 9 /p5991 0 /p5910 6 /p5961 5 /p5915 11 /p5911
19 58 M 2/18/75 0 10/1/77 0 0 /p5900 /p5900 /p5900 /p590N D
20 62 M 5/12/76 0 10/17/76 1 2 /p592N D N D 3 /p5932 /p592
21 71 F 10/22/76 3 12/12/76 1 10 /p5910 6 /p5960 /p5901 2 /p5912 0 /p590
22 44 M 6/6/77 3 10/1/77 0 10 /p5910 10 /p5910 0 /p5902 0 /p5920 0 /p590
23 69 M 6/21/76 0 10/13/76 1 0 /p5901 5 /p5915 25 /p5925 25 /p5925 0 /p590
24 56 M 6/7/77 2 10/1/77 0 0 /p5907 /p5970 /p5900 /p5900 /p590
25 57 M 11/16/76 0 12/10/76 1 11 /p5911 5 /p5950 /p5902 0 /p5920 0 /p590
26 69 M 5/10/77 0 7/25/77 1 0 /p5900 /p5900 /p5901 5 /p5915 0 /p590
27 60 M 6/29/77 0 7/7/77 1 0 /p5900 /p5900 /p5902 6 /p5926 0 /p590
28 60 M 7/21/75 3 10/1/77 0 11 /p5911 20 /p5920 10 /p5910 18 /p5918 0 /p590
29 72 M 7/19/75 0 10/18/75 1 10 /p5910 0 /p5907 /p5971 0 /p5910 0 /p590
30 42 F 3/3/75 0 4/23/75 1 0 /p590N D 0 /p5900 /p5900 /p590
31 57 M 2/24/77 2 10/1/77 0 5 /p5958 /p5980 /p5902 5 /p5915 0 /p590
32 66 M 6/15/77 3 10/1/77 0 0 /p5901 5 /p5915 0 /p5901 0 /p5910 0 /p590
33 59 M 3/4/77 0 4/2/77 1 0 /p5900 /p5900 /p5901 6 /p5916 0 /p590
Source: Data courtesy of Richard Ishmael.
/p630, no response; 1, complete response; 2, partial response; 3, stable.
/p640, alive; 1, dead.
/p65ND, not done.
49
gives the age, gender, date treatment began, response status, date of death
or last follow-up, survival status, and results of five pretreatment skintests. The investigator is interested in the response and survival of thepatients and in identifying prognosticfac tors. How would you analyze thedata?
3.2In a study undertaken to compare the treatments given to hyperneph-
roma patients and to relate response and survival to surgery, metastasis,
and treatment time, data from 58 patients were collected (Exercise Table
3.2). How would you analyze the data to answer these questions?
(a)Do patients who had nephrectomy have a higher response rate?
(b)Is the time of nephrectomy related to response and survival?
(c)Are there significant differences between the treatments?
(d)What are the most important variables related to response and
survival?
3.3Exercise Table 3.3 gives the age, gender, family history of melanoma,
remission duration, survival time, stage, and results of six pretreatmentskin tests (the larger diameter is given )of 102 stage 3 and 4 melanoma
patients (Lee et al., 1982 ).
(a)Study the immunocompetence of melanoma patients by investigating
skin test results.
(b)Determine if age, gender, or pretreatment skin test results are predic-
tive to remission and survival time.
(c)Find theoretical distributions that describe the survival and remission
patterns.
3.4One hundred and forty-nine diabeticpatients were followed for 17 years
(a subset of data from Lee et al., 1988 ). Exercise Table 3.4 gives the
survival time from baseline examination, survival status, and severalpotential prognosticfac tors at baseline: age, body mass index (BMI ), age
at diagnosis of diabetes, smoking status, systolicblood pressure (SBP ),
diastolicblood pressure (DBP ), electrocardiogram reading (ECG ), and
whether the patient had any coronary heart disease (CHD ). Identify the
important prognostic factors that are associated with survival.50
EXAMPLES OF SURVIVAL DATA ANALYSIS
Exercise Table 3.2 Data of 58 Patients with Hypernephroma
Time of Survival Lung Bone
Patient Gender Age Nephrectomy /p64Nephrectomy /p65Treatment /p66Response /p65Time Status /p68Metastasis /p64Metastasis /p64
12 5 3 1 0 . 0 1 17 7 01 0
21 6 9 1 4 . 0 1 21 8 10 1
31 6 1 0 /p579.0 1 0 8 1 1 0
42 5 2 1 2 . 0 1 26 8 11 0
51 4 6 1 2 . 0 1 23 5 10 161 5 5 1 0 . 0 1 0 8 11 072 6 2 1 0 . 0 1 22 6 11 081 5 3 1 0 . 0 1 28 4 10 191 7 0 0 /p579.0 1 0 17 1 1 0
10 1 48 1 0.0 1 3 52 1 1 0
11 1 58 1 1.5 1 0 26 1 1 112 1 61 1 5.0 1 1 108 0 1 013 1 77 1 4.0 1 0 18 1 1 014 1 56 1 0.0 1 3 72 1 0 115 1 55 1 0.0 1 2 38 1 1 1
16 1 50 1 4.0 1 3 /p5799 9 1 0
17 1 75 1 0.0 1 0 9 1 1 018 1 43 1 2.0 1 3 56 1 0 019 1 69 1 1.0 1 2 36 1 1 120 2 59 1 1.5 1 2 108 1 0 121 2 71 1 0.0 1 2 10 1 1 0
22 1 56 1 0.0 1 2 36 1 1 0
23 1 57 0 /p579.0 1 0 6 1 1 0
24 1 69 1 8.0 1 0 9 1 1 0
(Continued overleaf )
51
Exercise Table 3.2 Continued
Time of Survival Lung Bone
Patient Gender Age Nephrectomy /p64Nephrectomy /p65Treatment /p66Response /p65Time Status /p68Metastasis /p64Metastasis /p64
25 1 72 0 /p579.0 1 0 12 1 1 0
26 1 67 1 0.0 1 9 5 0 1 027 1 41 1 2.0 1 2 104 0 1 0
28 1 77 1 10.0 1 0 6 1 1 0
29 2 63 1 2.0 1 3 115 1 0 130 2 42 1 12.0 1 0 9 1 0 0
31 1 59 0 /p579.0 1 0 21 1 0 0
32 1 62 1 5.0 1 0 14 1 1 033 1 65 1 0.0 1 0 52 1 1 0
34 2 53 0 /p579.0 1 0 9 1 0 1
35 1 57 1 0.0 1 2 48 1 1 036 2 60 0 /p579.0 1 0 15 1 1 1
37 1 59 1 0.0 1 0 5 1 1 038 1 75 1 0.0 2 3 28 0 1 039 2 53 1 0.0 2 2 25 0 1 0
40 2 67 1 5.0 2 3 25 0 0 1
41 1 58 1 8.0 2 3 40 1 1 142 1 62 1 8.0 2 0 16 1 1 143 1 69 0 /p579.0 2 0 8 1 1 1
52
44 1 44 1 0.0 2 2 70 0 1 0
45 1 60 1 1.0 2 0 6 1 0 146 1 57 0 /p579.0 2 0 8 1 1 1
47 1 45 1 2.0 2 4 12 1 1 0
48 2 50 1 1.0 2 4 20 1 1 049 1 58 0 /p579.0 2 4 8 1 1 0
50 1 51 1 0.0 2 3 /p5799 0 0 1
51 1 59 1 3.0 2 3 12 1 1 052 1 53 1 0.0 2 1 181 0 1 0
53 1 70 1 0.5 2 0 20 1 1 1
54 1 69 1 3.0 2 0 14 1 1 055 1 62 1 0.0 2 3 26 1 0 056 1 52 1 2.0 2 0 16 1 1 057 1 77 1 2.0 2 2 30 1 1 058 1 61 1 8.0 2 0 20 1 1 0
Source: Data courtesy of Richard Ishmael.
/p631, male; 2, female.
/p641, yes; 0, no.
/p65Number of years prior to treatment; negative value—no nephrectomy.
/p661, combined chemotherapy and immunotherapy, 2, others.
/p670, no response; 1, complete response; 2, partial response; 3, stable; 4, increasing disease; 9, unknown.
/p681, dead; 0, alive; 9, unknown.
53
Exercise Table 3.3 Data of 102 Patients with Stages 3 and 4 Melanoma
Family Remission Survival Skin Tests /p68
History of Time Remission Time Survival
Patient Age Gender Melanoma /p64 (months )/p65Status /p66 (months )Status /p67Stage Monilia Mumps PPD PHA SK-SD Tricophyton
1 58 2 9 42.0 0 42.0 0 3B 18 16 0 20 14 99
25 0 2 0 3 . 3 13 . 9 1 3 B 7 8 0 0 70
37 6 1 0 6 . 1 1 1 0 . 5 1 3 B 0 1 5 0 0 9 90
46 6 2 9 2 . 3 16 . 0 0 3 B 8 0 0 1 0 005 33 1 9 5.1 1 20.6 1 3B 17 10 0 5 18 99
6 55 2 0 11.1 0 21.8 0 4B 99 99 99 99 99 99
7 25 2 0 36.5 0 36.5 0 3B 7 5 0 8 90 998 23 1 0 24.3 0 24.3 0 3B 10 20 0 7 30 0
9 30 1 0 28.7 0 28.7 0 3B 8 99 0 6 10 17
10 34 1 9 7.7 0 7.7 0 3B 15 99 0 0 0 011 34 1 0 29.3 0 29.3 0 3B 15 99 0 7 10 0
12 26 2 0 5.9 1 19.3 1 3AB 0 99 0 42 25 0
13 27 1 9 2.6 1 6.9 1 3AB 15 99 0 10 10 0
14 72 2 9 16.7 0 18.0 0 4B 10 99 99 15 0 7
15 70 2 0 14.6 0 14.6 0 3B 0 99 0 16 0 20
16 82 2 0 /p5799.0 9 23.6 0 4B 9 99 99 17 10 0
17 43 1 9 /p5799.0 9 3.9 1 4B 25 0 0 7 20 5
18 52 1 9 /p5799.0 9 7.3 1 4B 12 99 0 5 15 0
19 34 1 9 /p5799.0 9 9.8 1 4B 13 20 0 15 30 30
20 48 1 0 26.5 0 26.5 0 4A 10 99 0 88 5 0
21 62 1 9 18.0 1 25.4 0 3AB 5 99 0 10 0 24
22 49 1 9 4.3 1 8.0 1 3B 0 99 8 3 0 023 46 1 0 0.3 1 13.8 1 3B 0 5 0 7 0 0
24 53 2 1 21.5 0 21.5 0 3A 16 20 14 14 18 12
25 21 2 9 /p5799.0 9 9.3 1 4B 10 5 0 10 15 0
26 25 1 9 /p5799.0 9 1.2 1 4B 15 10 5 18 10 5
27 35 2 0 /p5799.0 9 20.0 1 4B 11 7 0 10 0 2
54
28 66 2 9 /p5799.0 9 12.5 0 4B 99 99 99 99 99 99
29 54 2 0 /p5799.0 9 7.4 1 4B 99 99 99 99 99 99
30 43 2 0 /p5799.0 9 4.7 0 4B 13 19 9 11 30 0
31 40 1 0 13.3 0 13.3 0 3B 0 5 25 12 10 032 16 1 0 0.0 0 0.0 0 3B 7 10 14 10 0 0
33 59 1 0 /p5799.0 9 25.8 0 4B 0 99 0 18 0 35
34 64 1 9 16.5 0 16.5 0 3B 8 7 0 0 0 035 52 1 9 /p5799.0 9 2.5 1 4B 9 14 0 12 5 0
36 /p5799 2 8 /p5799.0 9 13.8 1 4B 30 75 0 35 0 12
37 27 1 0 /p5799.0 9 4.2 1 4B 10 12 10 10 0 3
38 60 2 0 5.4 1 11.4 1 4B 20 10 0 10 0 0
39 73 1 9 /p5799.0 9 5.8 1 4B 0 9 0 40 0 30
40 50 2 0 13.5 0 13.5 0 3B 0 8 0 10 5 041 63 2 0 /p5799.0 9 2.7 1 4B 0 10 0 15 0 10
42 56 1 0 /p5799.0 9 0.9 1 4B 6 6 0 15 0 0
43 62 2 0 2.1 1 8.0 1 3A 0 8 0 32 0 2044 57 1 9 /p5799.0 9 0.0 0 3AB 0 24 11 20 0 0
45 56 2 0 12.1 1 16.1 0 3B 0 5 0 25 6 30
46 41 2 1 /p5799.0 9 13.3 1 4B 99 99 99 99 99 99
47 40 2 0 10.1 0 10.1 0 4A 0 15 0 15 17 20
48 81 2 0 /p5799.0 9 0.0 0 4B 99 99 99 99 99 99
49 61 1 0 8.4 0 8.4 0 3B 0 4 0 10 0 850 62 2 0 7.7 0 7.7 0 3AB 0 11 0 16 23 0
51 34 1 0 15.1 1 24.4 1 3B 0 9 15 5 20 0
52 62 2 9 1.1 1 10.5 1 4A 8 99 0 99 0 053 63 2 9 /p5799.0 9 22.2 1 4B 0 0 0 8 3 0
54 56 1 9 /p5799.0 9 7.4 1 4B 0 99 0 6 0 0
55 66 2 9 /p5799.0 9 1.3 1 4B 0 0· 0 22 0 0
56 62 1 0 11.1 0 20.5 1 4B 25 20 0 10 15 18
57 68 2 0 /p5799.0 9 13.8 1 4B 20 99 0 17 15 15
58 45 1 0 /p5799.0 9 6.3 1 4B 28 15 0 10 17 50
59 58 1 9 /p5799.0 9 8.5 0 3B 10 17 0 25 30 25
60 55 1 9 /p5799.0 9 5.8 0 4B 99 99 99 99 99 99
(Continued Overleaf )
55
Exercise Table 3.3 Continued
Family Remission Survival Skin Tests /p68
History of Time Remission Time Survival
Patient Age Gender Melanoma /p64 (months )/p65Status /p66 (months )Status /p67Stage Monilia Mumps PPD PHA SK-SD Tricophyton
61 63 2 1 7.3 1 8.7 0 3B 0 9 0 23 12 15
62 53 1 9 36.4 0 36.4 0 3B 5 35 20 17 6 063 45 1 0 /p5799.0 9 5.9 1 4B 0 0 0 0 0 0
64 41 1 0 /p5799.0 9 1.7 1 4B 0 99 0 5 0 6
65 43 1 9 /p5799.0 9 3.9 1 4B 25 0 0 7 20 5
66 80 1 0 5.8 0 11.0 1 4B 99 99 99 99 99 99
67 75 2 9 /p5799.0 9 3.8 1 4B 0 99 0 6 0 0
68 47 2 9 /p5799.0 1 15.9 0 3B 0 0 0 20 15 0
69 64 2 9 6.7 0 6.7 0 3AB 0 5 0 18 0 0
70 38 1 9 /p5799.0 9 1.6 1 4B 99 99 99 99 99 99
71 27 1 0 6.0 0 6.0 0 3B 8 15 20 27 20 1072 56 1 9 /p5799.0 9 4.1 0 4B 0 0 0 0 0 0
73 60 2 9 /p5799.0 9 2.8 0 3A 99 99 99 99 99 99
74 80 2 9 /p5799.0 9 0.2 0 4B 0 20 99 40 0 20
75 38 1 9 /p5799.0 9 7.0 0 4B 0 0 0 15 12 12
76 71 1 9 6.2 0 6.2 0 4A 99 99 99 99 99 99
77 57 2 0 6.1 0 6.1 0 4B 28 20 0 19 20 2078 69 1 0 /p5799.0 9 2.1 0 4B 15 15 15 10 0 0
79 17 2 9 4.9 0 4.9 0 3B 99 99 99 99 99 99
80 64 2 0 /p5799.0 9 1.6 0 4B 99 99 99 99 99 99
81 91 1 0 6.5 0 8.3 1 4A 99 99 99 99 99 99
82 40 2 0 1.7 1 4.6 0 3B 99 99 99 99 99 99
56
83 63 1 0 7.3 0 28.0 1 3A 99 99 99 99 99 99
84 40 1 9 /p5799.0 9 16.1 1 4B 99 99 99 99 99 99
85 53 1 9 /p5799.0 9 4.5 1 4B 99 99 99 99 99 99
86 41 1 0 21.2 0 21.2 0 3A 99 99 99 99 99 9987 27 1 9 /p5799.0 9 4.0 1 4B 99 99 99 99 99 99
88 /p5799 /p5790 /p5799.0 9 7.8 1 4B 99 99 99 99 99 99
89 45 2 9 /p5799.0 9 4.4 0 4B 99 99 99 99 99 99
90 50 2 9 /p5799.0 9 4.2 1 4B 99 99 99 99 99 99
91 47 1 9 /p5799.0 9 1.5 1 4B 99 99 99 99 99 99
92 63 1 9 /p5799.0 9 3.5 0 4B 99 99 99 99 99 99
93 52 1 9 /p5799.0 9 0.4 1 4B 99 99 99 99 99 99
94 53 1 9 /p5799.0 9 2.5 1 4B 99 99 99 99 99 99
95 60 2 9 /p5799.0 9 1.1 0 4B 99 99 99 99 99 99
96 35 1 9 /p5799.0 9 11.1 1 4B 99 99 99 99 99 99
97 24 2 9 1.2 0 1.2 0 3B 99 99 99 99 99 99
98 80 2 0 /p5799.0 9 1.9 0 3A 99 99 99 99 99 99
99 /p5799 /p579 0 4.6 1 6.7 0 3B 99 99 99 99 99 99
100 60 2 9 0.9 0 0.9 0 3AB 99 99 99 99 99 99
101 60 2 9 /p5799.0 9 4.3 0 4B 99 99 99 99 99 99
102 35 1 0 5.2 0 5.2 0 3B 99 99 99 99 99 99
Source: Lee et al. (1979 )
/p631 ,m a l e ;2 ,f e m a l e ; /p579, unknown.
/p641, yes; 0, No; 9, unknown.
/p65/p5799; never in remission during study period.
/p661, relapsed; 0, still in remission; 9, never in remission during study period.
/p671, dead; 0, still alive.
/p68In millimeters; 99, unknown.
57
Exe rciseTable3.4 Data of 149 Diabe tic Patie nts
Variable at Baseline
Survival Age at
Time Age Diagnosis Smoking SBP DBP
Patient Status /p63 (yr)( yr)BMI (yr) Status /p64 (mmHg )(mmHg )ECG /p65CHD /p66
1 1 12.4 44 34.2 41 0 132 96 1 0
2 1 12.4 49 32.6 48 2 130 72 1 03 1 9.6 49 22.0 35 2 108 58 1 14 1 7.2 47 37.9 45 0 128 76 2 15 1 14.1 43 42.2 42 2 142 80 1 06 1 14.1 47 33.1 44 0 156 94 1 0
7 1 12.4 50 36.5 48 0 140 86 2 1
8 1 14.2 36 38.5 33 2 144 88 1 09 1 12.4 50 41.5 47 1 134 78 1 1
10 1 14.5 49 34.1 45 0 102 68 1 011 1 12.4 50 39.5 48 2 142 84 1 012 1 10.8 54 42.9 43 0 128 74 1 0
13 0 10.9 42 29.8 36 2 156 86 1 0
14 1 10.3 44 33.2 43 2 102 58 1 015 0 13.6 40 27.5 26 2 146 98 1 016 1 11.9 48 25.3 48 0 120 68 2 117 1 12.5 50 31.6 44 1 142 76 1 018 1 5.9 47 26.3 38 1 144 82 1 0
19 1 12.4 38 32.4 36 2 150 98 2 1
20 1 14.1 35 47.0 33 1 134 78 1 021 0 9.8 51 26.5 47 2 130 76 1 022 1 7.2 40 43.9 34 0 122 92 1 0
58
23 1 3.5 54 32.3 52 1 132 80 1 0
24 1 0.0 53 34.5 47 2 150 88 3 125 0 12.1 45 18.9 40 1 134 98 1 0
26 1 1.9 41 32.0 31 1 142 90 2 1
27 1 8.6 34 33.9 30 2 124 66 1 028 1 14.0 38 23.7 28 0 102 60 1 029 1 14.3 43 24.8 43 0 134 80 1 030 1 12.4 45 26.6 41 2 118 66 2 131 1 12.4 40 39.2 35 2 192 108 1 0
32 1 14.4 44 32.7 36 2 122 78 1 0
33 1 14.2 48 33.5 43 1 122 92 1 034 1 14.5 51 32.2 49 2 112 74 1 035 1 12.4 36 24.2 30 2 142 90 1 036 1 14.3 52 31.6 48 1 152 96 1 037 0 13.7 41 30.7 39 2 112 74 1 0
38 1 13.4 49 28.0 35 2 118 84 1 0
39 1 12.5 44 32.0 29 0 152 88 1 040 1 14.4 37 32.7 36 2 136 88 1 0
41 1 12.6 51 24.2 42 2 134 90 1 0
42 1 13.8 47 18.7 42 0 130 78 2 143 1 14.0 45 25.6 36 0 108 72 1 0
44 1 6.8 38 22.8 27 2 126 66 2 1
45 1 12.4 35 30.1 33 0 132 78 1 046 1 12.9 50 27.7 49 1 144 88 1 047 1 8.9 53 27.6 49 2 126 68 1 048 1 12.4 48 28.1 47 1 128 70 1 049 1 14.5 40 31.7 37 2 132 82 1 0
50 1 13.0 43 26.1 42 2 128 80 1 0
51 1 13.4 54 30.8 54 1 142 80 2 152 1 10.6 52 36.9 50 1 132 80 2 153 1 13.9 69 24.2 63 1 148 78 1 0
(Continued overleaf )
59
Exe rciseTable3.4 Continued
Variable at Baseline
Survival Age at
Time Age Diagnosis Smoking SBP DBP
Patient Status /p63 (yr)( yr)BMI (yr) Status /p64 (mmHg )(mmHg )ECG /p65CHD /p66
54 1 16.9 38 27.5 26 2 170 100 1 0
55 1 3.6 50 27.3 44 1 140 90 1 056 1 10.2 64 30.1 58 0 138 76 2 157 1 15.7 44 36.1 41 0 112 78 1 058 1 12.0 38 43.1 39 2 140 78 1 059 0 6.7 62 34.6 58 0 138 78 3 1
60 1 11.6 47 39.0 45 0 130 82 1 0
61 0 2.0 78 28.7 77 0 178 86 2 162 1 10.2 49 28.2 43 2 158 80 1 063 1 3.6 63 25.1 46 1 168 88 3 164 1 15.4 71 26.0 59 0 146 88 1 065 1 11.3 51 32.0 49 2 128 76 1 0
66 1 10.3 59 28.1 57 1 132 76 1 1
67 1 5.8 50 26.1 49 1 154 80 1 068 0 8.0 66 45.3 49 0 154 92 1 069 1 14.6 42 30.0 41 1 122 80 1 070 1 11.4 40 35.7 36 2 144 76 2 171 1 7.2 67 28.1 61 0 178 96 1 0
72 1 5.5 86 32.9 61 0 162 60 1 0
73 1 11.1 52 37.6 46 1 142 80 1 074 1 16.5 42 43.4 37 0 120 76 1 075 1 10.9 60 25.4 60 0 124 64 1 076 1 2.5 75 49.7 57 1 174 82 2 1
60
77 0 10.8 81 35.2 81 0 142 88 1 0
78 1 4.7 60 37.3 39 0 160 78 1 079 0 5.5 60 26.0 42 0 122 68 3 1
80 1 4.5 63 21.8 60 2 162 98 1 1
81 1 9.0 62 18.2 43 0 132 72 2 182 1 6.8 57 34.1 41 2 116 60 3 183 0 3.6 71 25.6 54 1 152 84 3 184 1 12.1 58 35.1 45 0 144 68 2 185 1 8.1 42 32.5 28 1 98 68 3 1
86 1 11.1 45 44.1 40 0 138 76 1 1
87 0 7.0 66 29.7 59 1 138 78 1 088 1 1.5 61 29.2 54 0 184 80 2 189 1 11.7 48 25.2 30 2 158 98 1 090 1 0.3 82 25.3 50 0 176 96 1 191 1 13.6 35 25.8 34 1 118 72 1 0
92 1 15.0 57 33.7 57 2 172 98 1 0
93 1 11.2 56 39.5 55 1 182 100 1 194 1 3.0 49 32.9 48 0 144 90 2 1
95 1 13.7 50 37.1 50 0 142 80 1 0
96 1 10.2 53 35.3 53 2 154 76 1 097 1 12.4 71 29.3 70 0 122 60 1 0
98 1 1.1 55 22.1 33 2 222 102 2 1
99 1 16.3 69 23.6 43 0 150 80 1 1
100 1 6.7 59 26.1 55 2 142 66 1 0101 1 15.4 47 32.5 45 2 128 82 1 0102 0 7.6 75 29.8 67 0 122 76 3 1103 0 3.6 80 24.4 80 1 162 88 2 1
104 1 11.5 57 26.3 54 0 172 82 2 1
105 1 13.5 52 30.8 46 2 132 70 1 1106 1 10.6 48 29.4 46 0 112 68 1 0
(Continued overleaf )
61
Exe rciseTable3.4 Continued
Variable at Baseline
Survival Age at
Time Age Diagnosis Smoking SBP DBP
Patient Status /p63 (yr)( yr)BMI (yr) Status /p64 (mmHg )(mmHg )ECG /p65CHD /p66
107 0 6.5 57 29.1 47 1 138 92 2 1
108 0 14.3 58 30.1 56 0 128 74 1 0109 1 11.6 51 31.0 37 2 132 78 1 1110 1 15.4 33 34.0 33 2 120 78 1 0111 1 11.0 36 38.1 33 1 122 70 1 0112 0 11.0 52 37.0 46 0 140 98 1 0
113 0 4.8 64 31.2 57 2 172 88 3 1
114 1 14.8 31 38.8 29 1 136 76 1 0
115 1 1.8 69 22.3 56 0 152 74 3 1
116 1 15.8 59 25.0 58 0 126 80 1 0117 1 14.1 38 31.3 38 2 104 58 1 0118 1 4.6 49 59.7 49 1 142 82 1 0
119 1 15.5 49 34.0 41 0 128 76 1 0
120 0 7.2 68 29.4 66 1 122 58 3 1121 1 14.5 40 43.2 41 1 122 70 1 0122 1 10.5 36 35.1 32 2 122 68 1 0123 1 14.3 60 37.0 54 0 122 70 1 0124 0 2.2 74 27.1 54 1 168 84 2 1
125 1 5.0 61 27.6 51 0 162 82 1 0
126 1 12.4 54 25.2 51 0 116 76 1 0
62
127 1 1.1 35 25.8 34 2 126 82 1 0
128 1 15.4 46 32.2 42 2 180 98 1 0129 1 14.3 40 41.6 41 2 132 98 1 0
130 1 15.6 53 39.8 52 0 150 88 1 0
131 0 12.5 66 26.6 54 1 106 70 1 1132 1 12.3 61 33.3 55 0 154 88 1 0133 1 14.8 41 27.7 38 1 122 76 1 0134 1 10.2 64 26.6 51 2 130 68 1 0135 1 12.3 41 25.0 38 2 120 58 1 0
136 1 10.3 46 54.3 45 1 144 86 1 0
137 1 8.5 80 29.4 79 1 134 60 1 1138 1 10.2 63 33.1 60 1 148 80 2 1139 0 10.0 72 27.3 68 1 170 78 3 1140 1 7.3 41 36.9 33 0 160 92 2 1141 0 15.3 52 40.2 36 0 154 96 1 0
142 1 14.0 53 32.7 48 2 124 76 2 1
143 1 15.8 61 33.2 57 1 130 70 1 0144 1 11.4 53 41.4 47 1 156 78 1 0
145 0 5.5 75 35.8 66 0 162 78 1 0
146 1 11.0 40 34.0 38 2 132 76 1 0147 1 7.3 61 19.9 37 0 120 60 2 1
148 0 10.6 62 30.6 49 0 160 86 2 1
149 1 10.5 49 30.8 47 1 146 86 1 0
/p63Status: 0, dead; 1, alive.
/p640, no; ex-smoker; 2, current.
/p651, normal; 2, borderline; 3, abnormal.
/p660, no; 1, yes.
63
CHAPTER4
NonparametricMethodsof
EstimatingSurvivalFunctions
Inthischapterwediscussmethodsofestimatingthethreesurvival (survivor-
ship, density, and hazard )functions for censored data. Unfortunately, the
simple method of Example 2.1 cannot be applied if some of the patients arealive at the time of analysis and therefore their exact survival times areunknown. Nonparametric or distribution-free methods are quite easy tounderstand and apply. They are less efficient than parametric methods whensurvival times followa theoretical distribution and more efficient w hen nosuitable theoretical distributions are known. Therefore, we suggest usingnonparametric methods to analyze survival data before attempting to fit atheoreticaldistribution. If the main objective is to find a model for the data,estimates obtained by nonparametric methods and graphs can be helpful inchoosingadistribution.
Of the three survival functions, survivorship or its graphical presentation,
the survival curve, is the most widely used. Section 4.1 introduces theproduct-limit (PL)method of estimating the survivorshipfunction developed
byKaplanandMeier (1958 ).Withtheincreasedavailabilityofcomputers,this
method is applicable to small, moderate, and large samples. However, if thedatahavealreadybeengroupedintointervals,orthesamplesizeisverylarge,sayinthe thousands,orthe interest isin alarge population,it maybe moreconvenient to perform a life-table analysis. Section 4.2 is devoted to thediscussionofpopulationandclinicallifetables.ThePLestimatesandlife-tableestimatesofthesurvivorshipfunctionareessentiallythesame.Manyauthorsusetheterm life-table estimates forthePLestimates.Theonlydifferenceisthat
thePLestimateisbasedonindividualsurvivaltimes,whereasinthelife-tablemethod, survival times are grouped into intervals. The PL estimate can beconsidered as a special case of the life-table estimate where each intervalcontainsonlyoneobservation.
64
In Section 4.3 we discuss three other measures that describe the survival
experience: the relative survival rate, the five-year survival rate, and thecorrected survival rate. In Section 4.4 we describe two methods, direct andindirectstandardization,toadjustratestoeliminatetheeffectofdifferencesinpopulationcompositionwithrespecttoageandothervariables.Inaddition,itintroducesthestandardizedmortalityrateandstandardizedincidencerate.
4.1 PRODUCT-LIMIT ESTIMATES OF SURVIVORSHIP FUNCTION
Letusfirstconsiderthesimplecasewhereallthepatientsareobservedtodeath
sothat the survivaltimes areexact and known.Let t/p16,t/p17,...,t/p76be the exact
survivaltimesofthe nindividualsunderstudy.Conceptually,weconsiderthis
groupofpatientsasarandomsamplefromamuchlargerpopulationofsimilarpatients.We relabelthe nsurvival times t/p16,t/p17,...,t/p76in ascendingordersuch
thatt/p7/p16/p8/p45t/p7/p17/p8/p45···/p45t/p7/p76/p8.Following (2.1.2 )and (2.1.3 ),thesurvivorshipfunc-
tionatt/p7/p71/p8canbeestimatedas
S/p19(t/p7/p71/p8)/p58n/p57i
n
/p581/p57i
n(4.1.1)
wheren/p57iisthenumberofpeopleinthesamplesurvivinglongerthan t/p7/p71/p8.If
two or more t/p7/p71/p8are equal (tied observations ),the largest ivalue is used. For
example,if t/p7/p17/p8/p58t/p7/p18/p8/p58t/p7/p19/p8,then
S/p19(t/p7/p17/p8)/p58S/p19(t/p7/p18/p8)/p58S/p19(t/p7/p19/p8)/p58n/p574
n
Thisgivesaconservativeestimateforthetiedobservations.
Sinceeverypersonisaliveatthebeginningofthestudyandnoonesurvives
longerthan t/p7/p76/p8,
S/p19(t/p7/p15/p8)/p581 and S/p19(t/p7/p76/p8)/p580( 4 .1.2)
Inpractice, S/p19(t) iscomputedatevery distinctsurvivaltime.Wedonothaveto
worryabouttheintervalsbetweenthedistinctsurvivaltimesinwhichnoonediesand S/p19(t) remainsconstant.Equations (4.1.1 )and (4.1.2 )showthat S/p19(t)i s
a step function starting at 1.0 and decreasing in steps of 1/ n(if there are no
ties)to zero.When S/p19(t) isplottedversus t,the variouspercentilesofsurvival
timecanbereadfromthegraphorcalculatedfrom S/p19(t). Thefollowingexample
illustratesthemethod.
Example 4.1 Consideraclinicaltrialinwhich10lungcancerpatientsare
followedtodeath.Table4.1liststhesurvivaltimes tinmonths.Thefunction- 65
Table 4.1 Computation of S/p19(t) for 10 Lung Cancer
Patients
ti S /p19(t)
41 /p24/p16/p15/p580.9
52 /p23/p16/p15/p580.8
63 /p22/p16/p15/p580.7
84 /p19/p16/p15/p580.4
85 /p19/p16/p15/p580.4
86 /p19/p16/p15/p580.4
10 7 /p17/p16/p15/p580.2
10 8 /p17/p16/p15/p580.2
11 9 /p16/p16/p15/p580.1
12 10 /p15/p16/p15/p580.0
S/p19(t) iscomputedfollowing (4.1.1 )andplottedasastepfunctioninFigure4.1 a
andasasmoothcurveinFigure4.1 b.Theestimatedmediansurvivaltimeis8
months from Figure 4.1 aor 7.6 months from Figure 4.1 b. A more accurate
estimatecanbeobtainedusinglinearinterpolation:
tS /p19(t)
60 .7
m0.5
80 .4
8/p576
0.4/p570.7/p588/p57m
0.4/p570.5
m/p588/p572(0.1)
0.3/p587.3(months )
Theoretically, S/p19(t) should be plotted as a step function since it remains
constant between two observed exact survival times. However, when themediansurvivaltimemustbeestimatedfromasurvivalcurve,asmoothcurve(suchasFigure4.1 b)maygiveamuchbetterestimatethanastepfunction,as
indicatedintheexample.
Thismethodcanbeappliedonlyifallthepatientsarefollowedtodeath.If
someofthepatientsarestill aliveat the endofthestudy,adifferentmethodofestimating S/p19(t), suchasthePLestimategivenbyKaplanandMeier (1958 ),
isrequired.Therationalecanbeillustratedbythefollowingsimpleexample.
Suppose that 10 patients join a clinical study at the beginning of 2000;66
Figure 4.1Function S/p19(t) oflungcancerpatientsinExample4.1.
during that year 6 patients die and 4 survive. At the end of the year, 20
additional patients join the study. In 2001, 3 patients who entered in thebeginningof2000and15patientswhoenteredlaterdie,leavingoneandfivesurvivors, respectively. Suppose that the study terminates at the end of 2001and you want to estimate the proportion of patients in the populationsurvivingfortwoyearsormore,thatis, S(2).
The first group of patients in this example is followed for two years; the
second group is followed for only one year. One possible estimate, thereduced-sample estimate ,i sS/p19(2)/p581/10/p580.1, which ignores the 20 patients
whoarefollowedonlyforoneyear.KaplanandMeierbelievethatthesecondsample,underobservationforonlyoneyear,cancontributetotheestimateofS(2).
Patients who survived two years may be considered as surviving the first
yearandthensurvivingonemoreyear.Thus,theprobabilityofsurvivingfortwo years or more is equal to the probability of surviving the first year andthensurvivingonemoreyear.Thatis,
S(2)/p58P(survivingfirstyearandthensurvivingonemoreyear )
whichcanbewrittenas
S(2)/p58P(survivingtwoyearsgivenpatienthassurvivedfirstyear )
/p59P(survivingfirstyear )( 4.1.3 )
TheKaplan —Meierestimateof S(2)following (4.1.3 )is
S/p19(2)/p58
/p1proportion of patients surviving two years
given they survive for one year /p2
/p59(proportionofpatientssurvivingoneyear )(4.1.4 )- 67
For the data given above, one of the four patients who survived the first
yearsurvivedtwo years,so the first proportionin (4.1.4 )is/p16/p19. Fourof the 10
patients who entered at the beginning of 2000 and 5 of the 20 patients whoenteredattheendof2000survivedoneyear.Therefore,thesecondproportionin(4.1.4 )is(4/p595)/(10/p5920).ThePLestimateof S(2)is
S/p19(2)/p581
4
/p594/p595
10/p5920/p580.25/p590.3/p580.075
Thissimplerulemaybegeneralizedasfollows:Theprobabilityofsurviving
k(/p462) or more years from the beginning of the study is a product of k
observedsurvivalrates:
S/p19(k)/p58p/p16/p59p/p17/p59p/p18/p59···/p59p/p73(4.1.5 )
wherep/p16denotestheproportionofpatientssurvivingatleastoneyear, p/p17the
proportionofpatientssurvivingthesecondyearaftertheyhavesurvivedoneyear,p/p18the proportion of patients surviving the third year after they have
survived two years, and p/p73the proportion of patients surviving the kth year
aftertheyhavesurvived k/p571years.
Therefore, the PL estimate of the probability of surviving any particular
number of years from the beginning of study is the product of the sameestimate up to the preceding year, and the observed survival rate for theparticularyear,thatis,
S/p19(t)/p58S/p19(t/p571)p/p82(4.1.6)
ThePLestimatesaremaximumlikelihoodestimates.
Inpractice,thePLestimatescanbecalculatedbyconstructingatablewith
fivecolumnsfollowingtheoutlinebelow.
1. Column1containsallthesurvivaltimes,bothcensoredanduncensored,
in order from smallest to largest. Affix a plus sign to the censoredobservation.Ifacensoredobservationhasthesame valueasanuncen-soredobservations,thelattershouldappearfirst.
2. Thesecondcolumn,labeled i,consistsofthecorrespondingrankofeach
observationincolumn1.
3. The third column, labeled r, pertains to uncensored observations only.
Letr/p58i.
4. Compute( n/p57r)/(n/p57r/p591), orp/p71,foreveryuncensoredobservation t/p7/p71/p8incolumn4togivetheproportionofpatientssurvivinguptoandthen
throught/p7/p71/p8.68
Table 4.2 Calculation of the PL Estimate of S/p19(t) for Data in Example 4.2
RemissionTime Rank
ti r (n/p57r)/(n/p57r/p591) S/p19(t)
3.01 1 /p24/p16/p15/p24/p16/p15/p580.900
4.0/p59 2— — —
5.7/p59 3— — —
6.5 4 4 /p21/p22/p24/p16/p15/p59/p21/p22/p580.771 /p63
6.5 5 5 /p20/p21/p24/p16/p15/p59/p21/p22/p59/p20/p21/p580.643 /p63
8.4/p59 6— — —
10.0 7 7 /p18/p19/p24/p16/p15/p59/p21/p22/p59/p20/p21/p59/p18/p19/p580.482
10.0/p59 8— — —
12.0 9 9 /p16/p17/p24/p16/p15/p59/p21/p22/p59/p20/p21/p59/p18/p19/p59/p16/p17/p580.241
15.0 10 10 0 0
/p630.643isusedas S/p19(6.5).Itisaconservativeestimate.5. Incolumn5, S/p19(t) istheproductofallvaluesof( n/p57r)/(n/p57r/p591) upto
and including t. If some uncensored observations are ties, the smallest
S/p19(t) shouldbeused.
To summarizethis procedure,let nbe the total number of patientswhose
survival times, censored or not, are available. Relabel the nsurvival times in
orderofincreasingmagnitudesuchthat t/p7/p16/p8/p45t/p7/p17/p8/p45···/p45t/p7/p76/p8.Then
S/p19(t)/p58/p147
t/p7/p80/p8/p45tn/p57r
n/p57r/p591(4.1.7)
whererruns through those positive integers for which t/p7/p80/p8/p45tandt/p7/p80/p8is
uncensored.Thevaluesof rareconsecutiveintegers1,2,..., nifthereareno
censoredobservations;iftherearecensoredobservations,theyarenot.
Theestimatedmediansurvivaltimeisthe50thpercentile,whichisthevalue
oftatS/p19(t)/p580.50.Thefollowingexampleillustratesthecalculationprocedures.
Example 4.2 Supposethatthefollowingremissiondurationsareobserved
from10patients (n/p5810)withsolidtumors.Sixpatientsrelapseat3.0,6.5,6.5,
10, 12, and 15 months; 1 patient is lost to follow-up at 8.4 months; and 3patients are still in remission at the end of the study after 4.0, 5.7, and 10months.Thecalculationof S/p19(t) isshowninTable4.2.
Thesurvivorshipfunction S/p19(t) isplottedinFigure4.2;theestimatedmedian
remissiontimeis m/p589.8months.Fromthecalculationwenoticethat S/p19(t)a t- 69
Figure 4.2Function S/p19(t) ofExample4.2.
t/p58t/p7/p71/p8isrelatedto S/p19(t)a tt/p58t/p7/p71/p92/p16/p8and (4.1.6 )canberewrittenas
S/p19(t/p7/p71/p8)/p58S/p19(t/p7/p71/p92/p16/p8)n/p57i
n/p57i/p591(4.1.8 )
wheret/p7/p71/p8andt/p7/p71/p92/p16/p8areuncensoredobservations.Forexample,
S/p19(12) /p58S/p19(10)/p59/p16/p17/p580.482/p59/p16/p17/p580.241
Iftherearenocensoredobservationsorlossesbefore t,(4.1.7 )isequivalentto
(4.1.1 ).
ThevarianceofthePLestimateof S/p19(t) isapproximatedby
Var[S/p19(t)]/p60[S/p19(t)]/p17/p26
/p801
(n/p57r)(n/p57r/p591)(4.1.9)
whererincludes thosepositiveintegers for which t/p7/p80/p8/p45tandt/p7/p80/p8corresponds
toadeath.ForthedatainExample4.2,forexample,
Var[S/p19(10)] /p58(0.482)/p17/p11
9/p5910/p591
6/p597/p591
5/p596/p591
3/p594/p2
/p580.0352
andtheestimatedstandarderroris0.1876.InExample4.1,
Var[S/p19(6)]/p58(0.7)/p17/p11
9/p5910/p591
8/p599/p591
7/p598/p2/p580.021070
andtheestimatedstandarderroris0.145.Thevariancemaybeusedtoobtain
confidenceintervalsfor S(t).
CalculationofthePLestimateof S(t) inExample4.2canalsobeobtained
byusingstatisticalsoftware.Let tdenotetheobservedremissiontime (uncen-
sored or censored )in Table 4.2 and CENS denote an index (or dummy )
variablewithCENS /p580iftiscensoredand1otherwise.Assumethatthedata
havebeensavedin‘‘C: /p33D4d2.DAT’’asatextfile,whichcontainstwocolumns,
tandCENS,separatedbyaspace.
The following SAS code can be used to obtain the PL estimate of S(t)i n
Table4.2.Onecanadopt thiscode to obtainthePLestimateof S(t) forany
observeduncensoredorcensoredsurvivaltimedata.
dataw1;
infile‘c: /p33d4d2.dat’missover;
inputtcens;
run;proclifetestdata /p58w1outsurv /p58wa;
timet*cens (0);
run;
title’PLestimateofsurvivalfunction’;procprintdata /p58wa;
run;
IfBMDP1Lisused,thefollowingcodecanbeused.
/input file /p58‘c:/p33d4d2.dat’.
variables /p582.
format /p58free.
/variable names /p58t,cens.
/form time /p58t.
status /p58cens.
response /p581.
/estimate method /p58product.
Print.
/end
IftheSPSSKMprocedureisused,thefollowingcodecanbeused.
datalistfile /p58‘c:/p33d4d2.dat’free
/tcens.
kmt
/status /p58censevent (1)
/print.
Example 4.3 Consider the tumor-free time in days of the 30 rats on a
low-fatdietinTable3.4.Table4.3givesthecalculationsofthePLestimatesof
S(t) andthestandarderrorof S/p19(t). Theestimated S(t) isplottedinFigure3.3.
Themediantumor-freetimeisapproximately189days.- 71
Table 4.3 Calculation of S/p19(t) and Standard Error of S/p19(t) for 30 Rat s on a Low-FatDietin Table 3.4
Rat Tumor-free
Number Time ti rn/p57r
n/p57r/p591S/p19(t) StandardErrorof S/p19(t)
35 0 1 1 /p17/p24/p18/p15/p17/p24/p18/p15/p580.967 /p3(0.967 )/p17/p11
29/p5930/p2/p4/p16/p30/p17/p580.033
12 56 2 2 /p17/p23/p17/p240.967/p59/p17/p23/p17/p24/p580.933 /p3(0.933 )/p17/p11
29/p5930/p591
28/p5929/p2/p4/p16/p30/p17/p580.046
46 5 3 3 /p17/p22/p17/p230.933/p59/p17/p22/p17/p23/p580.900 /p3(0.900 )/p17/p11
29/p5930/p591
28/p5929/p591
27/p5928/p2/p4/p16/p30/p17/p580.055
13 66 4 4 /p17/p21/p17/p220.900/p59/p17/p21/p17/p22/p580.867 0.062
14 73 5 5 /p17/p20/p17/p210.833 0.068
97 7 6 6 /p17/p19/p17/p200.800 0.073
10 84 7 7 /p17/p18/p17/p190.767 0.077
58 6 8 8 /p17/p17/p17/p180.733 0.081
11 87 9 9 /p17/p16/p17/p170.700 0.084
15 119 10 10 /p17/p15/p17/p160.667 0.086
1 140 11 11 /p16/p24/p17/p150.633 0.088
16 140 /p5912 — — —
72
6 153 13 13 /p16/p22/p16/p230.633/p59/p16/p22/p16/p23/p580.598 0.090
2 177 14 14 /p16/p21/p16/p220.598/p59/p16/p21/p16/p22/p580.563 0.091
7 181 15 15 /p16/p20/p16/p210.528 0.092
8 191 16 16 /p16/p19/p16/p200.493 0.092
17 200 /p5917 — — — —
18 200 /p5918 — — — —
19 200 /p5919 — — — —
20 200 /p5920 — — — —
21 200 /p5921 — — — —
22 200 /p5922 — — — —
23 200 /p5923 — — — —
24 200 /p5924 — — — —
25 200 /p5925 — — — —
26 200 /p5926 — — — —
27 200 /p5927 — — — —
28 200 /p5928 — — — —
29 200 /p5929 — — — —
30 200 /p5930 — — — —
73
The mean survival time /afii9839can be shown to equal the area under the
estimatedsurvivorshipfunction.Toestimate /afii9839,wecanuse
/afii9839/p24/p58/p16/p27
/p15S/p19(t)dt
thatis, /afii9839/p24isequaltotheareaundertheestimatedsurvivorshipfunction.Thus,
if the times to death are ordered as t/p7/p16/p8/p45t/p7/p17/p8/p45···/p45t/p7/p75/p8(if there are m
uncensored observations )andt/p7/p75/p8is the largest observation of all nobserva-
tions[i.e., t/p7/p75/p8/p58t/p7/p76/p8whent/p7/p76/p8isanuncensoredobservation], /afii9839canbeestimated
as
/afii9839/p24/p581.000t/p7/p16/p8 /p59S/p19(t/p7/p16/p8)(t/p7/p17/p8 /p57t/p7/p16/p8)/p59S/p19(t/p7/p17/p8)(t/p7/p18/p8 /p57t/p7/p17/p8)/p59···
/p59S/p19(t/p7/p75/p92/p16/p8)(t/p7/p75/p8/p57t/p7/p75/p92/p16/p8) (4.1.10 )
whichisthesumoftheareasoftherectanglesunderthesurvivalcurveformed
by the uncensored observations. Consider the data in Example 4.2: m/p586,
t/p7/p16/p8 /p583.0,t/p7/p17/p8 /p586.5,t/p7/p18/p8 /p586.5,t/p7/p19/p8 /p5810,t/p7/p20/p8 /p5812, and t/p7/p21/p8 /p5815.The mean
survivaltimeisestimatedusing (4.1.10 )as
/afii9839/p24/p581.000/p593.0/p590.900 (6.5/p573.0)/p590.643 (10/p576.5)
/p590.482 (12/p5710)/p590.241 (15/p5712)
/p583.000 /p593.150 /p592.251 /p590.964 /p590.723
/p5810.088months
However, if the largest observation in the data is censored and is used as
t/p7/p75/p8in(4.10),/afii9839soobtainedmaybealowestimate.Insuchcases,Irwin (1949 )
suggeststhatinsteadofestimatingthemeansurvivaltime,oneshouldchooseatimelimit Landestimatethe‘‘meansurvivaltimelimitedtoatime L,’’say
/afii9839/p9/p42/p10, byusing Lfort/p7/p75/p8in(4.1.10 ).For example,if inExample4.2 thelargest
observationiscensored,thatis,15 /p59,andifwelet L/p5816,then
/afii9839/p9/p16/p21/p10/p583.000 /p593.150; /p592.251 /p590.964 /p590.241 (16/p5712)
/p5810.329months
whichisthemeansurvivaltimelimitedto16months.
Thevarianceof /afii9839/p24isestimatedby
Var(/afii9839/p24)/p58/p26
/p80A/p17/p80(n/p57r)(n/p57r/p591)(4.1.11)
whererrunsthroughthoseintegersfor which t/p80correspondsto adeath,and74
A/p80is theareaunderthecurve S/p19(t) to theright of t/p7/p80/p8. ThekthA/p80interms of
themuncensoredobservationsis
S/p19(t/p7/p73/p8)(t/p7/p73/p62/p16/p8 /p57k/p7/p73/p8)/p59S/p19(t/p7/p73/p62/p16/p8)(t/p7/p73/p62/p17/p8 /p57t/p7/p73/p62/p16/p8)/p59···/p59S/p19(t/p7/p75/p92/p16/p8)(t/p7/p75/p8/p57t/p7/p75/p92/p16/p8)
(4.1.12)
If thereare nocensoredobservations, (4.1.10 )reducesto thesamplemean
t/p16/p58/afii9814t/p71/n,and (4.1.11 )reducesto
Var(/afii9839/p24)/p58Var(t/p16)/p58/p26(t/p71/p57t/p16)/p17
n/p17(4.1.13)
whichisnotanunbiasedestimate.KaplanandMeiersuggestthat (4.1.11 )and
(4.1.13 )be multiplied by m/(m/p571) andn/(n/p571), respectively, to correct the
bias.
ConsiderthesurvivaltimesinExample4.1:Thesamplemeanis t/p16/p58/afii9839/p24/p588.2
months and the estimated variance of /afii9839/p24,b y (4.1.13 ), is 0.616. If the factor
n/(n/p571)/p5810/9ismultiplied,theestimatedvarianceof /afii9839becomes0.684.
Tocomputethevarianceof /afii9839/p24inExample4.2,wefirstcomputethefive A/p80’s:
A/p16,A/p19,A/p20,A/p22,andA/p24.Thefirst A/p80is
A/p16/p58S/p19(t/p7/p16/p8)(t/p7/p17/p8 /p57t/p7/p16/p8)/p59S/p19(t/p7/p17/p8)(t/p7/p18/p8 /p57t/p7/p17/p8)/p59···/p59S/p19(t/p7/p20/p8)(t/p7/p21/p8 /p57t/p7/p20/p8)
/p583.150/p592.251/p590.964/p590.723/p587.088
Thesecond A/p80is
A/p19/p58S/p19(t/p7/p17/p8)(t/p7/p18/p8 /p57t/p7/p17/p8)/p59···/p59S/p19(t/p7/p20/p8)(t/p7/p21/p8 /p57t/p7/p20/p8)
/p582.251/p590.964/p590.723/p583.938
Thethird,fourth,andfifth A/p80’sare,respectively,
A/p20/p582.251/p590.964/p590.723/p583.938
A/p22/p580.964/p590.723/p581.687
A/p24/p580.723
Thus,
Var/p19(/afii9839/p24)/p58(7.088 )/p17
9/p5910/p59(3.938 )/p17
6/p597/p59(3.928 )/p17
5/p596/p59(1.687 )/p17
3/p594/p59(0.723 )/p17
1/p592/p581.942
The estimated standard error of /afii9839/p24is 1.394. If the factor m/(m/p571)/p586/5i s
included,theseresultsbecome2.330and1.526,respectively.
The Kaplan —Meier method provides very useful estimates of survival
probabilitiesandgraphicalpresentationofsurvivaldistribution.Itisthemost- 75
Figure 4.3Kaplan—Meierestimateofmediansurvivaltime.widelyusedmethodinsurvivaldataanalysis.BreslowandCrowley (1974 )and
Meier (1975b )have shown that under certain conditions, the estimate is
consistent and asymptomatically normal. However, a few critical featuresshouldbementioned.
1. TheKaplan —Meierestimatesarelimitedtothetimeintervalinwhichthe
observations fall. If the largest observation is uncensored, the PLestimate at that time equals zero. Although the estimate may not bewelcomed by physicians, it is correct since no one in the sample liveslonger.Ifthelargestobservationiscensored,thePLestimatecanneverequalzeroandisundefinedbeyondthelargestobservation.
2. The most commonly used summary statistic in survival analysis is the
mediansurvivaltime.Asimpleestimateofthemediancanbereadfromsurvival curves estimated by the PL method as the time tat which
S/p19(t)/p580.5. However, the solution may not be unique. Consider Figure
4.3a,wherethe survivalcurveis horizontalat S/p19(t)/p580.5;anytvalue in
the interval t/p16tot/p17is a reasonableestimate of the median. A practical
solutionistotakethemidpointoftheintervalasthePLestimateofthemedian.Figure4.3 bpresentsadifferentcaseinwhichthestraightforward
estimate (t/p16)tendstooverestimatethemedian.Apracticalwaytohandle
thisproblemistoconnectthepointsandlocatethemedian.
3. If less than 50% of the observations are uncensored and the largest
observationiscensored,themediansurvivaltimecannotbeestimated.Apracticalwayto handlethesituationistouseprobabilitiesofsurvivinga given length of time, say 1, 3, or 5 years, or the mean survival timelimitedtoagiventime t.
4. ThePLmethodassumesthatthecensoringtimesareindependentofthe
survivaltimes.Inotherwords,the reasonan observationis censoredis
unrelatedtothecauseofdeath.Thisassumptionistrueifthepatientis76
still alive at the end of the study period. However, the assumption
is violated if the patient develops severe adverse effects from the treat-ment and is forced to leave the study before death or if the patientdied of a cause other than the one under study (e.g., death due to
automobile accidents in a cancer survival study ). When there is inap-
propriate censoring, the PL method is not appropriate. In practice,one way to alleviate the problem is to avoid it or to reduce it to aminimum.
5. Similar to other estimators, the standard error (S.E.)of the Kaplan —
Meierestimatorof S(t) givesanindicationofthepotentialerrorof S/p19(t).
The confidence interval deserves more attention than just the pointestimate S/p19(t). A95%confidenceintervalfor S(t)i sS/p19(t)/p591.96S.E.[S/p19(t)].
4.2 LIFE-TABLE ANALYSIS
Thelife-table methodisone of theoldesttechniquesfor measuringmortality
and describing the survival experience of a population. It has been used byactuaries, demographers, governmental agencies, and medical researchers instudies of survival, population growth, fertility, migration, length of marriedlife,lengthofworkinglife,andsoon.TherehasbeenadecennialseriesoflifetablesontheentireU.S.populationsince1900.Statesandlocalgovernmentsalsopublishlifetables.Theselifetables,summarizingthemortalityexperienceofaspecificpopulationfora specificperiodoftime,arecalled population life
tables.As clinical and epidemiologic research become more common, the
life-tablemethodhas been appliedto patientswitha given diseasewhohavebeen followed for a period of time. Life tables constructed for patients arecalledclinical life tables. Althoughpopulationandclinicallifetablesaresimilar
incalculation,thesourcesofrequireddataaredifferent.
4.2.1 Population Life Tables
Therearetwokindsofpopulationlifetables:thecohortlifetableandcurrent
life table. The cohort life table describes the survival or mortality experience
frombirthtodeathofaspecificcohortofpersonswhowerebornataboutthesametime,forexample,allpersonsbornin1950.Thecohorthastobefollowedfrom1950untilallofthemdie.Theproportionofdeath (survivor )isthenused
toconstructlifetablesforsuccessivecalendaryears.Thistypeoftable,usefulinpopulationprojectionandprospectivestudies,isnotoftenconstructedsinceitrequiresalongfollow-upperiod.
Thecurrent life table is constructed by applying the age-specific mortality
rates of a population in a given period of time to a hypothetical cohort of100,000or1,000,000persons.Thestartingpointisbirthatyear0.Twosourcesofdataarerequiredforconstructingapopulationlifetable: (1)censusdataon- 77
thenumberoflivingpersons ateach agefor agiven yearat midyearand (2)
vital statistics on the number of deaths in the given year for each age. Forexample, a current U.S. life table assumes a hypothetical cohort of 100,000persons that is subject to the age-specific death rates based on the observeddatafortheUnitedStatesinthe1990census.Thecurrentlifetable,basedonthelifeexperienceofanactualpopulationoverashortperiodoftime,givesagood summary of current mortality. This type of life table is regularlypublished by government agencies of different levels. One of the most often
reported statistics from current life tables is the life expectancy. The termpopulation life table isoftenusedtorefertothecurrentlifetable.
In the United States, the National Center for Health Statistics publishes
detailed decennial life tables after each decennial census. These complete lifetables use one-year age groups. Between censuses, annual life tables are alsopublished.The annuallife tables are often seen in five-year age intervals andarecalled abridged life tables .Tables 4.4and 4.5 are,respectively,acomplete
decennial life table for the total U.S. population for 1989 —1991 and an
abridged life table for the same population for 1998. The abridged table inTable4.5wasconstructedbasedonacompletelifetable.
Currentlifetablesusuallyhavethefollowingcolumns:
1.Age interval[xtox/p59t).Thisisthetimeintervalbetweentwoexactages
xandx/p59t;tis the length of the interval. For example, the interval
20—21 includes the time interval from the 20th birthday up to the 21st
birthday (butnotincludingthe21stbirthday ).
2.Proportionof persons aliveat beginning of age intervalbut dying during the
interval(/p82q/p86). The information is obtained from census data. For
example, (/p82q/p86)for age interval 20 —21 is the proportion of persons who
diedonoraftertheir20thbirthdayandbeforetheir21stbirthday.Itisanestimateoftheconditionalprobabilityofdyingin theintervalgiventhepersonisaliveatage x.Thiscolumnisusuallycalculatedfromdata
ofthedecennialcensusofpopulationanddeathsoccurringinthegiventimeinterval.Forexample,themortalityratesinTable4.4arecalculatedfromthedataofthe1990CensusofPopulationanddeathsoccurringinthe United States in the three years 1989 —1991. This column is the
foundation of the life table from which all of the other columns arederived.
3.Number living at beginning of age interval (l/p86).Theinitialvalueof l/p86,the
sizeofthehypotheticalpopulation,isusually100,000or1,000,000.Thesuccessivevaluesarecomputedusingtheformula
l/p86/p58l/p86/p92/p16(1/p57/p82q/p86/p92/p82)( 4.2.1 )
where1 /p57/p82q/p86/p92/p82istheproportionofpersonswhosurvivedthepreviousage
interval. For example, in Table 4.4, t/p581,l/p17/p15/p58l/p16/p24(1/p57/p16q/p16/p24)/p5878
98,314(1 /p570.00101) /p5898,215, which is the number of persons living at
thebeginningofage20.
4.Number dying during age interval (/p82d/p86)
/p82d/p86/p58l/p86(/p82q/p86)/p58l/p86/p57l/p86/p62/p16(4.2.2)
For example, the number of persons dying during age interval 20 —21,
/p82d/p17/p15/p5898,215(0.00104) /p58102(or/p16d/p17/p15/p5898,215 /p5798,113 /p58102).
5. Stationary population (/p82L/p86andT/p86).Here/p82L/p86isthetotalnumberofyears
livedinthe ithageintervalorthenumberofperson-yearsthat l/p86persons,
agedxexactly, live through the interval. For those who survive the
interval,theircontributionto/p82L/p86isthelengthoftheinterval, t.Forthose
whodieduringtheinterval,wemaynotknowexactlythetimeofdeathandthesurvivaltimemustbeestimated.Theconventionalassumptionisthattheyliveone-halfoftheintervalandcontribute t/2tothecalculation
of/p82L/p86.Thus,
/p82L/p86/p58t(l/p86/p62/p16/p59/p16/p17
/p82d/p86)( 4.2.3 )
Forexample,inTable4.4,/p16L/p17/p15/p5898,113 /p59102/2/p5898,164.Ifwedoknow
the exact survival time of those who die in the interval,/p82L/p86should be
computedaccordingly.
Thesymbol T/p86isthetotalnumberofperson-yearslivedbeyondage t
bypersonsaliveatthatage,thatis,
T/p86/p58/p26
j/p46x/p82L/p72(4.2.4)
and
T/p86/p58/p82L/p86/p59T/p86/p62/p82(4.2.5)
Forexample,inTable4.4, T/p15/p587,536,614,whichisthesumofall/p82L/p86values
incolumn5,and T/p16/p587,437,356,whichis
T/p15/p57/p16L/p15/p587,536,614 /p5799,258.
6.Average remaining lifetime or average number of years of life remaining at
beginning of age interval (e /p4/p71).Thisisalsoknownasthe life expectancy at
agivenage,whichisdefinedasthenumberofyears remainingtobelived
bypersonsatage x:
e /p4/p86/p58T/p86l/p86(4.2.6)
Theexpectedageatdeathofapersonaged xisx/p59e /p4/p86.Thee /p4/p86atx/p580is
the life expectancy at birth. For example, according to the U.S. life- 79
Table 4.4 Life Table for the Total Population, United States, 1989--1991
Proportion
Of100,000 StationaryAverage
Dying
BornAlive PopulationRemainingLifetime
Proportionof Average
Age PersonsAlive Number Number InThis Numberof
Interval atBeginningof Livingat Dying andAll YearsofLife
AgeInterval Beginning During Inthe Subsequent Remainingat
PeriodofLife DyingDuring ofAge Age Age Age BeginningofBetweenTwoAges Interval Interval Interval Interval Intervals AgeInterval
xtox/p59t/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p15
Days
0—1 0.00351 100,000 351 274 7,536,614 75.37
1—7 0.00135 99,649 134 1,637 7,536,340 75.63
7—28 0.00104 99,515 104 5,722 7,534,703 75.71
28—365 0.00349 99,411 347 91,625 7,528,981 75.74
Years
0—1 0.00936 100,000 936 99,258 7,536,614 75.37
1—2 0.00073 99,064 72 99,028 7,437,356 75.08
2—3 0.00048 98,992 48 98,968 7,338,328 74.13
3—4 0.00037 98,944 37 98,926 7,239,360 73.17
4—5 0.00030 98,907 30 98,892 7,140,434 72.19
5—6 0.00027 98,877 27 98,863 7,041,542 71.22
6—7 0.00025 98,850 24 98,839 6,942,679 70.23
7—8 0.00023 98,826 23 98,814 6,843,840 69.25
8—9 0.00020 98,803 20 98,794 6,745,026 68.27
9—10 0.00018 98,783 17 98,774 6,646,232 67.28
80
10—11 0.00016 98,766 16 98,758 6,547,458 66.29
11—12 0.00016 98,750 16 98,742 6,448,700 65.30
12—13 0.00022 98,734 21 98,723 6,349,958 64.31
13—14 0.00032 98,713 32 98,697 6,251,235 63.33
14—15 0.00047 98,681 46 98,658 6,152,538 62.35
15—16 0.00063 98,635 62 98,604 6,053,880 61.38
16—17 0.00077 98,573 76 98,534 5,955,276 60.41
17—18 0.00089 98,497 88 98,453 5,856,742 59.46
18—19 0.00096 98,409 95 98,362 5,758,289 58.51
19—20 0.00101 98,314 99 98,265 5,659,927 57.57
20—21 0.00104 98,215 102 98,164 5,561,662 56.63
21—22 0.00109 98,113 107 98,060 5,463,498 55.69
22—23 0.00112 98,006 110 97,951 5,365,438 54.75
23—24 0.00114 97,896 112 97,840 5,267,487 53.81
24—25 0.00116 97,784 113 97,727 5,169,647 52.87
25—26 0.00117 97,671 115 97,614 5,071,920 51.93
26—27 0.00119 97,556 115 97,499 4,974,306 50.99
27—28 0.00121 97,441 119 97,381 4,876,807 50.05
28—29 0.00126 97,322 123 97,261 4,779,426 49.11
29—30 0.00133 97,199 129 97,135 4,682,165 48.17
30—31 0.00140 97,070 136 97,002 4,585,030 47.23
31—32 0.00147 96,934 143 96,862 4,488,028 46.30
32—33 0.00154 96,791 149 96,717 4,391,166 45.37
33—34 0.00162 96,642 157 96,563 4,294,449 44.44
34—35 0.00170 96,485 163 96,404 4,197,886 43.51
35—36 0.00178 96,322 172 96,236 4,101,482 42.58
36—37 0.00188 96,150 181 96,060 4,005,246 41.66
37—38 0.00198 95,969 189 95,874 3,909,186 40.73
38—39 0.00207 95,780 199 95,681 3,813,312 39.81
39—40 0.00217 95,581 208 95,477 3,717,631 38.90
(Continued overleaf )
81
Table 4.4 Continued
Proportion
Of100,000 StationaryAverage
Dying
BornAlive PopulationRemainingLifetime
Proportionof Average
Age PersonsAlive Number Number InThis Numberof
Interval atBeginningof Livingat Dying andAll YearsofLife
AgeInterval Beginning During Inthe Subsequent Remainingat
PeriodofLife DyingDuring ofAge Age Age Age BeginningofBetweenTwoAges Interval Interval Interval Interval Intervals AgeInterval
xtox/p59t/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p15
40—41 0.00228 95,373 217 95,265 3,622,154 37.98
41—42 0.00240 95,156 228 95,042 3,526,889 37.06
42—43 0.00254 94,928 241 94,808 3,431,847 36.15
43—44 0.00271 94,687 257 94,559 3,337,039 35.24
44—45 0.00292 94,431 277 94,292 3,242,480 34.34
45—46 0.00318 94,154 299 94,005 3,148,188 33.44
46—47 0.00348 93,855 327 93,692 3,054,183 32.54
47—48 0.00380 93,528 355 93,350 2,960,491 31.65
48—49 0.00414 93,173 386 92,980 2,867,141 30.77
49—50 0.00449 92,787 417 92,579 2,774,161 29.90
50—51 0.00490 92,370 452 92,144 2,681,582 29.03
51—52 0.00537 91,918 494 91,671 2,589,438 28.17
52—53 0.00590 91,424 539 91,155 2,497,767 27.32
53—54 0.00647 90,885 588 90,591 2,406,612 26.48
54—55 0.00708 90.297 639 89.978 2,316,021 25.65
55—56 0.00773 89,658 693 89,311 2,226,043 24.83
82
56—57 0.00844 88,965 751 88,589 2,136,732 24.02
57—58 0.00926 88,214 817 87,806 2,048,143 23.22
58—59 0.01019 87,397 891 86,951 1,960,337 22.43
59—60 0.01120 86,506 679 86,021 1,873,386 21.66
60—61 0.01223 85,537 1,047 85,013 1,787,365 20.90
61—62 0.01328 84,490 1,122 83,930 1,702,352 20.15
62—63 0.01439 83,368 1,199 82,768 1,618,422 19.41
63—64 0.01560 82,169 1,282 81,527 1,535,654 18.69
64—65 0.01691 80,887 1,368 80,203 1,454,127 17.98
65—66 0.01827 79,519 1.453 78,793 1,373,924 17.28
66—67 0.01967 78,066 1,535 77,298 1,295,131 16.59
67—68 0.02121 76,531 1,624 75,719 1,217,833 15.91
69—69 0.02297 74,907 1,721 74,047 1,142,114 15.25
69—70 0.02499 73,186 1,829 72,272 1,068,067 14.59
70—71 0.02727 71,357 1,946 70,384 995,795 13.96
71—72 0.02979 69,411 2,067 68,377 925,411 13.33
72—73 0.03251 67,344 2,190 66,249 857,034 12.73
73—74 0.03534 65,154 2,302 64,003 790,785 12.14
74—75 0.03824 62,852 2,403 61,651 726,782 11.56
75—76 0.04126 60,449 2,494 59,201 665,131 11.00
76—77 0.04455 57,955 2,582 56,664 605,930 10.46
77—78 0.04819 55,373 2,669 54,039 549,266 9.92
78—79 0.05239 52,704 2,761 51,323 495,227 9.40
79—80 0.05723 49,943 2.859 48,514 443,904 8.89
80—81 0.06277 47,084 2,955 45,607 395,390 8.40
81—82 0.06885 44,129 3,038 42,609 349,783 7.93
82—83 0.07535 41,091 3,097 39,543 307,174 7.48
83—84 0.08207 37,994 3,118 36,435 267,631 7.04
84—85 0.08907 34,876 3,106 33,324 231,196 6.63
85—86 0.09705 31,770 3,083 30,228 197,872 6.23
(Continued overleaf )
83
Table 4.4 Continued
Proportion
Of100,000 StationaryAverage
Dying
BornAlive PopulationRemainingLifetime
Proportionof Average
Age PersonsAlive Number Number InThis Numberof
Interval atBeginningof Livingat Dying andAll YearsofLife
AgeInterval Beginning During Inthe Subsequent Remainingat
PeriodofLife DyingDuring ofAge Age Age Age BeginningofBetweenTwoAges Interval Interval Interval Interval Intervals AgeInterval
xtox/p59t/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p15
86—87 0.10627 28,687 3,049 27,163 167,644 5.84
87—88 0.11625 25,638 2,980 24,148 140,481 5.48
88—89 0.12688 22,658 2,875 21,220 116,333 5.13
89—90 0.13834 19,783 2,737 18,415 95,113 4.81
90—91 0.15135 17,046 2,580 15,757 76,698 4.50
91—92 0.16591 14,466 2,400 13,266 60,941 4.21
84
92—93 0.18088 12,066 2,182 10,975 47,675 3.95
93—94 0.19552 9,884 1,933 8,918 36,700 3.71
94—95 0.21000 7,951 1,669 7,116 27,782 3.49
95—96 0.22502 6,282 1,414 5,575 20,666 3.29
96—97 0.24126 4,868 1,174 4,281 15,091 3.10
97—98 0.25689 3,694 949 3,219 10,810 2.93
98—99 0.27175 2,745 746 2,372 7,591 2.77
99—100 0.28751 1,999 575 1,711 5,219 2.61
100—101 0.30418 1,424 433 1,208 3,508 2.46
101—102 0.32182 991 319 832 2,300 2.32
102—103 0.34049 672 229 557 1,468 2.19
103—104 0.36024 443 159 364 911 2.05
104—105 0.38113 284 109 229 547 1.93
105—106 0.40324 175 70 140 318 1.81
106—107 0.42663 105 45 83 178 1.70
107—108 0.45137 60 27 46 95 1.59
108—109 0.47755 33 16 25 49 1.49
109—110 0.50525 17 8 13 24 1.39
Source: U.S. Decennial Life Tables for 1989—1991,V o l.1 ,No .1 , U.S. Life Tables , DHHSPublication PHS-98-1150-1, National Center for HealthStatistics,
Washington,DC,1997.
85
Table 4.5 Abridged Life Table for the Total Population, United States, 1998
Stationary
ProportionDying NumberLiving NumberDying Stationary PopulationinThis LifeExpectancy
DuringAge atBeginningof DuringAge Populationinthe andAllSubsequent atBeginning
Interval, AgeInterval, Interval, AgeInterval, AgeIntervals, ofAgeInterval,
Age/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p86
0—1 0.00721 100,000 721 99,370 7,671,400 76.7
1—5 0.00139 99,279 138 396,786 7,572,030 76.3
5—10 0.00089 99,141 88 495,473 7,175,244 72.4
10—15 0.00110 99,053 109 495,057 6,679,771 67.4
15—20 0.00353 98,944 349 493,926 6,184,714 62.5
20—25 0.00476 98,595 469 491,820 5,690,788 57.7
25—30 0.00487 98,126 478 489,450 5,198,968 53.0
30—35 0.00600 97,648 586 486,840 4,709,518 48.2
35—40 0.00819 97,062 795 483,428 4,222,678 43.5
40—45 0.01176 96,267 1,132 478,670 3,739,250 38.8
45—50 0.01728 95,135 1,644 471,811 3,260,580 34.3
50—55 0.02564 93,491 2,397 461,839 2,788,769 29.8
55—60 0.04009 91,094 3,652 446,966 2,326,930 25.5
60—65 0.06302 87,442 5,511 424,280 1,879,964 21.5
65—70 0.09437 81,931 7,732 391,364 1,455,684 17.8
70—75 0.14239 74,199 10,565 345,660 1,064,320 14.3
75—80 0.20604 63,634 13,111 286,484 718,660 11.3
80—85 0.31641 50,523 15,986 213,526 432,176 8.6
85—90 0.46104 34,537 15,923 131,897 218,650 6.3
90—95 0.61502 18,614 11,448 62,020 86,753 4.7
95—100 0.75426 7,166 5,405 20,150 24,733 3.5
100/p59 1.00000 1,761 1,761 4,583 4,583 2.6
Source: U.S. Life Tables ,1998.NationalVitalStatisticsReports,Vol.48,No.18,NationalCenterforHealthStatistics,Washington,DC,2001.
86
tablefor1989 —1991thelifeexpectancyatbirthis75.37yearsandthatat
age40is37.98years.Thismeansthataccordingtothemortalityratesof1989—1991newbornsareexpectedtolive75.37yearsandthoseatage40
are expected to live another 37.98 years. The life expectancy of apopulationisageneralindicationofthecapabilityofprolonginglife.Itisusedtoidentifytrendsandtocomparelongevity.Table4.5showsthataccordingtothemortalityratesof1998,thenewbornsandthoseatage40areexpectedtolive76.7and38.8years,respectively.Theoveralllife
expectancy indicates an improvement in longevity in the United Statesoverthetimeperiod.
Population life tables can be constructed for various subgroups. For
example,therearepublishedlifetablesbygender,race,causeofdeath,aswellasthosewhicheliminatecertaincausesofdeath.
4.2.2 Clinical Life Tables
The actuarial life table method has been applied to clinical data for many
decades. Berkson and Gage (1950 )and Cutler and Ederer (1958 )give a
life-table method for estimating the survivorship function; Gehan (1969 )
providesmethodsforestimatingallthreefunctions (survivorship,density,and
hazard ).
Thelife-tablemethodrequiresafairlylargenumberofobservations,sothat
survival times can be grouped into intervals. Similar to the PL estimate, thelife-tablemethodincorporatesallsurvivalinformationaccumulateduptotheterminationofthe study.Forexample,incomputinga five-yearsurvivalrateof breast cancer patients, one need not restrict oneself only to those patientswhohaveenteredonstudyforfiveormoreyears.Patientswhohaveenteredfor four, three, two, and even one year contribute useful information to theevaluation of five-year survival. In this way, the life-table technique usesincomplete data such as losses to follow-up and persons withdrawn alive aswellascompletedeathdata.
Table 4.6 shows the format of the clinical life table. The columns are
describedbelow.
1.Interval[t/p71/p59t/p71/p62/p16). The first column gives the intervals intowhich the
survival times and times to loss or withdrawal are distributed. Theinterval is from t/p71up to but notincluding t/p71/p62/p16,i/p581,...,s. The last
intervalhasaninfinitelength.Theseintervalsareassumedtobefixed.
2.Midpoint (t/p75/p71).Themidpointofeachinterval,designated t/p75/p71,i/p581,...,
s/p571,isincludedforconvenienceinplottingthehazardandprobability
densityfunctions.Bothfunctionsareplottedas t/p75/p71.
3.Width (b/p71).Thewidth ofeach interval, b/p71/p58t/p71/p62/p16/p57t/p71,i/p581,...,,s/p571,
isneededforcalculationofthehazardanddensityfunctions.Thewidth- 87
Table 4.6 Formatof a Life Table
Number Number Number Number Conditional Conditional Cumulative Probability
Lostto Withdrawn Number Entering Exposed Proportion Proportion Proportion Density Hazard
Interval Midpoint Width Follow-up Alive Dying Interval toRisk Dying Surviving Surviving f(t/p75/p71)h/p19(t/p75/p71)
t/p57t/p17t/p75/p16b/p16l/p16w/p16d/p16n/p30/p16n/p16q/p24/p16p/p24/p16S/p19(t/p16)/p581.00f/p19(t/p75/p16)h/p19(t/p75/p71)
t/p17/p57t/p18t/p75/p17b/p17l/p17w/p17d/p17n/p30/p17n/p17q/p24/p17p/p24/p17S/p19(t/p17) f/p19(t/p75/p17)h/p19(t/p75/p17)
/p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36/p36/p36/p36 /p36
t/p71/p58t/p71/p62/p16t/p75/p71b/p71l/p71w/p71d/p71n/p30/p71n/p16q/p24/p71p/p24/p71s/p19(t/p71) f/p19(t/p75/p71)h/p19(t/p75/p71)
/p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36/p36/p36/p36 /p36
t/p81/p92/p16/p57t/p81t/p75/p11/p81/p92/p16b/p81/p92/p16l/p81/p92/p16w/p81/p92/p16d/p81/p92/p16n/p30/p81/p92/p16n/p81/p92/p16q/p24/p81/p92/p16p/p24/p81/p92/p16S/p19(t/p81/p92/p16)f/p19(t/p75/p11/p81/p92/p16)h/p19(t/p75/p11/p81/p92/p16)
t/p81/p57/p45—— l/p81w/p81d/p81n/p30/p81n/p8110 S/p19(t/p81)——
88
ofthelastinterval, b/p81,istheoreticallyinfinite;noestimateofthehazardor
densityfunctioncanbeobtainedforthisinterval.
4.Number lost to follow-up (l/p71).Thisisthenumberofpeoplewhoarelost
to observation and whose survival status is thus unknown in the ith
interval (i/p581,...,s).
5. Number withdrawn alive (w/p71).Peoplewithdrawnaliveinthe ithinterval
arethoseknowntobealiveattheclosingdateofthestudy.Thesurvival
time recorded for such persons is the length of time from entrance totheclosingdateofthestudy.
6.Number dying (d/p71). This is the number of people who die in the ith
interval.Thesurvivaltimeofthesepeopleisthetimefromentrancetodeath.
7.Number entering the i thinterval (n/p30/p71).Thenumberofpeopleenteringthe
first interval n/p30/p16is the total sample size. Other entries are determined
fromn/p30/p71/p58n/p30/p71/p92/p16/p57l/p71/p92/p16/p57w/p71/p92/p16/p57d/p71/p92/p16. That is, the number of persons
enteringthe ithintervalisequaltothenumberstudiedatthebeginning
of the preceding interval minus those who are lost to follow-up,withdrawnalive,orhavediedintheprecedinginterval.
8.Number exposed to risk (n/p71). This is the number of people who are
exposedtoriskinthe ithintervalandisdefinedas n/p71/p58n/p30/p71/p57/p16/p17
(l/p71/p59w/p71).
It is assumed that the times to loss or withdrawal are approximatelyuniformly distributed in the interval. Therefore, people lost or with-drawn in the interval are exposed to risk of death for one-half theinterval.Iftherearenolossesorwithdrawals, n/p71/p58n/p30/p71.
9.Conditional proportion dying (q/p24/p71). This is defined as q/p71/p58d/p71/n/p71for
i/p581,...,s/p571, andq/p24/p81/p581. It is an estimate of the conditional
probability of death in the ith interval given exposure to the risk of
deathinthe ithinterval.
10.Conditional proportion surviving (q/p24/p71).Thisisgivenby p/p24/p71/p581/p57q/p24/p71,which
is an estimate of the conditional probability of surviving in the ith
interval.
11.Cumulative proportion surviving [S/p19(t/p71)]. This is an estimate of the
survivorshipfunctionattime t/p71;itisoftenreferredtoasthe cumulative
survival rate . Fori/p581,S/p19/p24(t/p16/p71)/p581 and for i/p582,...,s,S/p19(t/p71)/p58
p/p24/p71/p92/p16S/p19(t/p71/p92/p16). It is the usual life-table estimate and is based on the fact
thatsurvivingtothestartofthe ithintervalmeanssurvivingtothestart
ofandthenthroughthe (i/p571)th interval.
12.Estimated probability density function [f/p19(t/p75)]. This is defined as the
probabilityofdying inthe ith intervalper unitwidth.Thus, anatural
estimateatthemidpointoftheintervalis
f/p19(t/p75)/p58S/p19(t/p71)/p57S/p19(t/p71/p92/p16)
b/p71
/p58S/p19(t/p71)q/p24/p71b/p71i/p581,...,s/p571( 4 .2.7)- 89
13. Hazard function [h/p19(t/p75/p71)]. The hazard function for the ith interval,
estimatedatthemidpoint,is
h/p19(t/p75/p71)/p58d/p71b/p71(n/p71/p57/p16/p17d/p71)/p582q/p24/p71b/p71(1/p59p/p24/p71)i/p581,...,s/p571( 4.2.8)
It is the number of deaths per unit time in the interval divided by
the average number of survivors at the midpoint of the interval.That is, h/p19(t/p75/p71) is derived from f/p19(t/p75/p71)/S/p19(t/p75/p71) andS/p19(t/p75/p71)/p58/p16/p17[S/p19(t/p71/p62/p16)
/p59S/p19(t/p71)] since S(t/p71) is defined as the probability of surviving at the
beginning,notthemidpoint,ofthe ithinterval:
h/p19(t/p75/p71)/p58f/p19(t/p75/p71)
S/p19(t/p75/p71)/p58S/p19(t/p71)q/p24/p71/b/p71/p16/p17S/p19(t/p71)(p/p24/p71/p591)(4.2.9)
whichreducesto (4.2.8 ).
Sacher (1956 )derivesanestimateofthehazardfunctionbyassuming
that hazard is constant within an interval but varies among intervals.Hisestimateis
h/p19(t/p75/p71)/p58/p57logp/p24/p71b/p71 (4.2.10 )
InaMonteCarlostudy,GehanandSiddiqui (1973 )showthat (4.2.9)is
lessbiasedthan (4.2.10 ).
Thelarge-sampleapproximatevariancesoftheestimatedsurvivalfunctions,
S/p19(t/p71),f/p19(t/p75/p71), andh/p19(t/p75/p71) intheithintervalare
Var[S/p19(t/p71)]/p60[S/p19(t/p71)]/p17/p71/p92/p16/p26
/p72/p14/p16q/p24/p72n/p72p/p24/p72(4.2.11)
Var[f/p19(t/p75/p71)]/p60[S/p19(t/p71)q/p24/p71]/p17
b/p71 /p1/p71/p92/p16/p26
/p72/p14/p16q/p24/p72n/p72p/p24/p72/p59p/p24/p71n/p72q/p24/p72/p2(4.2.12)
and
Var[h/p19(t/p75/p71)]/p60[h/p19(t/p75/p71)]/p17
n/p71q/p24/p71/p71/p57/p31
2h/p19(t/p75/p71)b/p71/p4/p17/p8(4.2.13)
Equation (4.2.11 )isgivenbyGreenwood (1926 );Gehan (1969 )derived (4.2.12 )
and(4.2.13 ).Thesemaybeusedtoobtainapproximateconfidenceintervalsfor
thevarioussurvivalfunctions.
The graph of S/p19(t/p71) can be used to find an estimate of the median. Or let
(t/p72,t/p72/p62/p16) be the interval such that S/p19(t/p72)/p460.5 andS/p19(t/p72/p62/p16)/p580.5. Then the90
mediansurvivaltime t/p75canbeestimatedbylinearinterpolation:
t/p19/p75/p58t/p72/p59[S/p19(t/p72)/p570.5]b/p72S/p19(t/p72)/p57S/p19(t/p72/p62/p16)/p58t/p72/p59S/p19(t/p72)/p570.5
f/p19(t/p75/p72)(4.2.14)
wheref/p19(t/p75/p72)isdefinedin (4.2.7 ).
Anotherinterestingmeasurethatcanbeobtainedfromthelifetableisthe
median remaining lifetime at timet/p72, denotedby t/p75/p80(i),i/p581,...,s/p571. If att/p71the proportion of individual survival is S/p19(t/p71), the proportion of individual
survivalat t/p75/p80(i)i s/p16/p17S/p19(t/p71). Thatis,one-halfofthepeoplewhoarealiveattime
t/p71areexpectedtobealiveattime t/p75/p80(i).Let (t/p72,t/p72/p62/p16) betheintervalinwhich
/p16/p17S/p19(t/p71)falls;thatis, S/p19(t/p72)/p46/p16/p17S/p19(t/p71)andS/p19(t/p72/p62/p16)/p58/p16/p17S/p19(t/p71).Thenanestimateof t/p75/p80(i)
is
t/p19/p75/p80(i)/p58(t/p72/p57t/p71)/p59b/p72[S/p19(t/p72)/p57/p16/p17S/p19(t/p71)]
S/p19(t/p72)/p57S/p19(t/p72/p62/p16)(4.2.15 )
HereS/p19(t/p72)istheestimatedproportionsurvivingbeyondthelowerlimitofthe
intervalcontainingthemedian.
Thevarianceof t/p75/p80(i) isapproximately
Var[t/p19/p75/p80(i)]/p58[S/p19(t/p71)]/p17
4n/p71[f/p19(t/p75/p72)]/p17(4.2.16 )
Example 4.4 The following survival data for 2418 males with angina
pectoris, originally reported by Parker et al. (1946 ), were also included in
Gehan’s (1969 )paper. Survival time is computed from time of diagnosis in
years.Thelifetableuses16intervalsofoneyear.Table4.7givesestimatesofthe various survival functions, the median remaining lifetime, and theirstandarderrors.Thesurvivorshipfunction, S/p19(t), isplottedat tandthehazard
anddensityfunctions, h/p19(t) andf/p19(t), areplottedatthemidpointoftheinterval
(Figure4.4 ).
The graph of the estimated hazard function shows that the death rate is
highestin the first year after diagnosis.From the end of the first year to thebeginning of the tenth year, the death rate remains relatively constant,fluctuatingbetween0.09and0.12.Thehazardrateisgenerallyhigherafterthetenth year. Hence, the prognosis for a patient who has survived one year isbetter than that for a newly diagnosed patient if factors such as age, gender,andracearenotconsidered.Asimilarinterpretationisreachedbyexaminingthe estimated median remaining lifetimes. Initially, the estimated medianremaininglifetimeis5.33years.Itreachesapeakof6.34yearsatthebeginningof the second year after diagnosis and then decreases. The median survivaltime,eitherreadfromthesurvivalcurveorusing (4.2.14 ),is5.33yearsandthe
five-yearsurvivalrateis0.5193withastandarderrorof0.0103.- 91
Table 4.7 Life-Table Analysis of 2418 Males with Angina Pectoris
Number Number Number Number Conditional Conditional
Yearafter Lostto Withdrawn Number Entering Exposed Proportion Proportion
Diagnosis Midpoint Width Follow-up Alive Dying Interval toRisk Dying Surviving S/p19(t/p71)f/p19(t/p75/p71)h/p19(t/p75/p71)/p40Var[S/p19(t/p71)]/p40Var[f/p19(t/p75/p71)]/p40Var[h/p19(t/p75/p71)]t/p19/p75/p80(i)/p40Var[t/p19/p75/p80(i)]
00 .51 .0 0 0 456 2418 2418 .00.1886 0 .8114 1 .0000 0.1886 0.2082 — 0 .0080 0 .0097 5 .33 0 .17
11 .51 .0 39 0 226 1962 1942 .50.1163 0 .8837 0 .8114 0.0944 0.1235 0 .0080 0 .0060 0 .0082 6 .35 0 .20
22 .51 .0 22 0 152 1697 1686 .00.0902 0 .9098 0 .7170 0.0646 0.0944 0 .0092 0 .0051 0 .0076 6 .34 0 .24
33 .51 .0 23 0 171 1523 1511 .50.1131 0 .8869 0 .6524 0.0738 0.1199 0 .0097 0 .0054 0 .0092 6 .23 0 .24
44 .51 .0 24 0 135 1329 1317 .00.1025 0 .8975 0 .5786 0.0593 0.1080 0 .0101 0 .0049 0 .0093 6 .22 0 .19
55 .51 .0 107 0 125 1170 1116 .50.1120 0 .8880 0 .5193 0.0581 0.1186 0 .0103 0 .0050 0 .0106 5 .91 0 .18
66 .51 .0 133 0 83 938 871 .50.0952 0 .9048 0 .4611 0.0439 0.1000 0 .0104 0 .0047 0 .0110 5 .60 0 .19
77 .51 .0 102 0 74 722 671 .00.1103 0 .8897 0 .4172 0.0460 0.1167 0 .0105 0 .0052 0 .0135 5 .17 0 .27
88 .51 .0 68 0 51 546 512 .00.0996 0 .9904 0 .3712 0.0370 0.1048 0 .0106 0 .0050 0 .0147 4 .94 0 .28
99 .51 .0 64 0 42 427 395 .00.1063 0 .8937 0 .3342 0.0355 0.1123 0 .0107 0 .0053 0 .0173 4 .83 0 .41
10 10 .51 .0 45 0 43 321 298 .50.1441 0 .8559 0 .2987 0.0430 0.1552 0 .0109 0 .0063 0 .0236 4 .69 0 .42
11 11 .51 .0 53 0 34 233 206 .50.1646 0 .8354 0 .2557 0.0421 0.1794 0 .0111 0 .0068 0 .0306 4 .00/p59—
12 12 .51 .0 33 0 18 146 129 .50.1390 0 .8610 0 .2136 0.0297 0.1494 0 .0114 0 .0067 0 .0351 3 .00/p59—
13 13 .51 .02 7 0 99 5 8 1 .50.1104 0 .8896 0 .1839 0.0203 0.1169 0 .0118 0 .0065 0 .0389 2 .00/p59—
14 14 .51 .02 3 0 65 9 4 7 .50.1263 0 .8737 0 .1636 0.0207 0.1348 0 .0123 0 .0080 0 .0549 1 .00/p59—
15 — — 0 0 0 30 30 .01.0000 0 .0000 0 .1429 — — 0 .0133 — — — —
Source:Gehan (1969 ).
92
Figure 4.4Survivalfunctionsofmalepatientswithanginapectoris.- 93
Assume that survival time t(year)from each of 2418 males with angina
pectorisinExample4.4hasthesameformatasthedatafile‘‘C: /p33D4d2.DAT’’
defined in Example 4.2 and is saved in ‘‘C: /p33D4d4.DAT’’. Then the following
SAScodecanbeusedtoproduceaclinicallifetablesuchasTable4.7.
dataw1;
infile‘c: /p33d4d4.dat’missover;
inputtcens;
run;proclifetestdata /p58w1outsurv /p58wamethod /p58lifeintervals /p580to15by1;
timet*cens (0);
run;title‘Lifetableofthesurvivaltimes’;
procprintdata /p58wa;
run;
IfBMDP1Lisused,therespectivecodeis
/input file /p58‘c:/p33d4d4.dat’.
variables /p582.
format /p58free.
/variable names /p58t,cens.
/form unit /p58year.
time /p58t.
status /p58cens.
response /p581.
/estimate method /p58life.
Print.
/end
IftheSPSSSURVIVALprocedureisused,therespectivecodeis
datalistfile /p58‘c:/p33d4d4.dat’free
/tcens.
survivaltables /p58t
/status /p58cens (1)fort
/intervals /p58thru15by1
/print.
4.3 RELATIVE, FIVE-YEAR, AND CORRECTED SURVIVAL RATES
Anotherapproachtolarge-scalesurvivaldataisthecalculationofthe relative
survival rate or annual survival ratio. The relative survival rate evaluates the
survivalexperienceofpatientsintermsofthegeneralpopulation.Greenwood(1926 )first suggested this approach for evaluating the efficacy of cancer
treatment:Iftheaveragesurvivaltimeofthepatientstreatedequalsthatofa94
randomsampleofpersonsofthesameage,gender,occupation,andsoon,the
patientscouldbeconsidered‘‘cured.’’Cutleretal. (1957,1959,1960 a,b,1967 )
adopted Greenwood’s idea of comparing the survival experience of cancerpatients with that of the general population to ascertain (1)the ratio of
observedtoexpectedsurvivalratesand (2)whether,intime,themortalityrate
declinestoa‘‘normal’’level.
The relative survival rate is defined as the ratio of the survival rate
(probabilityofsurvivingoneyear )forapatientunderstudy (observed rate )to
someoneinthegeneralpopulationofthesameage,gender,andrace (expected
rate)overaspecifiedperiodoftime.Toprovideamoreprecisemeasureofthe
relationshipof the observedand expectedsurvivalrates,Cutler et al. suggestcomputingtheratioforeachindividualfollow-upyear.Arelativerateof100%meansthatduring a specific follow-upyear the mortalityratesin the patientand in the general population are equal. A relative rate of less than 100%meansthatthemortalityrateinthepatientsishigherthanthatinthegeneralpopulation.Cutleretal.usethesurvivalratesintheConnecticutandU.S.lifetablesforthegeneralpopulation.
UsingthenotationsinTable4.6,thesurvivalrateobservedattime t/p71isp/p24/p71,
theexpectedsurvivalratecanbecomputedasfollows:Supposethatattime t/p71there are n/p30/p71individuals alive for whom age, gender, race, and time of
observationareknown.Let p*/p71/p72bethesurvivalrateofthe jthindividualfrom
generalpopulationlifetables (withcorrespondingage,gender,andrace ).The
expectedsurvivalrateis
p*/p71/p581
n/p30/p71
/p76/p89/p71/p26
/p72/p14/p16p*/p71/p72(4.3.1 )
Thentherelativesurvivalrateattime t/p71isdefinedby
r/p71/p58p/p24/p71p*/p71(4.3.2)
Example4.5taken from Cutler et al. (1957)illustrates the interpretation of
relative survival rates.
Example 4.5 A total of 9121 breast cancer cases were diagnosed in
Connecticuthospitalsfrom1935to1953.TheConnecticutlifetableforwhitefemales,1939 —1941,isusedincalculationoftheexpectedsurvivalrate.Table
4.8 gives the observed and expected survival rates as well as the relativesurvivalrates.Figure4.5 agraphicallyshowsthesedata:thesurvivalcurvesfor
the breast cancer patients and the general population. The relative survivalratesareplottedinFigure4.5 b.Forthisgroupofpatients,therelativesurvival
rates, although increasing during 13 successive years, are less than 100%throughout the 15 years of follow-up. During each of the 15 years, the,-, 95
Table 4.8 Relative Survival Rates of Breast Cancer
Patients in Connecticut, 1935--1953
SurvivalRates (%)Relative
Yearsafter SurvivalRate
Diagnosis Observed Expected (%)
0—1 82.9 97.2 85
1—2 83.3 97.1 86
2—3 85.9 96.9 89
3—4 86.8 96.7 90
4—5 89.2 96.6 92
5—6 90.0 96.4 93
6—7 89.9 96.4 93
7—8 91.6 96.2 95
8—9 92.0 96.1 96
9—10 92.7 96.1 96
10—11 92.9 95.9 97
11—12 94.0 95.8 98
12—13 94.1 95.3 99
13—14 91.5 95.3 96
14—15 90.6 94.9 95
Source:Cutleretal. (1957 ).
breast cancer patient mortality rate is greater than that of the general
population.
Othermeasuresofdescribingsurvivalexperienceofcancerpatientsarethe
five-year survival rate and the corrected rate. The five-year survival rate is
simply the cumulative proportion surviving at the end of the fifth year. Forexample, the five-year survival rate for the males with angina pectoris inExample 4.4 is 0.5193. The five-year survival rate is no longer a measure oftreatmentsuccessforpatientswithmanytypesofcancersincethesurvivalofcancerpatientshasimprovedconsiderablyinthelastfewdecades.
Berkson (1942 )suggestsusinga corrected survival rate. Thisisthesurvival
rate if the disease under study alone is the cause of death. In most survivalstudies, the proportion of patients surviving is usually determined withoutconsideringthecauseofdeath,whichmightbeunrelatedtothespecificillness.Ifp/p65denotesthesurvivalratewhencanceraloneisthecauseofdeath,Berkson
proposesthat
p/p65/p58p
p/p15
(4.3.3 )
wherepistheobservedtotalsurvivalrateinagroupofcancerpatientsand p/p15is the survival rate for a group of the same age and gender in the general96
Figure 4.5SurvivalratesofbreastcancerpatientsinConnecticut,1935 —1953.
population. Rate p/p65may be computed at any time after the initiation of
follow-up;it provides a measure of the proportionof patientsthat escaped adeathfromcancerupto thatpoint.Ifa five-yearsurvivalrateis0.5anditiscorrectedfornoncancerdeathsandifwefindthatfive-yearsurvivalrateofthegeneralpopulationis0.9,thecorrectedsurvivalrateis0.5/0.9,or0.56.
4.4 STANDARDIZED RATES AND RATIOS
Ratesandratiosareoftenusedindemographyandepidemiologyto describe
the occurrence of a health-related event. For example, the standardized
mortality (or morbidity )ratio (SMR )is frequently used in occupational
epidemiology as a measure of risk, and the standardized death rate iscommonlyusedincomparingmortalityexperiencesofdifferentpopulationsorthesamepopulationatdifferenttimes.
TheconceptoftheSMRisverysimilartothatoftherelativesurvivalrate
described above. It is defined as the ratio of the observed and the expectednumberofdeathandcanbeexpressedas
SMR /p58observed number of deaths in study population
expected number of deaths in study population
/p59100 (4.4.1 )
wherethe expectednumberof deathsisthe sumof the expecteddeathsfrom
the same age, gender, and race groups in the general population. Thestandardized morbidity ratio can similarly be calculated simply by replacingtheword deathsbydisease cases in(4.4.1 ).Ifonlynewcasesareofinterest,we
calltheratiothe standardized incidence ratio (SIR). 97
Table 4.9 Population and Deaths of Sunny City and Happy City by Age
SunnyCity HappyCity
Age-Specific Age-Specific
Rates Rates
Age Population Deaths (per1000 )Population Deaths (per1000 )
/p5825 25,000 25 1.00 55,000 110 2.0
25—44 40,000 50 1.25 20,000 50 2.5
45—64 20,000 200 10.00 21,000 315 15.0
/p4665 15,000 1,200 80.00 4,000 650 162.5
Total 100,000 1,475 100,000 1,125
Thestandardizeddeathrateisonlyoneofthemanyratesusedtodescribe
thehealth status of a populationorto comparethe healthstatus of differentpopulations. If the populations are similar with respect to demographicvariablessuchasage,gender,orrace,the crude rate,orratioofthenumberof
persons to whom the event under study occurred to the total number ofpersonsinthepopulation,cansafelybeusedforcomparison.
Thelevelofthecruderateisaffectedbydemographiccharacteristicsofthe
population for which the rate is computed. If populations have differentdemographiccompositions,acomparison ofthe cruderates may be mislead-ing.Asanexampleconsiderthetwohypotheticalpopulations,SunnyCityandHappy City, in Table 4.9. The crude death rate of Sunny City is 1000 (1475/
100,000 )or14.7 per 1000.Thecrude death rate of HappyCity is1000 (1125/
100,000 ),or11.25per1000,whichislowerthanthatofSunnyCityeventhough
allage-specificratesinHappyCityarehigher.Thisismainlybecausethereisa large proportion of older people in Sunny City. A crude death rate of apopulation may be relatively high merely because the population has a highproportionofolderpeople;itmayberelativelylowbecausethepopulationhasa high proportion of younger people. Thus, one should adjust the rate toeliminate the effects of age, gender, or other differences. The procedure ofadjustmentiscalled standardization andtherateobtainedafterstandardization
iscalledthe standardized rate .
Themostfrequentlyusedmethodsforstandardizationarethedirectmethod
andtheindirectmethod.
Direct Method
Inthis method a standardpopulationis selected. Thedistributionacross thegroups with different values of the demographic characteristic (e.g., different
age groups )must be known. Let r/p16,...,r/p73, wherekis the number of groups,
bethespecificratesofthedifferentgroupsforthepopulationunderstudy.Letp/p16,...,p/p73be the proportions of people in the kgroups for the standard
population.Thedirectstandardizedrateisobtainedbymultiplyingthespecific98
ratesr/p71byp/p71ineachgroup.Theformulaforthedirectstandardizedrateis
R/p3/p9/p18/p4/p2/p20/p58/p73/p26
/p71/p14/p16r/p71p/p71(4.3.2)
As an example, consider the data in Table 4.9. If we choose a standard
population whose distribution is shown in the second column of Table 4.10,
thedirectstandardizeddeathrateforSunnyCityandHappyCityis,respect-ively,9.37and17.84per1000.Thesestandardizedratesaremorereliablethanthecruderatesforcomparisonpurposes.
Indirect Method
Ifthespecificrates r/p71ofthepopulationbeingstudiedareunknown,thedirect
methodcannotbeapplied.Inthiscase,itispossibletostandardizetheratebyanindirectmethodifthefollowingareavailable:
1. Thenumberofpersonstowhomtheeventbeingstudiedoccurred (D)in
thepopulation.Forexample,ifthedeathrateisbeingstandardized, Dis
thenumberofdeaths.
2. The distribution across the various groups for the population being
studied,denotedby n/p16,...,n/p73.
3. The specific rates of the selected standard population, denoted by
s/p16,...,s/p73.
4. Thecruderateofthestandardpopulation,denotedby r.
Theformulaforindirectstandardizationis
R/p9/p14/p3/p9/p18/p4/p2/p20/p58D
/p26/p73/p71/p14/p16n/p71s/p71
r (4.3.3)
Thesummationin (4.3.3 )istheexpectednumberofpersonstowhomtheevent
occurredonthebasisofthespecificratesofthestandardpopulation.Thus,theindirectmethodadjuststhecruderateofthestandardpopulationbytheratiooftheobservedtoexpectednumberofpersonstowhomtheeventoccurredinthepopulationunderstudy.
Table 4.11 represents an example for the death rate in the states of
OklahomaandArizonain1960 (dataarefromGroveandHetzel,1963 ).The
U.S.populationin 1960is used as the standardpopulation.Thecrudedeathrate of Oklahoma (9.7 per thousand )is higher than that of Arizona (7.8 per
thousand ). However, the indirect standardized rates show a reverse relation-
ship (8.6 for Oklahoma and 9.6 for Arizona ). This, again, is because of the
differencesinagedistribution.Thereisahigherproportionofpeoplebelowtheageof25inArizonaandahigherproportionofpeopleabovetheageof54inOklahoma. 99
Table 4.10 Standardized Death Rates by Direct Method for Sunny City and Happy City
SunnyCity HappyCity
Age-Specific Age-Standardized Age-Specific Age-Standardized
Standard Proportion, DeathRates, DeathRates, DeathRates, DeathRates,
Age Population p/p71r/p71p/p71r/p71r/p71p/p71r/p71
/p5825 420,000 0 .42 1 .00 0 .42 2 .00 .84
25—44 280,000 0.28 1.25 0.35 2.5 0.70
45—64 220,000 0.22 10.00 2.20 15.0 3.30
/p4665 80,000 0.08 80.00 6.40 162.5 13.00
Total 1,000,000 9.37 17.84
(R/p3/p9/p18/p4/p2/p20)( R/p3/p9/p18/p4/p2/p20)
100
Table 4.11 Standardized Death Rates by Indirect Method for Oklahoma and Arizona, 1960
Oklahoma Arizona
StandardPopulation
(U.S.Population,1960 ) Expected Expected
Age-SpecificDeathRates, Population, Deaths, Population, Deaths,
Age s/p71n/p71n/p71s/p71n/p71n/p71s/p71
/p5810 .0270 49,103 1,325 .78 34,599 934 .17
1—4 0.0011 193,644 213.01 132,367 145.60
5—14 0.0005 454,972 227.49 285,830 142.92
15—24 0.0011 329,230 362.15 186,789 205.47
25—34 0.0015 279,327 418.99 169,873 254.81
35—44 0.0030 287,994 863.98 173,029 519.09
45—54 0.0076 269,147 2,045.52 136,573 1,037.95
55—64 0.0174 216,036 3,759.03 92,871 1,615.96
65—74 0.0382 157,385 6,012.11 63,634 2,430.82
75—84 0.0875 74,848 6,549.20 22,499 1,968.66
85/p59 0.1986 16,598 3,296.36 4,092 812.67
Total 2,328,284 25,074 1,302,161 10,068
Cruderates 9.5 9.7 7.8
(perthousand )
Observeddeaths 22,584 10,157Expecteddeaths /p63 25,074 10,068
Standardizedrate
/p122,584
25,074 /p29.5/p588.6 /p110,157
10,068 /p29.5/p589.6 (perthousand )
Source:DatafromGroveandHetzel (1963 ).
/p63/afii9814n/p71s/p71.
101
Resultsfor the adjustedratesdependon thestandardpopulationselected.
Hence,thisselectionshouldbedonecarefully.Whendiscussingdeathratebyage,Shryocketal. (1971 )suggestthatapopulationwithsimilaragedistribu-
tion to the various populationsunder study be selected as a standard. If thedeathrateoftwopopulationsisbeingcompared,itisbesttousetheaverageofthetwodistributionsasastandard.
Itshouldberememberedthatspecificratesarestillthemostaccurateand
essential indicators of the variations among populations. No matter which
methodis used, standardizedrates aremeaningful only whencompared withsimilarlycomputedrates.Kitagawa (1964 )alsocriticizesthestandardizedrate
becauseifthespecificratesvaryindifferentwaysbetweenthetwopopulationsbeing compared, standardization will not indicate the differences and some-timeswill evenmaskthe differences.Nevertheless,ifthespecificratesarenotavailable, if a single rate for a population is desired, or if the demographiccomposition of the population being compared is different, the standardizedrateisuseful.
Bibliographical Remarks
KaplanandMeier’s (1958 )PLmethodisthemostcommonlyusedtechnique
forestimatingthesurvivorshipfunctionforsamplesofsmallandmoderatesize.However,withtheaid ofacomputer,it isnotdifficult to usethe methodforlargesamplesizes.
Berkson (1942 ), Berkson and Gage (1950 ), Cutler and Ederer (1958 ), and
Gehan (1969 )have written classic reports on life-table analysis. Peto et al.
(1976 )published an excellent reviewof some statistical methods related to
clinicaltrials.Theterm life-tableanalysis thattheyuseincludesthePLmethod.
Otherreferencesonlifetablesare,forexample,Armitage (1971 ),Shryocketal.
(1971 ), Kuzma (1967 ), Chiang (1968 ), Gross and Clark (1975 ), and Elandt-
JohnsonandJohnson (1980 ).
RelativesurvivalratesandcorrectedsurvivalrateshavebeenusedbyCutler
andco-workersinaseriesofsurvivalstudiesoncancerpatientsinConnecticutinthe1950sand1960s (Cutleretal.,1957,1959,1960 a,b,1967;Edereretal.,
1961 ).DiscussionsofSMR,standardizedrates,andrelatedtopicscanbefound
inmanystandardepidemiologytextbooks:forexample,MausnerandKramer(1985 ),Kahn (1983 ),Kelseyetal. (1986 ),Shryocketal. (1971 ),Chiang (1961 ),
andMantelandStark (1968 ).
EXERCISES
4.1Considerthesurvivaltimeofthe30melanomapatientsinTable3.1.
(a)Compute and plot the PL estimates of the survivorship functions
S/p19(t) ofthetwotreatmentgroupsandcheckyourresultswithTable
3.2andFigure3.1.102
Exercise Table 4.1
Number
Timefrom NumberLost Withdrawn Number Number
Diagnosis toFollow-up, Alive, Dying, Entering,
(yr) l/p71w/p71d/p71n/p30/p71
0—5 18 0 731 949
5—10 16 0 52 200
10—15 8 67 14 132
15—20 0 33 10 43(b)Computethevarianceof S/p19(t) foreveryuncensoredobservation.
(c)Estimatethemediansurvivaltimesofthetwogroups.
4.2Do the same as in Exercise 4.1 for the remission durations of the two
treatmentgroupsinTable3.1.
4.3ComputeandplotthePLestimatesofthetumor-freetimedistributions
for the saturated fat and unsaturated fat diet groups in Table 3.4.CompareyourresultswithFigure3.4.
4.4Consider the remission data of 42 patients with acute leukemia in
Example3.3.
(a)ComputeandplotthePLestimatesof S(t) ateverytimetorelapse
forthe6-MPandplacebogroups.
(b)Compute the variances of S/p19(10) in the 6-MP group and of S/p19(3) in
theplacebogroup.
(c)Estimatethemedianremissiontimesofthetwotreatmentgroups.
4.5 (a)ComputethesurvivaltimeforeachpatientinExerciseTable3.1.
(b)Estimate and plot the overall survivorship function using the PL
method.Whatisthemediansurvivaltime?
(c)Divide the patients into two groups by gender. Compute and plot
thePLestimatesofthesurvivorshipfunctionsforeachgroup.Whatisthemediansurvivaltimeforeach?
4.6ConsidertheskintestresultsinExerciseTable3.1.Foreachofthefive
skintests:
(a)Divide patients into two groups according to whether they had a
positivereaction.Measurementslessthan10 /p5910(5/p595formumps )
areconsiderednegative.
(b)Estimateandplotthesurvivorshipfunctionsofthetwogroups.
(c)Can you tell from the plots if any skin tests might predict survival
time?
4.7Consider the data of patients with cancer of the ovary diagnosed in
Connecticutfrom1935to1944 (Cutleretal.1960b ).ExerciseTable4.1 103
Exercise Table 4.2 Survival Data of Female Patients with Angina Pectoris
YearAfter NumberEntering NumberLostto
Diagnosis Interval Follow-up NumberDying
0—1 555 0 82
1—2 473 8 30
2—3 435 8 27
3—4 400 7 22
4—5 371 7 26
5—6 338 28 25
6—7 285 31 20
7—8 234 32 11
8—9 191 24 14
9—10 153 27 13
10—11 113 22 5
11—12 86 23 5
12—13 58 18 5
13—14 35 9 2
14—15 24 7 3
15/p59 14 11 3
Source:R. L. Parker et al., JAMA,131(2),9 5—100 (1946 ). Copyright 1946. American Medical
Association.reproducesthe data in life-table format.Provide a life-table like Table
4.5.Whatdoyoufindout?
4.8Doacompletelife-tableanalysisforthetwosetsofdatagiveninTable
3.5.Plotthethreesurvivalfunctions.
4.9Doacompletelife-tableanalysisofthedatagiveninExerciseTable4.2.
Plotthethreesurvivalfunctions.
4.10ConsiderthesurvivaltimesofthemelanomapatientsinExerciseTable
3.4.Doacompletelife-tableanalysisofthesurvivaltime.Plotthethreesurvivalfunctions.
4.11Consider the data given in Exercise Table 4.3. Compute the direct
standardizeddeathrateforthestatesofOklahomaandMontanausingtheU.S.populationof1960asthestandard.
4.12GiventhepopulationofJapanandChile (ExerciseTable4.4 ),compute
theindirectstandardizeddeathrateforthetwocountriesusingtheU.S.deathrateof1960inTable4.11asthestandard.104
Exercise Table 4.3
OklahomaAverage MontanaAverage
DeathRate DeathRate
U.S.Population, Proportion, (per1000 )(per1000 )
Age 1960 (thousands ) p/p71r/p71r/p71
/p581 4,112 0 .023 25 .52 5 .8
1—4 16,209 0.091 1.2 1.2
5—14 35,465 0.198 0.5 0.5
15—24 24,020 0.134 1.2 1.6
25—34 22,818 0.127 1.6 1.8
35—44 24,081 0.134 2.9 3.1
45—54 20,486 0.114 6.9 7.5
55—64 15,572 0.087 14.8 16.3
65—74 10,997 0.061 32.4 37.3
75—84 4,634 0.026 79.0 87.3
85/p59 929 0.005 190.4 202.8
Total 179,323 1.000
Source:GroveandHetzel (1963 ).
Exercise Table 4.4
Population
(thousands )
Age Japan Chile
/p581 1,577 228
1—4 6,268 876
5—14 20,223 1,817
15—24 17,627 1,323
25—34 15,727 1,034
35—44 11,057 779
45—54 9,018 603
55—64 6,573 395
65—74 3,724 212
75—84 1,438 83
/p4585 188 22———————Total 93,419 7,374
Observeddeaths 706,599 95,486
Source:Shryocketal. (1971 ). 105
CHAPTER 5
Nonparametric Methods for
Comparing Survival Distributions
The problem of comparing survival distributions arises often in biomedicalresearch.A laboratory researcher may want to compare the tumor-free timesof two or more groups of rats exposed to carcinogens.A diabetologist maywish to compare the retinopathy-free times of two groups of diabetic patients.A clinical oncologist may be interested in comparing the ability of two or moretreatments to prolong life or maintain health.Almost invariably, the disease-free or survival times of the different groups vary.These differences can beillustrated by drawing graphs of the estimated survivorship functions, but thatgives only a rough idea of the difference between the distributions.It does notreveal whether the differences are significant or merely chance variations.Astatistical test is necessary.
In Section 5.1 we introduce five nonparametric tests that can be used for
data with and without censored observations.Section 5. 2 is devoted to theMantel—Haenszel test, which is particularly useful in stratified analysis, a
method commonly used to take account of possible confounding variables.InSection 5.3 we discuss the problem of comparing three or more survivaldistributions with or without censoring.
5.1 COMPARISON OF TWO SURVIVAL DISTRIBUTIONS
Suppose that there are n/p16andn/p17patients who receive treatments 1 and 2,
respectively.Let x/p16,...,x/p80/p129be ther/p16failure observations and x/p62/p80/p129/p62/p16,...,x/p62/p76/p129the
n/p16/p57r/p16censored observations in group 1.In group 2, let y/p16,...,y/p80/p130be ther/p17failure observations and y/p62/p80/p130/p62/p16,...,y/p62/p76/p130then/p17/p57r/p17censored observations.That
is, at the end of the study n/p16/p57r/p16patients who received treatment 1 and
n/p17/p57r/p17patients who received treatment 2 are still alive.Suppose that the
observations in group 1 are samples from a distribution with survivorshipfunction S/p16(t) and the observations in group 2 are samples from a distribution
106
with survivorship function S/p17(t).Then null hypothesis to consider is
H/p15:S/p16(t)/p58S/p17(t) (treatments 1 and 2 are equally effective )
against the alternative
H/p16:S/p16(t)/p57S/p17(t) (treatment 1 more effective than 2 )
or
H/p17:S/p16(t)/p58S/p17(t) (treatment 2 more effective than 1 )
or
H/p18:S/p16(t)/p34S/p17(t) (treatments 1 and 2 not equally effective )
When there are no censored observations, standard nonparametric tests can
be used to compare two survival distributions.For example, the Wilcoxon(1945 )test or the Mann —Whitney (1947 )U-test can test the equality of two
independent populations, and the sign test can be used for paired (or depend-
ent)samples (Marascuilo and McSweeney, 1977 ).In the following we introduce
five nonparametric tests: Gehan’s generalized Wilcoxon test (Gehan, 1965 a,b),
the Cox—Mantel test (Cox 1959, 1972; Mantel, 1966 ), the logrank test (Peto
and Peto, 1972 ), Peto and Peto’s generalized Wilcoxon test (1972 ), and Cox’s
F-test (1964 ).All the tests are designed to handle censored data; data without
censored observations can be considered a special case.
5.1.1 Gehan’s Generalized Wilcoxon Test
In Gehan’s generalized Wilcoxon test every observation x/p71orx/p62/p71in group 1 is
compared with every observation y/p72ory/p62/p72in group 2 and a score U/p71/p72is given
to the result of every comparison.For the purpose of illustration, let us assumethat the alternative hypothesis is H/p16:S/p16(t)/p57S/p17(t), that is, treatment 1 is more
effective than treatment 2.
Define
U/p71/p72/p58
/p7/p591i fx/p71/p57y/p72orx/p62/p71/p46y/p72
0i fx/p71/p58y/p72orx/p62/p71/p58y/p72ory/p62/p72/p58x/p71or (x/p62/p71,y/p62/p72)
/p571i fx/p71/p58y/p72orx/p71/p45y/p62/p72
and calculate the test statistic
W/p58/p76/p129/p26
/p71/p14/p16/p76/p130/p26
/p72/p14/p16U/p71/p72(5.1.1 )
where the sum is over all n/p16n/p16comparisons.Hence, there is a contribution to 107
the test statistic Wfor every comparison where both observations are failures
(except for ties )and for every comparison where a censored observation is
equal to or larger than a failure.The calculation of Wis laborious when n/p16andn/p17are large.Mantel (1967 )shows that it can be calculated in an alternative
way by assigning a score to each observation based on its relative ranking.InGehan’s computation each observation in sample 1 is compared with each insample 2.If the two samples are combined into a single pooled sample ofn/p16/p59n/p17observations, it is the same as comparing each observation with the
remaining n/p16/p59n/p17/p571.LetU/p71,i/p581,...,n/p16/p59n/p17, be the number of remaining
n/p16/p59n/p17/p571 observations that the ith is definitely greater than minus the
number that it is definitely less than.The n/p16/p59n/p17U/p71’s define a finite population
with mean 0 and it is true that Gehan’s
W/p58/p76/p129/p26
/p71/p14/p16U/p71(5.1.2 )
where summation is over the U/p71of sample 1 only.From either (5.1.1 )or(5.1.2 ),
it is clear that Wwould be a large positive number if H/p16is true.Mantel also
suggests that the permutational variance of Wbe used instead of the more
complicated variance formula derived by Gehan.The permutational distribu-tion ofWcan be obtained by considering all /p16
/p1n/p16/p59n/p17n/p17/p2/p58(n/p16/p59n/p17)!
n/p16!n/p17!
ways of selecting n/p16of theU/p71at random.The test statistic WunderH/p15can be
considered approximately normally distributed with mean 0 and variance /p17
Var(W)/p58n/p16n/p17/p76/p129/p62/p76/p130/p26
/p71/p14/p16U/p17/p71
(n/p16/p59n/p17)(n/p16/p59n/p17/p571)(5.1.3 )
SinceWis discrete, an appropriate continuity correction of 1 is ordinarily used
when there are neither ties nor censored observations.Otherwise, a continuitycorrection of 0.5 would probably be appropriate.
SinceWhas an asymptotically normal distribution with mean zero and
variance in (5.1.3 ),Z/p58W//p40Var(W)
has standard normal distribution.The
rejection regions are Z/p57Z/p63forH/p16, andZ/p58/p57Z/p63forH/p17, and /p34Z/p34/p57Z/p63/p30/p17for
H/p18whereP(Z/p57Z/p63/p34H/p15)/p58/afii9825.
/p16n! is read n factorial: n !/p58n(n/p571)(n/p572)/p373.2.1.
/p17This is called the permutational variance because it is obtained by considering the per mutational
distribution of all ( n/p16/p59n/p17)!/n/p16!n/p17!W’s108
The number U/p71can be computed in two stages.For each observation, the
first stage yields, unity plus the number of remaining observations that it isdefinitely larger than, that is, R/p16/p71.The second stage yields R/p17/p71, which is unity
plus the number of remaining observations that the particular observation isdefinitely less than.Then U/p71/p58R/p16/p71/p57R/p17/p71.The computations of R/p16/p71andR/p17/p71can
be accomplished systematically in steps, as illustrated in the following hypo-thetical example.
Example 5.1 Ten female patients with breast cancer are randomized to
receive either CMF (cyclic administration of cyclophosphamide, methatrexate,
and fluorouracil )or no treatment after a radical mastectomy.At the end of two
years, the following times to relapse (or remission times )in months are
recorded:
CMF (group 1 ): 23, 16 /p59,1 8/p59,2 0/p59,2 4/p59
Control (group 2 ): 15, 18, 19, 19, 20
The null hypothesis and the alternatives are
H/p15:S/p16/p58S/p17(the two treatments are equally effective )
H/p16:S/p16/p57S/p17(CMF more efficient than no treatment )
The computations of R/p16/p71,R/p17/p71, andU/p71are given in Table 5.1. Thus,
W/p581/p592/p595/p594/p596/p5818, Var (W)/p58(5)(5)(208) /[(10)(9)] /p5857.78, and
Z/p5818//p4057.78
/p582.368.Suppose that the significance level used is /afii9825/p580.05,
Z/p15/p13/p15/p20/p581.64; then the Zvalue computed is in the rejection region.Therefore,
we reject H/p15at 0.05 level and conclude that the data show that CMF is more
effective than no treatment.In fact, the approximate pvalue corresponding to
Z/p582.368 is 0.009.
Note that the sum of all n/p16/p59n/p17U/p71’s equals zero.This fact can be used to
check the computation.
5.1.2 Cox--Mantel Test
Lett/p7/p16/p8/p58···/p58t/p7/p73/p8be the distinct failure times in the two groups together and
m/p7/p71/p8be the number of failure times equal to t/p71, or the multiplicity of t/p71, so that
/p73/p26
/p71/p14/p16m/p7/p71/p8/p58r/p16/p59r/p17(5.1.4)
Further, let R(t) be the set of people still exposed to risk of failure at time
t, whose failure or censoring times are at least t.HereR(t) is called the riskset
at timet.Letn/p16/p82andn/p17/p82be the number of patients in R(t) that belong to 109
Table 5.1 Mantel’s Procedure of Calculating Uifor Gehan’s Generalized Wilcoxon Test
Observations of Two
Samples in Ascending
Order 15 16 /p62 18 18 /p62 19 19 20 20 /p62 23 24 /p62
Computation of R/p16/p71
Step 1.Rank from left to
right, omitting
censoredobservations 1 2 3 4 5 6
Step 2.Assign next-higher
rank to censoredobservations 2 3 6 7
Step 3.Reduce the rank
of tied observationsto the lower rankfor the value 3
Step 4.R/p16/p7112 2 3 3 3 5 6 6 7
Computation ofR/p17/p71
Step 5.Rank from right
to left 10 9 8 7 6 5 4 3 2 1
Step 6.Reduce the rank of
tied observations to
the lowest rank for
the value 5
Step 7.Reduce the rank
of censoredobservations to 1 1 1 1 1
Step 8.R/p17/p711 0 18 15 5 4 1 2 1
U/p71/p58R/p16/p71/p57R/p17/p71/p5791 /p63 /p5762 /p63 /p572/p5721 5 /p63 4/p63 6/p63
/p63From group 1.
treatment groups 1 and 2, respectively.The total number of observations,
failure or censored in R(t/p7/p71/p8), isr/p7/p71/p8/p58n/p16/p82/p59n/p17/p82.Define
U/p58r/p17/p57/p73/p26
/p71/p14/p16m/p7/p71/p8A/p7/p71/p8(5.1.5)
I/p58/p73/p26
/p71/p14/p16m/p7/p71/p8(r/p7/p71/p8/p57m/p7/p71/p8)
r/p7/p71/p8/p571A/p7/p71/p8(1/p57A/p7/p71/p8) (5.1.6 )
wherer/p7/p71/p8is the number of observations, failure or censored, in R(t/p7/p71/p8) andA/p7/p71/p8110
Table 5.2 Computations of Cox--Mantel Test
Number in Risk Set of:
Distinct Sample 1 Sample 2
Failure Time, t/p71m/p7/p71/p8n/p16/p82n/p17/p82r/p7/p71/p8A/p7/p71/p8
15 1 5 5 10 0 .5
18 1 4 4 8 0 .5
19 2 3 3 6 0 .5
20 1 3 1 4 0 .25
23 1 2 0 2 0is the proportion of r/p7/p71/p8that belong to group 2.An asymptotic two-sample test
is thus obtained by treating the statistic C/p58U//p40Ias a standard normal
variate under the null hypothesis (Cox, 1972 ).The following example illustrates
the procedure.
Example 5.2 Consider the remission data and the hypotheses in Example
5.1. There are k/p585 distinct failure times in the two groups, r/p16/p581 andr/p17/p585.
To perform the Cox —Mantel test, Table 5.2 is prepared for convenience:
U/p585/p57(0.5/p590.5/p592/p590.5/p590.25)
/p585/p572.25
/p582.75
I/p581/p599
9(0.5/p590.5)/p591/p597
7(0.5/p590.5)/p592/p594
5(0.5/p590.5)/p591/p593
3(0.25/p590.75)
/p580.25 /p590.25 /p590.4/p590.1875
/p581.0875
Therefore, C/p582.75//p401.0875 /p582.637/p57Z/p15/p13/p15/p20/p581.64 and we reject H/p15at 0.05
level and reach the same conclusion as in Example 5.1. The pvalue correspond-
ing toZ/p582.637 is approximately 0.004.
5.1.3 Logrank Test
Mantel’s (1966 )generalization of the Savage (1956 )test, often referred to as the
logranktest (Peto and Peto, 1972 ), is based on a set of scores w/p71assigned to
the observations.The scores are functions of the logarithm of the survival 111
function.Altshuler (1970 )estimates the log survival function at t/p7/p71/p8using
/p57e(t/p7/p71/p8)/p58/p57 /p26
j/p45t/p7/p71/p8m/p7/p72/p8r/p7/p72/p8(5.1.7)
wherem/p7/p72/p8andr/p7/p72/p8are as defined in Section 5.1.2. The scores suggested by Peto
and Peto are w/p71/p581/p57e(t/p7/p71/p8) for an uncensored observation t/p7/p71/p8and /p57e(T) for
an observation censored at T.In practice, for a censored observation t/p62/p71,
w/p71/p58/p57e(t/p7/p72/p8), where t/p7/p72/p8is the largest uncensored observation that t/p7/p72/p8/p45t/p62/p71.
Thus, the larger the uncensored observation, the smaller its score.Censoredobservations receive negative scores.The wscores sum identically to zero for
the two groups together.The logrank test is based on the sum Sof thewscores
of the two groups.The permutational variance of Sis given by
Var(S)/p58n/p16n/p17/p26/p76/p129/p62/p76/p130/p71/p14/p16w/p17/p71(n/p16/p59n/p17)(n/p16/p59n/p17/p571)(5.1.8)
which can be rewritten as
V/p58/p3/p73/p26
/p72/p14/p16m/p7/p72/p8(r/p7/p72/p8/p57m/p7/p72/p8)
r/p7/p72/p8 /p4n/p16n/p17(n/p16/p59n/p17)(n/p16/p59n/p17/p571)(5.1.9)
The test statistic L/p58S//p40Var(S)has an asymptotically standard normal
distribution under the null hypothesis.If Sis obtained from group 1, the critical
region is L/p58/p57Z/p63, and ifSis obtained from group 2, the critical region is
L/p57Z/p63, where /afii9825is the significance level for testing H/p15:S/p16/p58S/p17against
H/p16:S/p16/p57S/p17.The following example illustrates the computational procedures.
Example 5.3 Consider the data and hypotheses in Example 5.1. The test
statistic of the logrank test can be computed by tabulating m/p7/p71/p8,r/p7/p71/p8,m/p7/p71/p8/r/p7/p71/p8,
ande(t/p7/p71/p8) as in Table 5.3. Since every observation in the two samples, censored
or not, is assigned a score, it is convenient to list them in column 1.Columns2 to 5 pertain only to the failure times; e(t/p7/p71/p8) is the cumulative value of m/p7/p71/p8/r/p7/p71/p8,
Altshuler’s (1970 )estimate of the logarithm of the survivorship function
multipled by /p571.For example, at t/p7/p71/p8/p5818,e(t/p7/p71/p8)/p580.100 /p590.125 /p580.225; at
t/p7/p71/p8/p5819,e(t/p7/p71/p8)/p580.225/p590.333/p580.558.The last column, w/p71, gives the score for
every observation.For an uncensored observation w/p71/p581/p57e(t/p7/p71/p8), for example,
att/p71/p5818,w/p71/p581/p570.225/p580.775.Since e(t/p7/p71/p8) is an estimate of a function of
the survivorship function, which we assume to be constant between twoconsecutive failures, e(t/p62/p71) is equal to e(t/p7/p72/p8) fort/p7/p72/p8/p45t/p62/p71.Thusw/p71for censored
observations t/p62/p71equals /p57e(t/p7/p72/p8), where t/p7/p72/p8/p45t/p62/p71.For example, w/p71for 16 /p62is
/p57e(15), or /p570.100, and that for 18 /p62is/p57e(18), or /p570.225. Tied observations
like the two 19’s receive the same score: 0.442. The 10 scores w/p71sum to zero,
which can be used to check the computation.112
Table 5.3 Computations of Logrank Test
Remission Times
in Both Samples,
t/p71m/p7/p71/p8r/p7/p71/p8m/p7/p71/p8/r/p7/p71/p8e(t/p7/p71/p8)w/p71
15 1 10 0 .100 0 .100 0 .900 /p63
16/p59 —— — — /p570.100
18 1 8 0 .125 0 .225 0 .775 /p63
18/p59 —— — — /p570.225
19 2 6 0 .333 0 .558 0 .442 /p63
20 1 4 0 .250 0 .808 0 .192 /p63
20/p59 —— — — /p570.808
23 1 2 0 .500 1 .308 /p570.308
24/p59 —— — — /p571.308
/p63From sample 2.
The statistic S/p580.900/p590.775/p590.442/p590.442/p590.192/p582.751.The vari-
ance of S, computed by (5.8)is 1.210. Hence, the test statistic L/p582.751/
/p401.210/p582.5 and the pvalue is approximately 0.0064, data showing that CMF
treatment is superior.The logrank statistic Scan be shown to equal the sum
of the failures observed minus the conditional failures expected computed ateach failure time, or simply the difference between the observed and expectedfailures in one of the groups.A similar version of the logrank test is achi-square test which compares the observed number of failures to the expectednumber of failures under the hypothesis.Let O/p16andO/p17be the observed
numbers and E/p16andE/p17the expected numbers of death in the two treatment
groups.The test statistic
X/p17/p58(O/p16/p57E/p16)/p17
E/p16
/p59(O/p17/p57E/p17)/p17
E/p17(5.1.10)
has approximately the chi-square distribution with 1 degree of freedom.A large
X/p17value (e.g.,/p46X/p17/p16/p11/p13/p15/p20)would lead to the rejection of the null hypothesis in
favor of the alternative that the two treatments are not equally effective(/afii9825/p580.05).
To compute E/p16andE/p17, we arrange all the uncensored observations in
ascending order and compute the deaths expected at each uncensored time andsum them.The number of deaths expected at an uncensored time is obtainedby multiplying the deaths observed at that time by the proportion of patientsexposed to risk in the treatment group.Let d/p16be the number of deaths at time
tandn/p16/p82andn/p17/p82be the numbers of patients still exposed to risk of dying at
time up to tin the two treatment groups.The deaths expected for groups 1 113
Table 5.4 Computation of E1of Logrank Test
Relapse time, td/p82n/p16/p82n/p17/p82e/p16/p82e/p17/p82
15 1 5 5 0.5 0.5
18 1 4 4 0.5 0.519 2 3 3 1.0 1.020 1 3 1 0.75 0.2523 1 2 0 1.0 0
Total 3.75 2.25and 2 at time tare
e/p16/p82/p58n/p16/p82n/p16/p82/p59n/p17/p82/p59d/p82e/p17/p82/p58n/p17/p82n/p16/p82/p59n/p17/p82/p59d/p82(5.1.11 )
Then the total numbers of deaths expected in the two groups are
E/p16/p58/p26
/p0/p12/p12/p82e/p16/p82E/p17/p58/p26
/p0/p12/p12/p82e/p17/p82
In practice, we only need to compute the total number of deaths expected
in one of the two groups, for example, E/p16, sinceE/p17is the total observed number
of deaths minus E/p16.The following example illustrates the calculation pro-
cedure.
Example 5.4 Let us use the hypothetical data in Example 5.1 again. The
remission times in months are:
CMF (group 1 ): 23, 16 /p59,1 8/p59,2 0/p59,2 4/p59
Control (group 2 ): 15, 18, 19, 19, 20.
Consider the following null and alternative hypotheses:
H/p15:S/p16/p58S/p17(the two treatments are equally effective )
H/p16:S/p16/p34S/p17(the two treatments are not equally effective )
Table 5.4 gives the calculation of E/p16.For example, at t/p5818, four patients
in group 1 and four in group 2 are still exposed to the risk of relapse, and thereis one relapse.Thus, d/p82/p581,n/p16/p82/p58n/p17/p82/p584, ande/p16/p82/p580.5.
The total number of relapses expected is E/p16/p583.75.Since there are a total of
six deaths ( O/p16/p581,O/p17/p585)in the two groups, E/p17/p586/p573.75/p582.25.Using114
(5.1.10 ), we have
X/p17/p58(1/p573.75)/p17
3.75/p59(5/p572.25)/p17
2.25/p585.378
Using Table C-2, the pvalue corresponding to this X/p17value is less 0.05
(p/p600.02).Therefore, we reach the same conclusion: that there is a significant
difference in remission duration between the CMF and control groups.
Computer software is available to perform a number of two-sample tests
with censored observations.For example, SAS, SPSS, and BMDP provideprocedures for the logrank and Cox —Mantel tests.We use the remission time
of the 10 breast cancer patients in Example 5.1 to illustrate the use of thesesoftware packages.To compare the two groups, we create the following threevariables: t, remission time; CENS /p580i ftis censored and 1 otherwise; and
TREAT /p581 if receiving CMF and /p582 if no treatment.Assume that the data
have been saved in ‘‘C: /p33D5d1.DAT’’ as a text file, which contains three
columns, separated by a space (tis in the first column, CENS the second
column, and TREAT the third column ), and the data in each row are for the
same patient.The following SAS code can be used to perform the logrank test.
data w1;
infile ‘c: /p33d5d1.dat’ missover;
input t cens treat;
run;
proc lifetest data /p58w1;
time t*cens (0);
strata treat;
run;
If BMDP procedure 1L is used, the following code can be used to perform
the Cox—Mantel test.
/input file /p58‘c:/p33d5d1.dat’ .
variables /p583.
format /p58free.
/variable names /p58t, cens, treat.
/form time /p58t.
status /p58cens.
response /p581.
/group codes (treat )/p581, 2.
Names (treat )/p58treated, control.
/estimate method /p58product.
Group /p58treat.
Stat /p58mantel.
/end 115
If procedure KM in SPSS is used, the following code can be used to perform
the Cox—Mantel test.
data list file /p58‘c:/p33d5d1.dat’ free
/ t cens treat.
km t by treat
/status /p58cens event (1)
/test /p58logrank.
These codes can be modified to perform tests comparing more than two groups
simply by replacing TREAT in the codes with the group variable defined.
5.1.4 Peto and Peto’s Generalized Wilcoxon Test
Another generalization of Wilcoxon’s two-sample rank sum test is described by
Peto and Peto (1972 ).Similar to the logrank test, this test assigns a score to
every observation.For an uncensored observation t, the score is u/p71/p58
S/p19(t/p59)/p59S/p19(t/p57)/p571, and for an observation censored at T, the score is
u/p71/p58S/p19(T)/p571, where S/p19is the Kaplan —Meier estimate of the survival function.
If we use the notation of Section 5.1.2, the score for an uncensored observationt/p7/p71/p8isu/p71/p58S/p19(t/p7/p71/p8)/p59S/p19(t/p7/p71/p92/p16/p8)/p571 andS/p19(t/p7/p15/p8)/p580 and that for a censored observa-
tion ist/p62/p72isu/p72/p58S/p19(t/p7/p71/p8)/p571, where t/p7/p71/p8/p45t/p62/p72.These generalized Wilcoxon scores
sum to zero.The test procedure after the scores are assigned is the same as forthe logrank test.The following example illustrates the computational pro-cedures.
Example 5.5 Using the same data and hypotheses as in Example 5.1, the
calculations of the scores u/p71for Peto and Peto’s generalized Wilcoxon test are
given in Table 5.5. Using the scores of group 1, we obtain
S/p58/p57 0.100/p570.212/p570.605/p570.408/p570.803/p58/p57 2.128
Var(S)/p58(5)(5)(0.9)/p17/p59/p37/p59(/p570.803) /p17
10/p599
/p580.765
Thus,Z/p58/p57 2.128//p400.765/p58/p57 2.433/p58/p57Z/p15/p13/p15/p20/p58/p57 1.64.We reject H/p15at the
0.05 level and reach the same conclusion as in the last three examples: that thedata show that CMB is more effective than no treatment.
5.1.5 Cox’s F-test
Cox’sF-test (Cox, 1964 )is based on ordered scores from the exponential
distribution.It is for singly censored or complete samples; it is not applicable
to progressively censored data.The procedure is as follows:116
Table 5.5 Computations of Peto and Peto’s
Generalized Wilcoxon Test
t/p7/p71/p8S/p19(t) u/p71
15 0.900 1 /p590.900 /p571/p580.900
16/p59 — 0.900 /p571/p58/p57 0.100 /p63
18 0.788 0.900 /p590.788 /p571/p580.688
18/p59 — 0.788 /p571/p58/p57 0.212 /p63
19 0.657 0.788 /p590.657 /p571/p580.445
19 0.526 0.526 /p590.657 /p571/p580.183
20 0.395 0.395 /p590.526 /p571/p58/p57 0.079
20/p59 — 0.395 /p571/p58/p57 0.605 /p63
23 0.197 0.197 /p590.395 /p571/p58/p57 04.08 /p63
24/p59 — 0.197 /p571/p58/p57 0.803 /p63
/p63Group 1.
1.Rank the observations in the combined sample.
2.Replace the ranks by the corresponding expected order statistics in
sampling the unit exponential distribution [ f(t)/p58e/p92/p82].Denote by t/p80/p76the
expected value of the rth observation in increasing order of magnitude,
t/p80/p76/p581
n/p59/p37/p591
n/p57r/p591r/p581,...,n (5.1.12)
wherenis the total number of observations in the two samples.In
particular,
t/p16/p76/p581
n
t/p17/p76/p581
n/p591
n/p571
/p36
t/p76/p76/p581
n/p591
n/p571/p59/p37/p591(5.1.13)
Fornnot too large, they can easily be computed by using tables of
reciprocals.When two or more observations are tied, the average of thescores is used.
3.For data without censored observations, the entire set of nobservations
is replaced by the set of scores /p43t/p80/p76/p44so obtained.The sample mean scores
denoted by t/p16/p16andt/p16/p17of the two samples with n/p16,n/p17observations are then
computed.The ratio t/p16/p16/t/p16/p17has been shown to follow an Fdistribution
with (2n/p16,2n/p17)degrees of freedom.Critical regions for testing H/p15:S/p16/p58S/p17 117
againstH/p16(S/p16/p57S/p17),H/p17(S/p16/p58S/p17), andH/p18(S/p16/p34S/p17) are, respectively,
t/p16/p16/t/p16/p17/p57F/p17/p76/p129/p11/p17/p76/p130/p11/p63,t/p16/p16/t/p16/p17/p58F/p17/p76/p129/p17/p76/p130/p11/p16/p92/p63, andt/p16/p16/t/p16/p17/p57F/p17/p76/p129/p11/p17/p76/p130/p11/p63/p30/p17ort/p16/p16/
t/p16/p17/p58F/p17/p76/p129/p11/p17/p76/p130/p11/p16/p92/p63/p30/p17.
4.The calculation of Fis slightly different for singly censored data.Let r/p16andr/p17be the number of failures and n/p16/p57r/p16andn/p17/p57r/p17the number of
censored observations in the two samples.Then there are p/p58r/p16/p59r/p17failures in the combined sample and n/p57pcensored observations.Cox
(1964 )suggests using the scores t/p16/p76,...,t/p78/p76as before for the failures and
t/p7/p78/p62/p16/p8/p76for all censored observations.The mean score, for example, for the
first group is
t/p16/p16/p58r/p16t/p16/p30/p16/p59(n/p16/p57r/p16)t/p7/p78/p62/p16/p8/p76r/p16(5.1.14 )
wheret/p16/p30/p16is the mean score of the failures.The mean score for the second
group is calculated in a similar way.The F-statistic t/p16/p16/t/p16/p17, has an
approximate F-distribution with (2r/p16,2r/p17)degrees of freedom.
This test is for the hypothesis that the two samples are from populations
with equal means.It can also determine if the second population mean is k
times the first population mean, for a given k, by dividing the observations in
the second sample by kbefore ranking and applying the test.The set of all
valuesknot rejected in such a significance test forms a confidence interval.The
following example illustrates the computation.
Example 5.6 In an experiment comparing two treatments (A and B )for
solid tumor, suppose that the question is whether treatment B is better thantreatment A.Six mice are assigned to treatment A and six to treatment B.Theexperiment is terminated after 30 days.The following survival times in days arerecorded.Our null and alternative hypotheses are H/p15:S/p31/p58S/p32and
H/p16:S/p31/p58S/p32.
Treatment A: 8, 8, 10, 12, 12, 13
Treatment B: 9, 12, 15, 20, 30 /p59,3 0/p59
That is, all the mice receiving treatment A die within 13 days and two mice
receiving treatment B are still alive at the end of the study.Do the data providesufficient evidence that treatment B is more effective than treatment A?
To compute the test statistic, it is convenient to set up a table like Table 5.6.
The first column lists all the observations in the two samples.The secondcolumn contains the ordered exponential scores t/p80/p76.In this case, n/p16/p586,n/p17/p586,
n/p5812,r/p16/p586, andr/p17/p584.The scores are computed following (5.1.12 )and
(5.1.13 ).For example, t/p80/p76fort/p71/p5810 is equal to 1/12 /p591/11 /p591/10 /p591/9
or simply the previous t/p80/p76plus 1/9, that is, 0.274 /p591/9/p580.385. The
tied observations receive an average score: for example, for t/p71/p5812,118
Table 5.6 Computations of Cox’s F-Test for Data in Example 5.6
t/p80/p76of t/p80/p76of
t/p71t/p80/p76Sample A Sample B
8 /p16/p16/p17/p580.0831 0.129 —
8 /p16/p16/p17/p59/p16/p16/p16/p580.174/p440.1290.129 —
9 /p16/p16/p17/p59/p16/p16/p16/p59/p16/p16/p15/p580.174 /p590.100 /p580.274 — 0.274
10 0.274 /p59/p16/p24/p580.385 0.385 —
12 0.385 /p59/p16/p23/p580.510 0.661 —
12 0.510 /p59/p16/p22/p580.661/p440.661 0.661 —
12 0.653 /p59/p16/p21/p580.820 — 0.661
13 0.820 /p59/p16/p20/p581.020 1.020 —
15 1.020 /p59/p16/p19/p581.270 — 1.270
20 1.270 /p59/p16/p18/p581.603 — 1.603
30/p59 1.603 /p59/p16/p17/p582.103 — 2.103
30/p59 2.103 — 2.103
——— ———
2.985 8.014
t/p80/p76/p58/p16/p18(0.510/p590.653/p590.820) /p580.661.The last two columns of Table 5. 6 give
the scores for the two samples and the sums are entered at the bottom.Thust/p16/p31/p582.985/6/p580.498 and t/p16/p32/p588.014/4/p582.004 according to (5.1.14 )and
F/p58t/p16/p31t/p16/p32/p580.498
2.004/p580.249
with (12, 8 )degrees of freedom.The critical region is F/p58F/p16/p17/p11/p23/p11/p15/p13/p24/p20/p58
1/F/p23/p11/p16/p17/p11/p15/p13/p15/p20/p581/2.8486 /p580.351 for /afii9825/p580.05. /p18Hence, the data provide strong
evidence that treatment B is superior to treatment A.
5.1.6 Comments on the Tests
The tests presented in Sections 5.1.1 to 5.1.5 are based on rank statistics
obtained from scores assigned to each observation.The first four tests areapplicable to data with progressive censoring.They can be further groupedinto two categories: generalization of the Wilcoxon test (Gehan’s and Peto and
Peto’s )and the non-Wilcoxon test (Cox—Mantel and logrank test ).In the
logrank test, if the statistic Sis the sum of wscores in group 2, it is the same
asUof the Cox —Mantel test.This can be seen in Examples 5. 2 (U/p582.75) and
5.3(S/p582.751); the small discrepancy is due to rounding-off errors.
/p18F/p80/p129/p11/p80/p130/p11/p63/p581/F/p80/p130/p11/p80/p129/p11/p16/p92/p63. 119
The only reason to choose one test over another in a given circumstance is
if it will be more powerful, that is, more likely to reject a false hypothesis.Whensample sizes are small (n,n/p17/p4550), Gehan and Thomas (1969 )show that Cox’s
F-test is more powerful than Gehan’s generalized Wilcoxon test if samples are
from exponential or Weibull distributions and if there are no censoredobservations or the observations are singly censored.Comparisons of Gehan’sWilcoxon test to several other tests are reported by Lee et al. (1975 ).They
show that when samples are from exponential distributions, with or without
censoring, the Cox —Mantel and logrank tests are more powerful and more
efficient than the generalized Wilcoxon tests of Gehan and Peto and Peto.There is little difference between the Cox —Mantel and logrank tests and
between the two generalized Wilcoxon tests.When the samples are taken fromWeibull distributions with a constant hazard ratio (i.e., the ratio of the two
hazard functions does not vary with time ), the results are essentially the same
as in the exponential case.However, when the hazard ratio is nonconstant,the two generalizations of the Wilcoxon test have more power than the othertests.Thus, the logrank test is more powerful than the Wilcoxon tests indetecting departures when the two hazard functions are parallel (proportional
hazards )or when there is random but equal censoring and when there is no
censoring in the samples (Crowley and Thomas, 1975 ).The generalized
Wilcoxon tests appear to be more powerful than the logrank test for detectingmany other types of differences, for example, when the hazard functions are notparallel and when there is no censoring and the logarithm of the survival timesfollow the normal distribution with equal variance but possibly differentmeans.
The generalized Wilcoxon tests give more weight to early failures than later
failures, whereas the logrank test gives equal weight to all failures.Therefore,the generalized Wilcoxon tests are more likely to detect early differences in thetwo survival distributions, whereas the logrank test is more sensitive todifferences in the right tails.Prentice and Marek (1979 )show that Gehan’s
Wilcoxon statistic is subject to a serious criticism when censoring rates arehigh.If heavy censoring exists, the test statistic is dominated by a small numberof early failures and has very low power.
There are situations in which neither the logrank nor Wilcoxon test is very
effective.When the two distributions differ but their hazard functions orsurvivorship functions cross, neither the Wilcoxon nor logrank test is verypowerful, and it will be sensible to consider other tests.For example, Taroneand Ware (1977 )discuss general statistics of similar form (using scores )and
Fleming and Harrington (1979 )and Fleming et al. (1980 )present a two-sample
test based on the maximum of a Smirnov-type statistic designed to measure themaximum distance between estimates of two distributions.The latter approachis shown to be more effective than the logrank or Wilcoxon tests when twosurvival distributions differ substantially for some range of tvalues, but not
necessarily elsewhere.These statistics have not been widely applied.Interestedreaders are referred to the original papers.120
5.2 MANTEL--HAENSZEL TEST
The Mantel —Haenszel (1959 )test is particularly useful in comparing survival
experience between two groups when adjustments for other prognostic factorsare needed.The test has been used in many clinical and epidemiological studiesas a method of controlling the effects of confounding variables.For example,in comparing two treatments for malignant melanoma, it would be importantto adjust the comparison for a possible confounding variable such as stage of
the disease.In studying the association of smoking and heart disease, it wouldbe important to control the effects of age.To use the Mantel —Haenszel test,
the data are stratified by the confounding variable and cast into a sequence of2/p592 tables, one for each stratum.
Letsbe the number of strata, n/p72/p71be the number of individuals in group j,
j/p581, 2, and stratum i,i/p581,...,s, andd/p72/p71be the number of deaths (or failures )
in group jand stratum i.For each of the sstrata, the data can be represented
by a 2/p592 contingency table:
Number of Number of
Group Deaths Survivors Total
1 d/p16/p71n/p16/p71/p57d/p16/p71n/p16/p712 d/p17/p71n/p17/p71/p57d/p17/p71n/p17/p71Total D/p71S/p71T/p71
The null hypothesis to be tested can be stated as
H/p15:p/p16/p16/p58p/p16/p17
p/p17/p16/p58p/p17/p17
/p36
p/p81/p16/p58p/p81/p17
wherep/p71/p72/p58P(death /p34groupj, stratum i).Thus, the test permits simultaneous
comparison over all the scontingency tables of the difference in survival or
death probabilities for the two groups.
The chi-square test statistic without continuity correction /p19is given by
X/p17/p58[/afii9814/p81/p71/p14/p16d/p16/p71/p57/afii9814/p81/p71/p14/p16(d/p16/p71)]/p17
/afii9814/p81/p71/p14/p16Var(d/p16/p71)(5.2.1)
/p19According to Grizzle (1967 ), the distribution of X/p17without continuity correction is closer to the
chi-square distribution than the X/p17with continuity correction.His simulations show that the
probability of Type I error (rejecting a true hypothesis )is better controlled without the continuity
correction at /afii9825/p580.01, 0.05.— 121
where
E(d/p16/p71)/p58n/p16/p71D/p71T/p71(5.2.2 )
Var(d/p16/p71)/p58n/p16/p71n/p17/p71D/p71S/p71T/p17/p71(T/p71/p571)(5.2.3 )
are the mean and variance, respectively, of the number of deaths in group i
computed conditionally on the contingency table marginal totals.This statisticfollows the chi-square distribution with 1 degree of freedom.Thus, a computedchi-square value larger than the table chi-square value for the significance levelchosen indicates a significant difference in survival between the two groups.The following two examples illustrate the use of the test.
Example 5.7 Five hundred and ninety-five persons participate in a case
control study of the association of cholesterol and coronary heart disease(CHD ).Among them, 300 persons are known to have CHD and 295 are free
of CHD.To find out if elevated cholesterol is significantly associated withCHD, the investigator decides to control the effects of smoking.The studysubjects are then divided into two strata: smokers and nonsmokers.
The following tables give the data for smokers:
Elevated
Cholesterol? With CHD Without CHD Total
Yes 120 20 140No 80 60 140
Total 200 80 280
and for nonsmokers:
Elevated
Cholesterol? With CHD Without CHD Total
Yes 30 60 90
No 70 155 255
Total 100 215 315122
Using (5.2.2 )and (5.2.3 ), we obtain
E(d/p16/p16)/p58140/p59200
280/p58100E(d/p16/p17)/p5890/p59100
315/p5828.571
Var(d/p16/p16)/p58140/p59140/p59200/p5980
(280) /p17(280 /p571)/p5814.337
Var(d/p16/p17)/p5890/p59225/p59100/p59215
(315) /p17(315 /p571)/p5813.974
Using (5.2.1 )andd/p16/p16/p58120,d/p16/p17/p5830, we have
X/p17/p58(150 /p57128.571) /p17
14.337/p5913.974/p5816.220
which is significant at the 0.001 level. Thus, elevated cholesterol is significantly
associated with CHD after adjusting for the effects of smoking.
Example 5.8 Table 5.7 gives survival data in life-table format of male cases
with localized cancer of the rectum in Connecticut for 1935 —1944 and
1945—1954.We use Mantel and Haenszel’s chi-square test to see if the survival
distribution of patients diagnosed in 1935 —1944 is the same as for patients
diagnosed in 1945 —1954.The null hypothesis is that the two survival distribu-
tions are the same.It is not necessary to set up 10 contingency tables for the10 intervals.The chi-square value is easily calculated by constructing columns7 to 12 directly from the life table.Using the sums in columns 1, 10, and 12,we obtain
X/p17/p58(330.0/p57246.50)/p17
132.491
/p5852.624
which is significant at the 0.001 level. Thus, the data show a significant
difference between the survival distributions of patients diagnosed in 1935 —
1944 and 1945 —1954.
It should be noted that this chi-square test statistic, when applied to life
tables, gives more weight to those deaths that occur in an early time intervalrather than later.That is, if the two groups are subject to the same probabilityof surviving through the entire study period, (5.2.1 )—(5.2.3 )will give high
mortality for the group in which early deaths occur.Mantel (1966 )gives the
following illustration.
Consider two groups of 100 persons each.Both have 50 deaths.In group 1
all deaths occur in the first interval, and in group 2 all deaths occur in the— 123
Table 5.7 Computational Procedure for Comparing Two Survival Distributions: Male Cases with Localized Cancer of Rectum in Connecticut
(1) (2) (3) (4) (5) (6) (7) (8) (9) (10) (11) (12)
Combined
1935—1944 1945 —1954 Time Periods E(d/p16/p71) n/p17/p71S/p71/T/p71Var(d/p16/p71)
Deaths, Survivors, Total, Deaths, Survivors, Total, Deaths, Survivors, Total,
Internal d/p16/p71n/p16/p71/p57d/p16/p71n/p16/p71d/p17/p71n/p17/p71/p57d/p17/p71n/p17/p71D/p71S/p71T/p71(3)/p59(7)
(9)(6)/p59(8)
(9)(10)/p59(11)
(9)/p571.0
1 167 220 .0 387 .0 185 559 .0 744 .0 352 779 .01 1 3 1 .01 2 0 .45 512 .45 54 .624
2 45 173 .5 218 .5 88 461 .0 549 .0 133 634 .57 6 7 .53 7 .86 453 .86 22 .418
3 45 127 .5 172 .5 55 396 .0 451 .0 100 523 .56 2 3 .52 7 .67 378 .67 16 .832
4 19 108 .0 127 .0 43 343 .0 386 .0 62 451 .05 1 3 .01 5 .35 339 .35 10 .174
51 7 9 1 .0 108 .0 32 299 .0 331 .0 49 390 .04 3 9 .01 2 .05 294 .05 8 .090
61 1 7 9 .59 0 .5 31 235 .0 266 .0 42 314 .53 5 6 .51 0 .66 234 .66 7 .036
78 7 1 .07 9 .0 20 170 .0 190 .0 28 241 .02 6 9 .08 .22 170 .22 5 .221
85 6 6 .07 1 .0 7 132 .0 139 .0 12 198 .02 1 0 .04 .06 131 .06 2 .546
96 5 9 .56 5 .5 6 101 .5 107 .5 12 161 .01 7 3 .04 .54 100 .04 2 .641
10 7 52 .05 9 .06 7 1 .07 7 .0 13 123 .01 3 6 .05 .64 69 .64 2 .909—— ——— ——— —330 246.50 132 .491
124
second interval.The contingency table for the first interval is:
Group Deaths Survivors Total
1 50 50 100
2 0 100 100
Total 50 150 200
and for the second interval is:
Group Deaths Survivors Total
1 0 50 50
2 50 50 100
Total 50 100 150
From these two tables we have E(d/p16/p16)/p58100/p5950/200/p5825 and E(d/p16/p17)/p58
100/p5950/150/p5816.67.The total deaths expected is 25 /p5916.67 /p5841.67, so the
50 deaths in group 1 is 20% larger than expected.Thus, a significant chi-squarevalue may be obtained if early survival patterns differ significantly in the twogroups.
5.3 COMPARISON OF K( K /p572) SAMPLES
In this section the two-sample problem is generalized to a situation in which
the data consist of K(K/p572)samples, one sample from each of the K
treatment populations.The problem is to decide whether the Kindependent
samples can be regarded as coming from the same population, or in practicalterms, to see if the survival data from patients receiving the Ktreatments
provide enough evidence to conclude that the Ktreatments are not equally
effective.This problem has been considered by many statisticians: for example,Kruskal and Wallis (1952 ), Mantel and Haenszel (1959 ), Breslow (1970 ), and
Peto and Peto (1972 ).In this section two nonparametric tests for the problem
are presented.The first is Kruskal and Wallis’s (1952 )H-test for uncensored
data.The second is a generalization of the H-test for censored data (Peto and
Peto, 1972 ).Both use ranks instead of the original observations and are simple
to apply.
5.3.1 Kruskal--Wallis Test
The Kruskal —WallisH-test (Kruskal and Wallis, 1952; Hollander and Wolfe,
1973; Marascuilo and McSweeney, 1977 ), analogous to the F-test in the usual
analysis of variance, uses ranks rather than original observations; it is alsocalled the Kruskal —Wallis one-way analysis of variance by ranks.It assumesCOMPARISON OF K(K/p572) SAMPLES 125
that the variable (survival time )under study has an underlying continuous
distribution.
LetNbe the total number of independent observations in the Ksamples,
n/p72the number of observations in the jth sample, j/p581,...,K, andt/p71/p72theith
observation in the jth sample.The null hypothesis H/p15states that the Ksamples
come from the same population (or clinically, the Ktreatments are equally
effective ).
In computation of the Kruskal —WallisH-test, we first rank all Nobserva-
tions from smallest to largest.Let r/p71/p72be the rank of t/p71/p72.Compute, for j/p581,...,K,
R/p72/p58/p76/p72/p26
/p71/p14/p16r/p71/p72R/p16/p72/p58R/p72n/p72R/p16/p581
2(N/p591) (5.3.1 )
whereR/p72andR/p16/p72are, respectively, the sum of the ranks and the average rank
of thejth treatment, and R/p16is the overall average rank.Then the Kruskal —
WallisH-statistic is
H/p5812
N(N/p591)/p41/p26
/p72/p14/p16n/p72(R/p16/p72/p57R/p16)/p17 (5.3.2)
/p5812
N(N/p591)/p41/p26
/p72/p14/p16R/p17/p72n/p72/p573(N/p591) (5.3.3 )
Under the null hypothesis, Hhas an asymptotic (n/p72’s approaching infinity or
n/p72’s are large )chi-square distribution with K/p571 degrees of freedom.Thus, for
largen/p72’s, the approximate test procedure at the /afii9825level is to reject H/p15if
H/p46/afii9851/p17/p7/p73/p92/p16/p8/p11/p63.WhenK/p583 and the number of observations in each of the three
samples is 5 or fewer, the chi-square approximation is not sufficiently close.Forsuch cases, exact permutational distributions of Hare available and percentage
points /afii9851/p73are given in Table B-4 of Appendix B.The test procedure is to reject
H/p15ifH/p46/afii9851/p73/p11/p63, where /afii9851/p73/p11/p63satisfies the equation P(H/p46/afii9851/p73/p11/p63/p34H/p15)/p58/afii9825.
When there are tied observations, each is assigned the average of the ranks.
To correct for the effects of ties, His computed by (5.3.3 )and then divided by
1/p571
N/p18/p57N/p69/p26
/p72/p14/p16T/p72(5.3.4 )
wheregis the number of tied groups, and T/p72/p58t/p18/p72/p57t/p72, witht/p72being the number
of tied observations in a tied group.In counting g, an untied observation is
considered as a tied group of size 1.Thus, a general expression of Hcorrected
for ties is
H/p58[12/N(N/p591)]/afii9814/p73/p72/p14/p16(R/p17/p72/n/p72)/p573(N/p591)
1/p57/afii9814/p69/p72/p14/p16T/p72/(N/p18/p57N)(5.3.5)126
Table 5.8 Cholesterol Values of 12 Subjects on Three
Different Diets
Diet 1 Diet 2 Diet 3
229 145 231176 181 208187 147 217
208 187 199
Table 5.9 Computation of Hfor Data in Example 5.9
Ordered Ranks of Ranks of Ranks of
Observations Ranks Diet 1 Diet 2 Diet 3
145 1 — 1
147 2 — 2176 3 3181 4 — 4
187 5.5 — 5.5
187 5.5 5.5199 7 — — 7208 8.5 8.5208 8.5 — — 8.5217 10 — — 10
229 11 11
231 12 — — 12
R/p7228 12.5 37.5Note that when there are no ties, g/p58N,t/p72/p581 for allj, andT/p72/p580, and (5.3.5 )
reduces to (5.3.3 ).The following example illustrates the use of the test.
Example 5.9 In a study of the relationship between cholesterol level and
diet, three diets are given randomly to 12 men whose initial cholesterol levelsare almost the same.Table 5. 8 shows the cholesterol levels of the 12 peopleafter having their assigned diet for a given period of time.The purpose of thestudy is to decide if the three diets are equally effective in controllingcholesterol level.
The null hypothesis H/p15states that there is no difference in cholesterol level
of men having the three diets, and the alternative H/p16says that the cholesterol
levels of men having the three different diets are different.To compute theH-statistic, we first rank the observations as in Table 5.9 and compute R/p72.In
this case N/p5812,n/p16/p58n/p17/p58n/p18/p584,g/p5810, and T/p72/p580 except for j/p585, 7.COMPARISON OF K(K/p572) SAMPLES 127
Hence, /afii9814/p69/p72/p14/p16T/p72/p582(8/p572)/p5812, and by (5.3.5 ),
H/p58[12/(12/p5913)](784/4/p59156.25/4/p591406.25/4)/p573(13)
1/p5712/(1728 /p5712)/p586.168
From Table B-4 we find that P(H/p466.038/p34H/p15)/p580.037 and P(H/p466.269/p34H/p15)
/p580.033; we reject H/p15at the /afii9825/p32 0.035 level. There is evidence of significant
differences among the diets.
5.3.2 Multiple Comparisons Based on the Kruskal--Wallis Test
If the null hypothesis that the Ksamples are from the same distribution is
rejected, we might ask which particular samples are from different distribu-tions.In Example 5. 9 we reject the null hypothesis that the three diets aresimilar.The investigator may also be interested in knowing which particulardiets differ from one another.In this section we introduce some nonparametricmethods for multiple comparison based on Kruskal —Wallis rank sums.An
excellent treatment of multiple comparisons is given by Miller (1966 ).
To decide which treatments differ from one another, there are /p16/p17
K(K/p571)
decisions to make, one for each pair of treatments.The null hypothesis can bewritten as
H/p15: samples iandjare from the same population for
i/p581,...,K/p571,j/p58i/p591,...,K,i/p58j
Let the probability of at least one wrong decision when H/p15is true be controlled
by/afii9825and the probability of making all correct decisions when H/p15is true be
1/p57/afii9825.To make the /p16/p17K(K/p571) decisions, we introduce the following compari-
son procedures.
1.When sample sizes are equal, that is, n/p16/p58n/p17/p58/p37/p58n/p41/p58n, andnis
small, we reject the hypothesis that the ith andjth samples, i/p58j, are from
the same distribution if
/p34R/p71/p57R/p72/p34/p46y(/afii9825,K,n) (5.3.6 )
wherey(/afii9825,K,n)satisfies the equation
P(/p34R/p71/p57R/p72/p34/p46y(/afii9825,K,n)/p34H/p15,i/p58j)/p581/p57/afii9825 (5.3.7)
andR/p16,R/p17,...,R/p41are given in (5.3.1 ).Some approximate values of yare
given in Table B-5.
2.When sample sizes are equal to nandnis large, we introduce Miller’s
(1966 )procedure, that is, to reject the hypothesis that the ith andjth128
Table 5.10 Multiple Comparisons for Data in
Example 5.9
ij /p34R/p71/p57R/p72/p34 Decision
12 /p3428/p5712.5/p34/p5815.5 Not significant
13 /p3428/p5737.5/p34/p589.5 Not significant
23 /p3412.5 /p5737.5/p34/p5825.0 Significant
samples, i/p58j, are from the same distribution if
/p34R/p16/p71/p57R/p16/p72/p34/p46q(a,K)[/p16/p16/p17K(Kn/p591)]/p16/p30/p17 (5.3.8)
whereR/p16/p16,...,R/p16/p41are given in (5.3.1 )andq(/afii9825,K)is the upper /afii9825percentile
point of the range of Kindependent standard normal variables.Table
B-6 gives the q(/afii9825,K)values for some Kand/afii9825.
3.For cases of small unequal sample sizes n/p16,...,n/p41, a conservative
procedure is to reject the hypothesis that the ith andjth samples, i/p58j,
are from the same distribution if
/p34R/p16/p71/p57R/p16/p72/p34/p46(x/p63/p11/p41)/p16/p30/p17/p31
12N(N/p591)/p4/p16/p30/p17/p11
n/p71/p591
n/p72/p2/p16/p30/p17(5.3.9)
whereNis the total number of observations.Values of x/p63/p11/p41are given in
Table B-4.
4.When n/p16,...,n/p41are large, Dunn (1964 )suggests the following procedure.
Reject the hypothesis that the ith andjth samples, i/p58j, are from the
same distribution if
/p34R/p16/p71/p57R/p16/p72/p34/p46Z/p63/p30/p9/p41/p7/p41/p92/p16/p8/p10/p31
12N(N/p591)/p4/p16/p30/p17/p11
n/p71/p591
n/p72/p2/p16/p30/p17(5.3.10)
whereZ/p63/p30/p9/p41/p7/p41/p92/p16/p8/p10is the upper 100 /afii9825/[K(K/p571)] percentage point of the
standard normal distribution (see Table B-1 ).
Example 5.10 Let us use the data in Example 5.9. To examine which
particular diets differ from one another, we apply (5.3.6 ).SinceK/p583, there are
three possible comparisons.The calculation is shown in Table 5. 10.For K/p583,
n/p584, and from Table B-5, y(0.045, 3, 4) /p5824; hence for ( i,j)/p58(2, 3), (5.3.6 )is
satisfied.Thus, at /afii9825/p600.045, we conclude that diets 2 and 3 are significantly
different.COMPARISON OF K(K/p572) SAMPLES 129
Table 5.11 Initial Remission Times of Leukemia Patients
12 3
4, 5, 9, 10, 12, 13, 10, 8, 10, 10, 12, 14, 8, 10, 11, 23, 25, 25,
23, 28, 28, 28, 29 20, 48, 70, 75, 99, 103, 28, 28, 31, 31, 40,31, 32, 37, 41, 41, 162, 169, 195, 220, 48, 89, 124, 143,57, 62, 74, 100, 139, 161 /p59, 199 /p59, 217 /p59,1 2 /p59, 159 /p59, 190 /p59, 196 /p59,
20/p59, 258 /p59, 269 /p59 245/p59 197/p59, 205 /p59, 219 /p595.3.3 Test for Censored Data
In Section 5.1 we introduced three nonparametric tests based on scores for
comparing two samples with censored observations; Gehan’s generalizedWilcoxon test (if Mantel’s procedure is used ), Peto and Peto’s generalized
Wilcoxon test, and the logrank test.The K-sample test discussed in this section
can be considered an extension of these tests and the Kruskal —Wallis test.
Suppose that we have a set of Nscoresw/p16,w/p17,...,w/p44obtained according
to the manner of scoring in one of the three tests mentioned above.The sumof theNscores is zero.Let S/p72be the sum of the scores in the jth sample.The
null hypothesis H/p15states that the Ksamples are from the same distribution.
To testH/p15we calculate
X/p17/p58/afii9814/p41/p72/p14/p16(S/p17/p72/n/p72)
s/p17
(5.3.11)
where
s/p17/p58/afii9814/p44/p71/p14/p16w/p17/p71N/p571(5.3.12)
Under the null hypothesis X/p17has approximately chi-square distribution with
K/p571 degrees of freedom (Peto and Peto, 1972 ).Thus, we reject H/p15ifX/p17
exceeds the upper 100 /afii9825percentage point of the chi-square distribution with
K/p571 degrees of freedom, that is, if X/p17/p46/afii9851/p17/p7/p73/p92/p16/p8/p11/p63.
Example 5.11, using the scoring method of Mantel (1967 )for Gehan’s
generalized Wilcoxon test, illustrates the K-sample test for censored data.
Example 5.11 Consider the initial remission times of leukemia patients (in
days )induced by three treatments as given in Table 5.11. In this case, K/p583,
N/p5866,n/p16/p5825,n/p17/p5819, andn/p18/p5822.A table similar to Table 5. 1 may be set
up to compute the score for every observation.The computation is left to thereader as an exercise.The sums of scores in the three samples are S/p16/p58/p57 273,
S/p17/p58170, and S/p18/p58103.The sum of squares of the scores /afii9814w/p17/p71/p5889,702.
Hence, from (5.3.11 )and (5.3.12 ),X/p17/p58 3.612.From Table B-2, /afii9851/p17/p17/p11/p15/p13/p15/p20/p585.991;130
thus we do not reject H/p15.The data do not show significant differences among
the three initial treatments.
Bibliographical Remarks
Gehan’s test was first proposed in 1965.The Cox —Mantel test was first
discussed by Cox in 1959, then by Mantel in 1966, and finally, by Cox againin 1972.The scores for the logrank test was proposed by Peto and Peto in 1972
along with another generalization of the Wilcoxon test.In the same paper, theyalso discuss the K-sample test for censored data.The logrank test is also
discussed in Peto et al. (1977 ).Cox’sF-test was developed in 1964 and the
Mantel—Haenszel chi-square test can be found in Mantel and Haenszel (1959 )
and Mantel (1966 ).The Kruskal —Wallis one-way analysis of variance can be
found in most standard textbooks under nonparametric methods.Readers whoare interested in the theoretical development or more properties of these testsshould read the original papers cited above.Applications of these tests aregiven in the original papers or can easily be found in medical and epidemiologi-cal journals.
EXERCISES
The first five exercises are continuations of Exercises 4.1 to 4.4 and 4.6.5.1 For the survival times given in Table 3.1, compare the survival distribu-
tions of the two treatment groups using:
(a) Gehan’s generalized Wilcoxon test
(b) The Cox —Mantel test
5.2 For the remission data given in Table 3.1, compare the remission time
distributions of the two treatment groups using:
(a) The logrank test
(b) Peto and Peto’s generalized Wilcoxon test
5.3 For the data given in Table 3.4, compare the tumor-free time distribu-
tions of the three diet groups.
5.4 For the remission data of 42 leukemia patients given in Example 3.3, use
the two generalized Wilcoxon tests to see if 6-MP is more effective thanplacebo in prolonging remission time.
5.5 For the first four skin tests given in Exercise Table 3.1, use the
Cox—Mantel and logrank tests to see if there is a significant difference in
survival between patients with positive (/p465 mm for mumps, /p4610 mm for
others )and negative (/p585 mm for mumps, /p5810 mm for others )reactions. 131
Exercise Table 5.1
Percentage of Male Female
Standard
BMI /p63 Case Control Case Control
/p58140 130 160 60 65
/p46140 85 55 55 50
/p63Standard BMI: male, 22.1; female, 20.6. Percentage of standard
BMI /p58(observed BMI/standard BMI )/p59100.
Exercise Table 5.2
12 3
10.5 10.0 12.0
9.0 12.0 13.0
9.5 12.5 15.59.0 11.0 14.08.5 12.0 12.5
10.0 10.5 15.05.6 Compute the test statistic Wof Gehan’s generalized Wilcoxon test by
using (5.1.1 )for the data in Example 5.1. Do you get the same result as
in Example 5.1?
5.7 Consider the data in Example 5.11. Use Mantel’s procedure for Gehan’s
generalized Wilcoxon test to compute a score for each observation andthe sum of scores for each of the three treatment groups.
5.8 Using the data in Table 3.1, compare the age distributions of the two
treatment groups using Cox’s F-test.
5.9 Consider the data in Exercise Table 5.1. Is elevated percent standard
BMI associated with renal cell carcinoma after controlling the effects ofgender?
5.10 Consider the survival data of men with angina pectoris in Table 4.6 and
women with the same disease in Exercise Table 4.2. Is there a significantdifference between the survival distributions of men and women?
5.11 In a study of noise level and efficiency, 18 students were given a very
simple test under three different noise levels.It is known that under132
Exercise Table 5.3
12 2 3
41 3 5
54 7 1 59 9 14 20
12 12 20 31
20/p59 15 27 39
25 23 30 4730/p59 30 32 /p59 55/p59
50/p59 67/p59normal conditions, they should be able to finish the test in 10 minutes.
The students were randomly assigned to the three levels.Exercise Table5.2 gives the time required to finish the test. Are the three noise levelssignificantly different? If they are, determine which levels differ from oneanother.
5.12 Exercise Table 5.3 gives the survival time in weeks of 30 brain tumor
patients receiving four different treatments.Are the four treatments
equally effective? 133
CHAPTER6
SomeWell-KnownParametric
SurvivalDistributionsandTheirApplications
Usually,therearemanyphysicalcausesthatleadtothefailureordeathofapersonataparticulartime.Itisverydifficult,ifnotimpossible,toisolatethesephysicalcausesandaccountmathematicallyforallofthem.Therefore,choos-ingatheoreticaldistributiontoapproximatesurvivaldataisasmuchanartasascientifictask.Inthischapter,severaltheoreticaldistributionsthathavebeenusedwidelytodescribesurvivaltimearediscussed,theircharacteristicssummarized,andtheirapplicationsillustrated.
6.1 EXPONENTIAL DISTRIBUTION
The simplest and most important distribution in survival studies is the
exponential distribution. In the late 1940s, researchers began to choose theexponentialdistributiontodescribethelifepatternofelectronicsystems.Davis(1952 )givesanumberofexamples,includingbankstatementandledgererror,
payroll check errors, automatic calculating machine failure, and radar setcomponent failure, in which the failure data are well described by theexponentialdistribution.EpsteinandSobel (1953 )reportwhytheyselectthe
exponentialdistributionoverthepopularnormaldistributionandshowhowtoestimatetheparameterwhendataaresinglycensored.Epstein (1958 )also
discussesinsomedetailthejustificationfortheassumptionofanexponentialdistribution.Theexponentialdistributionhassincecontinuedtoplayaroleinlifetimestudiesanalogoustothatofthenormaldistributioninotherareasofstatistics.
Theexponentialdistributionisoftenreferredtoasapurelyrandomfailure
pattern.Itisfamousforitsunique‘‘lackofmemory,’’whichrequiresthattheage of the animal or person does not affect future survival. Although many
134
Figure 6.1Exponentialdistribution: (a)survivorshipfunction;( b) probabilitydensity
function;( c) hazardfunction.
survivaldatacannotbedescribedadequatelybytheexponentialdistribution,
anunderstandingofitfacilitatesthetreatmentofmoregeneralsituations.
Theexponentialdistributionischaracterizedbyaconstanthazardrate /afii9838,
itsonlyparameter.Ahigh /afii9838valueindicateshighriskandshortsurvival;alow
/afii9838valueindicateslowriskandlongsurvival.Figure6.1depictsthesurvivorship
function, the density function, and the hazard function of the exponentialdistributionwithparameter /afii9838.When /afii9838/p581,thedistributionisoftenreferredto
astheunit exponential distribution.
When the survival time Tfollows the exponential distribution with a
parameter /afii9838,theprobabilitydensityfunctionisdefinedas
f(t)/p58/p7/afii9838e/p92/p72/p82
0t/p460,/afii9838/p570
t/p580(6.1.1 )
Thecumulativedistributionfunctionis
F(t)/p581/p57e/p92/p72/p82t/p460 (6.1.2 )
andthesurvivorshipfunctionisthen
S(t)/p58e/p92/p72/p82t/p460 (6.1.3 )
Sothat,by (2.2.1 ),thehazardfunctionis
h(t)/p58/afii9838t/p460( 6 .1.4)
aconstant,independentof t.Figure6.1givesthegraphicalpresentationofthe
threefunctions.
Becausetheexponentialdistributionischaracterizedbyaconstanthazard
rate,independentoftheageoftheperson,thereisnoagingorwearingout, 135
and failure or death is a random event independent of time. When natural
logarithmsofthesurvivorshipfunctionaretaken,log S(t)/p58/p57/afii9838t,whichisa
linearfunctionof t.Thus,itiseasytodeterminewhetherdatacomefroman
exponentialdistributionbyplottinglog S/p19(t) against t,where S/p19(t) isanestimate
ofS(t). A linear configurationindicates that the data follow an exponential
distributionandtheslopeofthestraightlineisanestimateofthehazardrate /afii9838.
Themeanandvarianceoftheexponentialdistributionwithparameter /afii9838are,
respectively,1/ /afii9838and1/ /afii9838/p17.Themedianis (1//afii9838)log2.Thecoefficientofvariation
is1.
A moregeneral formof the exponentialdistributionis the two-parameter
exponential distribution withprobabilitydensityfunction
f(t)/p58/p7/afii9838e/p92/p72/p7/p82/p92/p37/p8
0t/p46G
t/p58G(6.1.5 )
Then
F(t)/p58/p71/p57e/p92/p72/p7/p82/p92/p37/p8
0t/p46G
t/p58G(6.1.6 )
S(t)/p58/p7e/p92/p72/p7/p82/p92/p37/p8
0t/p46G
t/p58G(6.1.7)
and
h(t)/p58/p70
/afii98380/p45t/p58G
t/p58G(6.1.8)
Theterm Gisaguarantee time withinwhichnodeathsorfailurescanoccur,
or a minimum survival time. If G/p580, (6.1.5)—(6.1.8 )reduce to (6.1.1 )—
(6.1.4 )for the one-parameter exponential. The mean and the median of the
two-parameter exponential distribution are, respectively, G/p591//afii9838and
(log2 /p59/afii9838G)//afii9838.
Example 6.1 In a study of new anticancer drugs in the L1210 animal
leukemiasystem,Zelen (1966 )usedtheexponentialdistributionsuccessfullyas
themodelforsurvivaltime.Thesystemconsistsofinjectingatumorinoculuminto inbred mice. These tumor cells then proliferate and eventually kill theanimal, but survival time may be prolonged by an active drug. Figure 6.2shows the survival curve in a semilogarithmic scale of the untreated miceinoculatedatdifferentcelldilutions.Twenty-fivemicewereinoculatedateachdilution. The reasonably linear configurations suggest that the survival dis-tributionsfollowtheexponentialdistributionquitewell.Thefourstraightlinesfitted to the points are almost parallel, indicating that the hazard rate wasindependentoftheinoculumsize.Table6.1givestheestimatedvaluesofthe136 -
Figure 6.2Survivalcurvesofuntreatedmiceinoculatedwithserial10-folddilutionsof
leukemiaL1210: (/p42)10/p20cells; (I)10/p19cells; (/p41)10/p18cells; (G)10/p17cells. (FromZelen,
1966. )
Table 6.1 Estimates of G,/afii9838, and Mean Survival Time
for Untreated Mice with Serial Leukemia Dilutions
Dilution G/p19/afii9838 /p19MeanSurvivalTime
10/p208.0 0.78 9.3
10/p1910.0 0.78 11.3
10/p1811.9 0.76 13.2
10/p1713.9 0.67 15.4
Source:Zelen (1966 ).guaranteetimes G,hazardrates /afii9838,andthemeansurvivaltimes (indays )for
thevariousdilutions. (EstimationproceduresarediscussedinChapter7. )Note
thattheestimatedhazardrates, /afii9838,areveryclose.
Theprobabilitythat amousereceiving10 /p20cells ofinoculumwill survive
morethan20daysis,from (6.1.7 ),
S(20)/p58e/p92/p15/p13/p22/p23/p7/p17/p15/p92/p23/p8 /p600.0001
andthemediansurvivaltimeis8.9days. 137
Figure 6.3Survival curves of mice treated with cyclophosphamide on day 3 after
inoculationof10 /p20tumorcells: (/p42)control; (I)80mg/kg; (G)160mg/kg. (FromZelen,
1966. )
Figure6.3givesthesurvivalcurvesofmicetreatedwithdifferentdosesof
cyclophosphamideonday3afterreceivinga10 /p20tumorinoculum.Table6.2
gives the estimates of G,/afii9838, and the mean survival time. Mice receiving 160
mg/kg of the drug show a remarkably improved surivival pattern over thecontrolgroup.
Theprobabilitythatamousereceiving80mg/kgofcyclophosphamidewill
survivemorethan20daysis,from (6.1.7 ),
S(20)/p58e/p92/p15/p13/p17/p24/p7/p17/p15/p92/p16/p20/p13/p21/p8 /p58 0.279
Themediansurvivaltimeisapproximately18days.
6.2 WEIBULL DISTRIBUTION
The Weibull distribution is a generalization of the exponential distribution.
However,unliketheexponentialdistribution,it does notassumea constanthazard rate and therefore has broader application. The distribution wasproposedbyWeibull (1939 )anditsapplicabilitytovariousfailuresituations
discussedagainbyWeibull (1951 ).Ithasthenbeenusedinmanystudiesof
reliabilityandhumandiseasemortality.
TheWeibulldistributionischaracterizedbytwoparameters, /afii9828and/afii9838.The
valueof /afii9828determinesthe shapeof thedistributioncurveandthevalueof /afii9838138 -
Table 6.2 Estimates of G,/afii9838, and Mean Survival Time
for Treated Mice on Day 3
Dose (mg/kg )G/p19/afii9838 /p19MeanSurvivalTime
Control 8.7 1.12 9.6
80 15.6 0.29 19.0
160 21.5 0.10 31.5
Source:Zelen (1966 ).
Figure 6.4HazardfunctionsofWeibulldistributionwith /afii9838/p581.determinesits scaling. Consequently, /afii9828and/afii9838are called the shapeandscale
parameters,respectively.Therelationshipbetweenthevalueof /afii9828andsurvival
timecanbeseenfromFigure6.4,whichshowsthehazardrateoftheWeibulldistributionwith /afii9828/p580.5,1,2,4.When /afii9828/p581,thehazardrateremainsconstant
astimeincreases;thisistheexponentialcase.Thehazardrateincreaseswhen/afii9828/p571anddecreaseswhen /afii9828/p581astincreases.Thus,theWeibulldistribution
maybeusedtomodelthesurvivaldistributionofapopulationwithincreasing,decreasing, or constant risk. Examples of increasing and decreasing hazardrates are, respectively, patients with lung cancer and patients who undergosuccessfulmajorsurgery.
Theprobabilitydensityfunctionandcumulativedistributionfunctionsare,
respectively,
f(t)/p58/afii9838/afii9828(/afii9838t)/p65/p92/p16e/p92/p7/p72/p82/p8
/p65t/p460,/afii9828,/afii9838/p570 (6.2.1 ) 139
Figure 6.5DensitycurvesofWeibulldistributionwith /afii9838/p581.and
F(t)/p581/p57e/p92/p7/p72/p82/p8/p65(6.2.2)
Thesurvivorshipfunctionis,therefore,
S(t)/p58e/p92/p7/p72/p82/p8/p65(6.2.3)
andthehazardfunction,theratioof (6.2.1 )to(6.2.3 ),is
h(t)/p58/afii9838/afii9828(/afii9838t)/p65/p92/p16 (6.2.4)
Figure6.5givestheWeibulldensityfunctionwithscaleparameter /afii9838/p581and
severaldifferentvaluesoftheshapeparameter /afii9828.
Forthesurvivalcurve,itissimpletoplotthelogarithmof S(t),
logS(t)/p58/p57(/afii9838t)/p65 (6.2.5)
Figure6.6giveslog S(t) for/afii9838/p581and /afii9828/p581,/p571,/p581.When /afii9828/p581isastraight
line with negative slope. When /afii9828/p581, negative aging, log S(t) decreasesvery
slowly from 0 and then approaches a constant value. When /afii9828/p571, positive
aging,log S(t)decreasessharplyfrom0as tincreases.Equation (6.2.5 )canalso140 -
Figure 6.6Curvesoflog/p67S(t) ofWeibulldistributionwith /afii9838/p581.
bewrittenas
log[/p57logS(t)]/p58/afii9828log/p67/afii9838/p59/afii9828log/p67t (6.2.6 )
ThemeanoftheWeibulldistributionis
/afii9839/p58/afii9772(1/p591//afii9828)
/afii9838(6.2.7 )
andthevarianceis
/afii9846/p17/p581
/afii9838/p17/p3/afii9772/p11/p592
/afii9828/p2/p57/afii9772/p17/p11/p591
/afii9828/p2/p4(6.2.8 )
where /afii9772(/afii9828)isthewell-knowngammafunctiondefinedas
/afii9772(/afii9828)/p58/p16/p27
/p15x/p65/p92/p16e/p92/p86dx
/p58(/afii9828/p571)! when /afii9828isapositiveinteger (6.2.9 )
Valuesof /afii9772(/afii9828)canbefoundinAbramowitzandStegun (1964 ).Thecoefficient
ofvariationisthen
CV/p58/p3/afii9772(1/p592//afii9828)
/afii9772/p17(1/p591//afii9828)/p571/p4/p16/p30/p17(6.2.10 )
The Weibull distribution can also be generalized to take into account a
guarantee time Gduring which no deaths or failures can occur. The three- 141
Figure 6.7SurvivalcurvesofratsexposedtocarcinogenDMBA. (FromPike,1966.
ReproducedwithpermissionoftheBiometricsSociety. )parameterWeibullprobabilitydensityfunctionis
f(t)/p58/afii9838/p65/afii9828(t/p57G)/p65/p92/p16exp[ /p57/afii9838/p65(t/p57G)/p65] (6.2.11 )
Consequently,
S(t)/p58exp[ /p57/afii9838/p65(t/p57G)/p65]( 6 .2.12)
and
h(t)/p58/afii9838/p65/afii9828(t/p57G)/p65/p92/p16 (6.2.13)
Example 6.2 Pike (1966 )appliedtheWeibulldistributiontoatwo-group
experimentonvaginalcancerinratsexposedtothecarcinogenDMBA.Thetwogroupsweredistinguishedbypretreatmentregime.Thetimesindays,afterthestartoftheexperiment,atwhichthecarcinomawasdiagnosedforthetwogroupsofratswereasfollows:
Group1: 143,164,188,188,190,192,206,209,213,216,220,227,230,
234,246,265,304,216 /p59,244/p59
Group2: 142,156,173,198,205,232,232,233,233,233,233,239,240,
261,280,280,296,296,323204 /p59,344/p59
Assumingthat G/p58100and /afii9828/p583,Pikeobtained /afii9838/p19/p16/p58(4.51/p5910/p92/p22)/p16/p30/p18for
group1and /afii9838/p19/p17/p58(2.38/p5910/p92/p22)/p16/p30/p18forgroup2.Analyticalestimationprocedure
isdiscussedinChapter7.Figure6.7plotsthesurvivalcurvesofthetwogroups.ThestepfunctionsarenonparametricestimatessimilartotheKaplan —Meier142 -
Table 6.3 Calculation of Survivorship Functions for Group 2 of Rats Exposed to DMBA
Kaplan—Modified
Meier Kaplan —Meier S/p19(t) Obtained
Estimates, Estimates /p63, fromWeibull
Time rn /p57r/p591n/p57r
n/p57r/p591 S/p19(t) S/p19(t) Plotted Fit
142 1 21 0.9524 0.9524 0.9546 0.9825
156 2 20 0.9500 0.9048 0.9091 0.9590
163 3 19 0.9474 0.8572 0.8637 0.9421
198 4 18 0.9444 0.8095 0.8182 0.7990204/p59— 17 1.0000 0.8095 0.8182 0.7647
205 6 16 0.9375 0.7589 0.7700 0.7588232 7 15 0.9333 0.7083 0.7218 0.5778232 8 14 0.9286 0.6577 0.6737 0.5778
232 9 13 0.9231 0.6071 0.6255 0.5706
233 10 12 0.9167 0.5565 0.5773 0.5706233 11 11 0.9091 0.5059 0.5292 0.5706233 12 10 0.9000 0.4553 0.4810 0.5706239 13 9 0.8888 0.4047 0.4320 0.5271240 14 8 0.8750 0.3541 0.3847 0.5198
261 15 7 0.8571 0.3035 0.3365 0.3697
280 16 6 0.8333 0.2529 0.2883 0.2489280 17 5 0.8000 0.2023 0.2402 0.2489296 18 4 0.7500 0.1517 0.1920 0.1660296 19 3 0.6667 0.1011 0.1439 0.1660323 20 2 0.5000 0.0506 0.0958 0.0710
344/p59— 1 1.0000 0.0506 0.0958 0.0313
/p63Insteadof( n/p57r)/(n/p57r/p591), Pikeuses( n/p57r/p591)/(n/p57r/p592) intheKaplan —Meierproduct-
limitestimatetoavoid( n/p57r)/(n/p57r/p591)/p580.
Source:Pike (1966 ).ReproducedwithpermissionoftheBiometricSociety.
estimate. The smooth curves are obtained from the Weibull fits. Table 6.3
showsthecalculationoftheplottingpointsforgroup2.
It is obvious that the Weibull distributions with G/p58100, /afii9828/p583,
/afii9838/p19/p16/p58(4.51/p5910/p92/p22)/p16/p30/p18, and /afii9838/p19/p16/p58(2.38/p5910/p92/p22)/p16/p30/p18fitthecarcinoma-freetimeof
thetwogroupsofratsverywell.
6.3 LOGNORMAL DISTRIBUTION
Initssimplestformthelognormaldistributioncanbedefinedasthedistribu-
tionofavariablewhoselogarithmfollowsthenormaldistribution.Itsorigin
maybetracedasfarbackas1879,whenMcAlister (1879 )describedexplicitly
atheoryofthedistribution.Mostofitsaspectshavesincebeenunderstudy.Gaddum (1945a,b)gave a review of its application in biology, followed by 143
Figure 6.8Hazardofthelognormaldistributionwithdifferentparameters.Boag’s (1949 )applicationsincancerresearch.Itshistory,properties,estimation
problems,andusesineconomicshavebeendiscussedindetailbyAitchisonandBrown (1957 ).Later,otherinvestigatorsalsoobservedthattheageatonsetof
Alzheimer’sdiseaseandthedistributionofsurvivaltimeofseveraldiseasessuchasHodgkin’sdisease and chronic leukemia couldbe rather closelyapproxi-matedbyalognormaldistributionsincetheyaremarkedlyskewedtotherightandthelogarithmsofsurvivaltimesareapproximatelynormallydistributed.
Considerthesurvivaltime Tsuchthatlog Tisnormallydistributedwith
mean /afii9839andvariance /afii9846/p17.Wethensaythat Tislognormallydistributedand
writeTas/afii9806(/afii9839,/afii9846/p17).Itshouldbenotedthat /afii9839and/afii9846/p17arenotthemeanand
varianceofthelognormaldistribution.Figure6.8givesthehazardfunctionofthelognormaldistributionwithdifferentvaluesfortheparameters.Thehazardfunctionincreasesinitiallytoamaximumandthendecreases (almostassoon
asthemedianispassed )tozeroastimeapproachesinfinity (WatsonandWells
1961 ). Therefore, the lognormal distribution is suitable for survival patterns
withaninitiallyincreasingandthendecreasinghazardrate.Byacentrallimittheorem,itcanbeshownthatthedistributionoftheproductof nindependent
positive variates approaches a lognormal distribution under very generalconditions: for example, the distribution of the size of an organism whosegrowth is subject to many small impulses, the effect of each of which isproportionaltothemomentarysizeoftheorganism.
Thepopularityofthelognormaldistributionisdueinparttothefactthat
the cumulative values of y/p58logtcan be obtained from the tables of the
standardnormaldistributionandthecorrespondingvaluesof tarethenfound
bytakingantilogs.Thus,thepercentilesofthelognormaldistributionareeasytofind.144 -
Theprobabilitydensityfunctionandsurvivorshipfunctionare,respectively,
f(t)/p581
t/afii9846/p402/afii9843exp/p3/p571
2/afii9846/p17(logt/p57/afii9839)/p17/p4t/p570,/afii9846/p572 (6.3.1 )
and
S(t)/p581
/afii9846/p402/afii9843/p16/p27
/p821
xexp/p3/p571
2/afii9846/p17(logx/p57/afii9839)/p17/p4dx (6.3.2 )
Leta/p58exp(/p57/afii9839).Then /p57/afii9839/p58loga,(6.3.1 )and (6.3.2 )canbewrittenas
f(t)/p581
t/afii9846/p402/afii9843exp/p3/p571
2/afii9846/p17(logat)/p17/p4(6.3.3)
and
S(t)/p581
/afii9846/p402/afii9843/p16/p27
/p821
xexp/p3/p571
2/afii9846/p17(logax)/p17/p4dx(6.3.4)
/p581/p57G/p1logat
/afii9846/p2(6.3.5)
whereG(y) is the cumulative distribution function of a standard normal
variable
G(y)/p581
/p402/afii9843/p16/p87
/p15e/p92/p83/p130/p30/p17du (6.3.6)
Thelognormaldistributionisspecifiedcompletelybythetwoparameters /afii9839
and/afii9846/p17.TimeTcannotassumezerovaluessincelog Tisnotdefinedfor T/p580.
Figure6.9givesthelognormalfrequencycurvesfor /afii9839/p580,/afii9846/p17/p580.1,0.5,2,from
which an idea of the flexibility of the distribution may be obtained. It isobviousthatthedistributionispositivelyskewedandthatthegreaterthevalueof/afii9846/p17, the greater the skewness. Figure 6.10 shows the frequency curves for
/afii9846/p17/p580.5,/afii9839/p580, 0.5, 1. It is obvious that /afii9839and/afii9846/p17are, respectively, scale
parametersand notlocationandscaleparametersasinthenormaldistribu-
tion.Thehazardfunction,from (6.3.3 )and (6.3.5 ),hastheform
h(t)/p58(1/t/afii9846/p402/afii9843
)exp[ /p57(logat)/p17/2/afii9846/p17]
1/p57G(logat//afii9846)(6.3.7 )
andisplottedinFigure6.8. 145
Figure 6.9Lognormaldensitycurveswith /afii9839/p580.
Figure 6.10Lognormaldensitycurveswith /afii9846/p17/p580.5.The meanand variance of the two-parameterlognormaldistribution are,
respectively, exp (/afii9839/p59/p16/p17/afii9846/p17)and [exp (/afii9846/p17)/p571]exp (2/afii9839/p59/afii9846/p17). The coefficient of
variationofthedistributionisthen[exp (/afii9846/p17)/p571]/p16/p30/p17.Themedianis e/p73andthe
modeisexp (/afii9839/p59/afii9846/p17).
The two-parameter lognormal distribution can also be generalized to a
three-parameter distribution by replacing twitht/p57Gin(6.3.1 ). In other
words,thesurvivaltimelog( T/p57G)followsthenormaldistributionwithmean
/afii9839andvariance /afii9846/p17.Incertainsituationsthevalueof Gmaybedetermineda
priori and should not be regarded as an unknown parameter that requiresestimation.Ifthisisso,thevariable T/p57Gmaybeconsideredinplaceof T
and the distribution of T/p57Ghas all the properties of the two-parameter
lognormaldistribution.However,theestimationproceduresdevelopedforthe
two-parametercasearenotdirectlyapplicabletothedistributionof T/p57G.146 -
Figure 6.11Lognormalprobabilityplotofthesurvivaltimeof234malepatientswith
chroniclymphocyticleukemia. (FromFeinleibandMacMahon,1960.Reproducedby
permissionofthepublisher. )
Example 6.3 Inastudyofchroniclymphocyticandmyelocyticleukemia,
FeinleibandMacMahon (1960 )appliedthelognormaldistributiontoanalyze
survivaldataof649whiteresidentsofBrooklyndiagnosedfrom1943to1952.Theanalysisofseveralsubgroupsofpatientsfollows.Thesurvivaltimeofeachpatientiscomputedfromthedateofdiagnosisinmonths.Analyticalmethodisusedtofitthelognormaldistributiontothedata.ThemethodisdiscussedinChapters7and8.
Figure 6.11 gives the probability plot of the survival time of 234 male
patientswithchroniclymphocyticleukemia,inwhichthehorizontalaxisfor
the survival time is in logarithmic scale and the vertical axis is in normalprobabilityscale.Whenplotting1 /p57S(t) onthisgraphpaper,astraightlineis
obtained when the data follow a two-parameter lognormal distribution. Aninspection of the graph shows that the distribution is concave. Gaddum(1945a,b)haspointedoutthatsuchadeviationcanbecorrectedbysubtracting
an appropriate constant from the survival times. In other words, the three-parameterlognormaldistributioncanbeused.Figure6.12givesasimilarplotinwhichthesurvivaltimeofeverypatientplus4isplotted.Theconfigurationis linear and hence empiricallyit seems valid to assume that the lognormaldistributionisappropriate.
Similargraphsformalepatientswithchronicmyelocyticleukemiaandfor
femalepatientswithchroniclymphocyticormyelocyticleukemiaaregiveninFigures6.13and6.14.Parametersofthelognormaldistributionareestimated.FeinleibandMacMahonreportthattheagreementbetweentheobservedandcalculated distributions is striking for each group except for women withchroniclymphocyticleukemia.Thecorresponding pvaluesforthechi-square 147
Figure 6.12Lognormalprobabilityplotofthesurvivaltimeinmonthsplus4of234
male patients with chronic lymphocytic leukemia. (From Feinleib and MacMahon,
1960.Reproducedbypermissionofthepublisher. )
goodness-of-fittestareasfollows:
ChronicMyelocytic ChronicLymphocytic
Male 0.86 0.73
Female 0.57 0.016
Since a large pvalue indicates close agreement, it is concluded that the
three-parameterlognormaldistribution adequatelydescribesthe distributionofsurvivaltimesforeachsubgroupexceptwomenwithchroniclymphocyticleukemia.Theshapeoftheobserveddistributionforthelattergroupsuggeststhat itmight actuallybe composedof two dissimilargroups,each of whosesurvivaltimesmightfitalognormaldistribution.
6.4 GAMMA AND GENERALIZED GAMMA DISTRIBUTIONS
The gamma distribution, which includes the exponential and chi-square
distribution,wasusedalongtimeagobyBrownandFlood (1947 )todescribe
the life of glass tumblers circulating in a cafeteria and by Birnbaum andSaunders (1958 )asastatisticalmodelforlifelengthofmaterials.Sincethen,
thisdistributionhasbeenusedfrequentlyasamodelforindustrialreliabilityproblemsandhumansurvival.148 -
Figure 6.13Lognormalprobabilityplotofthesurvivaltimeinmonthsplus4of162
malepatientswithchronicmyelocyticleukemia. (FromFeinleibandMacMahon,1960.
Reproducedbypermissionofthepublisher. )
Figure 6.14Lognormalprobabilityplotofthesurvivaltimeinmonthsplus4offemale
patientswithtwotypesofleukemia. (FromFeinleibandMacMahon,1960.Reproduced
bypermissionofthepublisher. )Suppose that failure or death takes place in nstages or as soon as n
subfailureshavehappened.Attheendofthefirststage,aftertime T/p16,thefirst
subfailureoccurs;afterthatthesecondstagebeginsandthesecondsubfailureoccursaftertime T/p17;andsoon.Totalfailureordeathoccursattheendofthe
nth stage, when the nth subfailure happens. The survival time, T, is then
T/p16/p59T/p17/p59/p37/p59T/p76.Thetimes T/p16,T/p17,...,T/p76spentineachstageareassumedto 149
Figure 6.15Gammahazardfunctionswith /afii9838/p581.be independently exponentially distributed with probability density function
/afii9838exp(/p57/afii9838t/p71),i/p581,...,n. That is, the subfailures occur independently at a
constantrate /afii9838.Thedistributionof Tisthencalledthe Erlangian distribution .
There is no need for the stages to have physical significance since we canalways assume that death occurs in the n-stage process just described. This
idea, introduced by A. K. Erlang in his study of congestion in telephonesystems,hasbeenusedwidelyinqueuingtheoryandlifeprocesses.
A natural generalization of the Erlangian distribution is to replace the
parameter nrestrictedtotheintegers1,2,...byaparameter /afii9828takinganyreal
positivevalue.Wethenobtainthe gamma distribution.
Thegammadistributionischaracterizedbytwoparameters, /afii9828and/afii9838.When
0/p58/afii9828/p581,thereisnegativeagingandthehazardratedecreasesmonotonically
from infinity to /afii9838as time increases from 0 to infinity. When /afii9828/p571, there is
positiveagingandthehazardrateincreasesmonotonicallyfrom0to /afii9838astime
increasesfrom0toinfinity.When /afii9828/p581,thehazardrateequals /afii9838,aconstant,
asintheexponentialcase.Figure6.15illustratesthegammahazardfunctionfor/afii9838/p581 and /afii9828/p581,/afii9828/p581, 2, 4. Thus, the gamma distribution describes a
different type of survival pattern where the hazard rate is decreasing orincreasingtoaconstantvalueastimeapproachesinfinity.
Theprobabilitydensityfunctionofagammadistributionis
f(t)/p58/afii9838
/afii9772(/afii9828)
(/afii9838t)/p65/p92/p16e/p92/p72/p82t/p570,/afii9828/p570,/afii9838/p570 (6.4.1 )
where /afii9772(/afii9828)is defined as in (6.2.9 ). Figures 6.16 and 6.17 show the gamma
densityfunctionwithvariousvaluesof /afii9828and/afii9838.Itisseenthatvarying /afii9828changes
the shape of the distribution while varying /afii9838changes only the scaling.
Consequently, /afii9828and/afii9838are shape and scale parameters, respectively. When
/afii9828/p571,thereisasinglepeakat t/p58(/afii9828/p571)//afii9838.150 -
Figure 6.16Gammadensityfunctionswith /afii9838/p581.
Figure 6.17Gammadensityfunctionswith /afii9828/p583.Thecumulativedistributionfunction F(t) hasamorecomplexform:
F(t)/p58/p16/p82
/p15/afii9838
/afii9772(/afii9828)(/afii9838x)/p65/p92/p16e/p92/p72/p86dx (6.4.2)
/p581
/afii9772(/afii9828)/p16/p72/p82
/p15u/p65/p92/p16e/p92/p83du
/p58I(/afii9838t,/afii9828)( 6 .4.3)
where
I(s,/afii9828)/p581
/afii9772(/afii9828)/p16/p81
/p15u/p65/p92/p16e/p92/p83du (6.4.4)
knownasthe incomplete gamma function ,istabulatedinPearson (1922,1957 ). 151
FortheErlangiandistribution,itcanbeshownthat
F(t)/p581/p57/p76/p92/p16/p26
/p73/p14/p15e/p92/p72/p82(/afii9838t)/p73
k!(6.4.5 )
Thus,thesurvivorshipfunction1 /p57F(t)i s
S(t)/p58/p16/p27
/p82/afii9838
/afii9772(/afii9828)(/afii9838x)/p65/p92/p16e/p92/p72/p86dx (6.4.6)
forthegammadistributionor
S(t)/p58e/p92/p82/p76/p92/p16/p26
/p73/p14/p15(/afii9838t)/p73
k!(6.4.7 )
fortheErlangiandistribution.
Sincethehazardfunctionistheratioof f(t)toS(t),itcanbecalculatedfrom
(6.4.1 )and (6.4.7 ).When /afii9828isaninteger n,
h(t)/p58/afii9838(/afii9838t)/p76/p92/p16
(n/p571)!/afii9814/p76/p92/p16/p73/p14/p15(1/k!)(/afii9838t)/p73(6.4.8)
When /afii9828/p581,thedistributionisexponential.When /afii9838/p58/p16/p17and/afii9828/p58/p16/p17/afii9840,where /afii9840
isaninteger,thedistributionischi-squarewith /afii9840degreesoffreedom.Themean
andvarianceofthestandardgammadistributionare,respectively, /afii9828//afii9838and/afii9828//afii9838/p17,
sothatthecoefficientofvariationis1/ /p40/afii9828.
Manysurvivaldistributionscanberepresented,atleastroughly,bysuitable
choiceoftheparameters /afii9838and/afii9828.Inmanycases,thereisanadvantageinusing
theErlangiandistribution,thatis,intaking /afii9828integer.
Theexponential,Weibull,lognormal,andgammadistributionsarespecial
casesofageneralizedgammadistributionwiththreeparameters, /afii9838,/afii9825,and /afii9828,
whosedensityfunctionisdefinedas
f(t)/p58/afii9825/afii9838/p63/p65
/afii9772(/afii9828)t/p63/p65/p92/p16exp[ /p57(/afii9838t)/p63]t/p570,/afii9828/p570,/afii9838/p570,/afii9825/p570( 6.4.9)
It is easily seen that this generalized gamma distribution is the exponential
distribution if /afii9825/p58/afii9828/p581, the Weibull distribution if /afii9828/p581; the lognormal
distributionif /afii9828/p59/p45,andthegammadistributionif /afii9825/p581.
In later chapters (e.g., Chapters 7 and 9 ), we discuss several parametric
proceduresforestimationandhypothesistesting.TouseavailablecomputersoftwaresuchasSAStocarryoutthecomputation,weusethedistributionsadoptedbythesoftware.OneoftheveryfewsoftwarepackagesthatincludethegammaorgeneralizedgammadistributionisSAS.InSAS,thegeneralized152 -
Table 6.4 Lifetimes of 101 Strips of Aluminum Coupon
370
706
716746785797
844
855
858886886930960
988
990
1000101010161018
1020
10551085
1102
1102110811151120
1134
1140
11991200120012031222
1235
12381252125812621269
1270
12901293
1300
1310131313151330
1355
1390
14161419142014201450
1452
14751478148114851502
1505
15131522
1522
1530154015601567
1578
1594
16021604160816301642
1674
17301750175017631768
1781
17821792
1820
1868188118901893
1895
1910
19231940194520232100
2130
221522682440
Source:BirnbaumandSaunders (1958 ).
gammadistributionisdefinedashavingthefollowingdensityfunction:
f(t)/p58/p34/afii9825/p34/afii9828/p65/afii9838/p63/p65
/afii9772(/afii9828)t/p63/p65/p92/p16exp[ /p57/afii9828(/afii9838t)/p63],t/p570,/afii9828/p570,/afii9838/p570( 6.4.10)
To differentiatethisformof thegeneralizedgammadistributionfromthe
generalized gamma in (6.4.9 ), we refer to this distribution as the extended
generalized gamma distribution .Itcanbeshownthattheextendedgeneralized
gammadistributionreducestotheWeibulldistributionwhen /afii9825/p570and /afii9828/p581,
thelognormaldistributionwhen /afii9828/p59/p45,thegammadistributionwhen /afii9825/p581,
andtheexponentialdistributionwhen /afii9825/p58/afii9828/p581.
Example 6.4 BirnbaumandSaunders (1958 )reportanapplicationofthe
gammadistributiontothelifetimeofaluminumcoupon.Intheirstudy,17setsofsixstripswereplacedinaspeciallydesignedmachine.Periodicloadingwasappliedto the stripswith a frequencyof 18 hertzand a maximumstress of21,000poundspersquareinch.The102stripswererununtilallofthemfailed.One of the 102 strips tested had to be discarded for an extraneous reason,yielding101observations.ThelifetimedataaregiveninTable6.4inascendingorder. From the data the two parameters of the gamma distribution were 153
Figure 6.18Graphical comparison of observed and fitted cumulative distribution
functionsofdatainExample6.4. (FromBirnbaumandSaunders,1958. )
estimated (estimation methods are discussed in Chapter 7 ). They obtained
/afii9828/p24/p5811.8and /afii9838/p19/p581/(118.76/p5910/p18).
Agraphicalcomparisonoftheobservedandfittedcumulativedistribution
function is given in Figure 6.18, which shows very good agreement. Achi-squaregoodness-of-fittest (discussedinChapter9 )yieldeda /afii9851/p17valueof
4.49for6degreesoffreedom,correspondingtoasignificancelevelbetween0.5
and0.6.Thus,itwasconcludedthatthegammadistributionwasanadequatemodelforthelifelengthofsomematerials.
6.5 LOG-LOGISTIC DISTRIBUTION
The survival time Thas a log-logistic distribution if log( T) has a logistic
distribution.The density, survivorship,hazard, and cumulative hazard func-tionsofthelog-logisticdistributionare,respectively,
f(t)/p58/afii9825/afii9828t/p65/p92/p16
(1/p59/afii9825t/p65)/p17
(6.5.1)
S(t)/p581
1/p59/afii9825t/p65(6.5.2)154 -
h(t)/p58/afii9825/afii9828t/p65/p92/p16
1/p59/afii9825t/p65(6.5.3)
H(t)/p58log(1/p59/afii9825t/p65)( 6 .5.4)
t/p460,/afii9825/p570,/afii9828/p570
Thelog-logisticdistributionischaracterizedbytwoparameters /afii9825,and /afii9828.The
medianofthelog-logisticdistributionis /afii9825/p92/p16/p30/p65.Figure6.19 (a)to(c)showthe
log-logistichazard,density,andsurvivorshipfunctionswith /afii9825/p581andvarious
valuesof /afii9828/p582.0,1,and0.67.
When /afii9828/p571,thelog-logistichazardhasthevalue0attime0,increasestoa
peakat t/p58(/afii9828/p571)/p16/p30/p65//afii9825/p16/p30/p65,andthendeclines,whichissimilartothelognormal
hazard.When /afii9828/p581,thehazardstartsat /afii9825/p16/p30/p65andthendeclinesmonotonically.
When /afii9828/p581,thehazardstartsatinfinityandthendeclines,whichissimilarto
the Weibull distribution. The hazard function declines toward 0 as tap-
proachesinfinity.Thus,thelog-logisticdistributionmaybeusedtodescribeafirst increasing and then decreasing hazard or a monotonically decreasinghazard.
Example 6.5 Byers et al. (1988 )used the log-logistic distribution to
describetherateofspreadofHIVbetween1978and1986.Between1978and1980,over6700homosexualandbisexualmeninSanFranciscowereenrolledinstudiesoftheprevalenceandincidenceofsexuallytransmittedhepatitisBvirus (HBV )infections.Bloodspecimenswerecollectedfromtheparticipants.
Four hundred and eighty-eight men who were HBV-seronegative were ran-domly selected to participate in a study of HIV infection later. These menagreed to allow the investigators to test the specimens collected previouslytogether with a current specimen. For those who convert to positive, theinfectiontimeisonlyknowntohaveoccurredbetweenthepreviousnegativetestandthetimeofthefirstpositiveone.Theexacttimeisunknown.Thetimetoinfectionisthereforeintervalcensored.Theinvestigatorstriedtofitseveraldistributions to the interval-censored data, including the Weibull and log-logisticbymaximumlikelihoodmethods (discussedinChapter7 ).Basedon
the Akaike information criterion (discussed in Chapter 9 ), the log-logistic
distribution was found to provide the best fit to the data. The maximumlikelihoodestimatesofthetwoparametersare /afii9825/p24/p580.003757and /afii9828/p24/p581.424328.
Basedonthelog-logisticmodel,themedianinfectiontimeisestimatedtobe50.4months,andthehazardfunctionapproachesitspeakat27.6months.
6.6 OTHER SURVIVAL DISTRIBUTIONS
Manyotherdistributionscanbeusedasmodelsofsurvivaltime,threeofwhich
wediscussbrieflyinthissection:thelinearexponential,theGompertz (1825 ), 155
(a)
(b)
(c)
Figure 6.19(a) Hazardfunctionofthelog-logisticdistribution;( b) densityfunctionof
thelog-logisticdistribution;( c) Survivorshipfunctionofthelog-logisticdistribution.
156
Figure 6.20Hazardfunctionoflinear-exponentialmodel.andadistributionwhosehazardrateisastepfunction.Thelinear-exponential
model and the Gompertz distribution are extensions of the exponentialdistribution.Bothdescribesurvivalpatternsthathaveaconstantinitialhazardrate. The hazard rate varies as a linear function of time or age in thelinear-exponentialmodelandasanexponentialfunctionoftimeorageintheGompertzdistribution.
Indemonstratingtheuseofthelinear-exponentialmodel,Broadbent (1958 ),
uses as an example the service of milk bottles that are filled in a dairy,circulatedtocustomers,andreturnedemptytothedairy.ThemodelwasalsousedbyCarboneetal. (1967 )todescribethesurvivalpatternofpatientswith
plasmacyticmyeloma.Thehazardfunctionofthelinear-exponentialdistribu-tionis
h(t)/p58/afii9838/p59/afii9828t (6.6.1)
where /afii9838and/afii9828can be values such that h(t) is nonnegative.The hazardrate
increasesfrom /afii9838withtimeif /afii9828/p570,decreasesif /afii9828/p580,andremainsconstant (an
exponentialcase )if/afii9828/p580,asdepictedinFigure6.20.
Theprobabilitydensityfunctionandthesurvivorshipfunctionare,respec-
tively,
f(t)/p58(/afii9838/p59/afii9828t)exp[ /p57(/afii9838t/p59/p16/p17
/afii9828t/p17)] (6 .6.2)
and
S(t)/p58exp[ /p57(/afii9838t/p59/p16/p17/afii9828t/p17)] (6 .6.3)
Themeanofthelinear-exponentialdistributionis /p57(/afii9838//afii9828)/p59(/afii9828/2)/p92/p16/p30/p17L(/afii9838/p17/2/afii9828),
where
L(x)/p58e/p86/p16/p27
/p86y/p16/p30/p17e/p92/p87dy 157
Table 6.5 Values of L(x) and G(x)
xL (x) G(x)
00 .886 /p45
0.10 .951 2 .015
0.21 .012 1 .493
0.31 .067 1 .223
0.41 .119 1 .048
0.51 .168 0 .923
0.61 .214 0 .828
0.71 .258 0 .753
0.81 .300 0 .691
0.91 .341 0 .640
11 .381 0 .596
21 .712 0 .361
31 .987 0 .262
Source:Broadbent (1958 ).
Figure 6.21Gompertzhazardfunction.istabulatedinTable6.5.Aspecialcaseofthelinear-exponentialdistribution,
theRayleighdistribution,isobtainedbyreplacing /afii9828by/p16/p17/afii9828(Kodlin,1967 ).That
is,thehazardfunctionoftheRayleighdistributionis h(t)/p58/afii9838/p59/p16/p17/afii9828t.
TheGompertzdistributionisalsocharacterizedbytwoparameters, /afii9838and
/afii9828.Thehazardfunction,
h(t)/p58exp(/afii9838/p59/afii9828t)( 6 .6.4)
isplottedinFigure6.21.When /afii9828/p570,thereispositiveagingstartingfrom e/p72;
when /afii9828/p580,thereisnegativeaging;andwhen /afii9828/p580,h(t)reducestoaconstant,
e/p72.ThesurvivorshipfunctionoftheGompertzdistributionis
S(t)/p58exp/p3/p57e/p72
/afii9828(e/p65/p82/p571)/p4(6.6.5)158 -
Figure 6.22Stephazardfunction.andtheprobabilitydensityfunction,from (6.6.4 )and (2.2.5 ),isthen
f(t)/p58exp/p3(/afii9838/p59/afii9828t)/p571
/afii9828(e/p72/p62/p65/p82/p57e/p72)/p4(6.6.6)
ThemeanoftheGompertzdistributionis G(e/p72//afii9828)/e/p72, where
G(x)/p58e/p86/p16/p27
/p86y/p92/p16e/p92/p87dy
istabulatedinTable6.5.
Finally,weconsideradistributionwherethehazardrateisastepfunction:
h(t)/p58/p7a/p16
a/p17
/p36
a/p73/p92/p16
a/p730/p45t/p58t/p16
t/p16/p45t/p58t/p17
t/p73/p92/p17/p45t/p58t/p73/p92/p16
t/p46t/p73/p92/p16(6.6.7)
wheret/p16,t/p17,...,t/p73aredifferenttimepoints.Figure6.22showsatypicalhazard
functionofthisnaturefor k/p585.Using (2.2.4 ),thesurvivorshipfunctioncan
bederived:
S(t)/p58/p7exp(/p57a/p16t)0 /p45t/p58t/p16
exp[ /p57a/p16t/p16/p57a/p17(t/p57t/p16)] t/p16/p45t/p58t/p17
/p36
exp[ /p57a/p16t/p16/p57a/p17(t/p17/p57t/p16)/p57/p37/p57a/p73(t/p57t/p73/p92/p16)t/p46t/p73/p92/p16(6.6.8) 159
Theprobabilitydensityfunction f(t) canthen be obtained from (6.6.7 )and
(6.6.8 )using (2.2.5 ):
f(t)/p58/p7a/p16exp(/p57a/p16t)0 /p45t/p58t/p16
a/p17exp[ /p57a/p16t/p16/p57a/p17(t/p57t/p16)] t/p16/p45t/p58t/p17
/p36
a/p73exp[ /p57a/p16t/p16/p57a/p17(t/p17/p57t/p16)/p57/p37/p57a/p73(t/p57t/p73/p92/p16)t/p46t/p73/p92/p16(6.6.9)
One application of this distribution is the life-table analysis discussed in
Chapter4.Inalife-tableanalysis,timeisdividedintointervalsandthehazardrateisassumedtobeconstantineachinterval.However,theoverallhazardrateisnotnecessarilyconstant.
The nine distributions described above are, among others, reasonable
modelsforsurvivaltimedistribution.Allhavebeendesignedbyconsideringabiologicalfailure,adeathprocess,oranagingproperty.Theymayormaynotbe appropriate for many practical situations, but the objective here is toillustratethevariouspossibletechniques,assumptions,andargumentsthatcanbeusedtochoosethemostappropriatemodel.Ifnoneofthesedistributionsfitsthedata,investigatorsmighthavetoderiveanoriginalmodeltosuittheparticulardata,perhapsbyusingsomeoftheideaspresentedhere.
Bibliographical Remarks
Inadditiontothepapersonthedistributionscitedinthischapter,Mannetal.
(1974 ), Hahn and Shapiro (1967 ), Johnson and Kotz (1970a,b), Elandt-
JohnsonandJohnson (1980 ),Lawless (1982 ),Nelson (1982 ),CoxandOakes
(1984 ), Gertsbakh (1989 ), and Klein and Moeschberger (1997 )also discuss
statisticalfailuremodels,includingtheexponential,Weibull,gamma,lognor-mal,generalizedgamma,andlog-logisticdistributions.Applicationsofsurvivaldistributionscanbefoundeasilyinmedicalandepidemiologicaljournals.Thefollowing are a few examples: Dharmalingam et al. (2000 ), Riffenburgh and
Johnstone (2001 ),andMafartetal. (2002 ).
EXERCISES
6.1Summarize the distributions discussed in this chapter, answering the
followingquestions.
(a)Whatdistributionsdescribeconstanthazardrates?Givetherangeof
parametervalues.
(b)Whatdistributionsdescribeincreasinghazardrates?Iftherearemore
thanone,discussthedifferencesbetweenthem.
(c)Whatdistributionsdescribedecreasinghazardrates?Iftherearemore
thanone,discussthedifferencesbetweenthem.160 -
6.2Supposethatthesurvivaldistributionofagroupofpatientsfollowsthe
exponentialdistributionwith G/p580(year),/afii9838/p580.65. Plot the survivor-
shipfunctionandfind:
(a)Themeansurvivaltime
(b)Themediansurvivaltime
(c)Theprobabilityofsurviving1.5yearsormore
6.3Supposethatthesurvivaldistributionofagroupofpatientsfollowsthe
exponentialdistributionwith G/p585(years )and/afii9838/p580.25.Plotthesurviv-
orshipfunctionandfind:
(a)Themeansurvivaltime
(b)Themediansurvivaltime
(c)Theprobabilityofsurviving6yearsormore
6.4ConsiderthefollowingtwoWeibulldistributionsassurvivalmodels:
(i)G/p580,/afii9838/p581,/afii9828/p580.5
(ii)G/p580,/afii9838/p580.5,/afii9828/p582
For each distribution, plot the survivorship function and the hazard
functionandfind:
(a)Themean
(b)Thevariance
(c)Thecoefficientofvariation
Whichdistributiongivesthelargerprobabilityofsurvivingatleast3units
oftime?
6.5Suppose that the survival timefollows the lognormaldistribution with
/afii9839/p581and /afii9846/p580.5.Find:
(a)Themeansurvivaltime
(b)Thevariance
(c)Thecoefficientofvariation
(d)Themedian
(e)Themode
6.6Supposethatpainrelieftimefollowsthegammadistributionwith /afii9838/p581,
/afii9828/p580.5.Find:
(a)Themean
(b)Thevariance
(c)Thecoefficientofvariation
6.7Suppose that the survival distribution is (1)Gompertz and (2)linear-
exponential,and /afii9838/p581,/afii9828/p582.0.Plotthehazardfunctionandfind:
(a)Themean
(b)Theprobabilityofsurvivinglongerthan1unitoftime
6.8ConsiderthesurvivaltimesofhypernephromapatientsgiveninExercise
Table 3.1. From the plot you obtained in Exercise 4.5, suggest adistributionthatmightfitthedata. 161
CHAPTER 7
Estimation Procedures for
Parametric Survival Distributionswithout Covariates
In this chapter we discuss some analytical procedures for estimatingthe mostcommonlyusedsurvival distributionsdiscussed in Chapter6. We introducethemaximumlikelihoodestimates (MLEs )oftheparametersof thesedistributions.
The general asymptotic likelihood inference results that are most widely usedfor these distributions are given in Section 7.1. We begin to used the generalsymbol b/p58(b/p16,b/p17,...,b/p78) to denote a set of parameters. For example, in
discussingthe Weibull distribution, b/p16could be /afii9838andb/p17could be /afii9828, andp/p582.
bis called avectorin linear algebra. Readers who are not familiar with linear
algebra or are not interested in the mathematical details may skip this sectionand proceed to Section 7.2 without loss of continuity. In Sections 7.2 to 7.7 weintroducethe MLEs for the parameters of the exponential, Weibull, lognormal,gamma, log-logistic, and Gompertz distributions for data with and withoutcensored observations. The related BMDP or SAS programming codes thatmay be used to obtain the MLE are given in the respective sections.
7.1 GENERAL MAXIMUM LIKELIHOOD ESTIMATION
PROCEDURE
7.1.1 Estimation Procedures for Data with Right-Censored Observations
Suppose that persons were followed to death or censored in a study. Let t/p16,
t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76be the survival times observed from the nindividuals,
withrexact times and (n/p57r)right-censored times. Assume that the survival
times follow a distribution with the density function f(t,b)and survivorship
functionS(t,b), where b/p58(b/p16,...,b/p78) denotes unknown pparameters
b/p16,...,b/p78in the distribution. As shown in Chapter 6, an exponential distribu-
tion has one (p/p581)parameter /afii9838, the Weibull distribution has two (p/p582)
162
parameters /afii9838and /afii9828, and so on. If the survival time is discrete (i.e., it is observed
atdiscretetime only ),f(t,b)representstheprobabilityof observing tandS(t,b)
represents the probability that the survival or event time is greater than t.I n
other words, f(t,b)andS(t,b)represent the information that can be obtained
from an observed uncensored survival time and an observed right-censoredsurvival time, respectively. Therefore, the product /afii9811/p76/p71/p14/p16f(t/p71,b)represents
the joint probability of observingthe uncensored survival times, and/afii9811/p76/p71/p14/p80/p62/p16S(t/p62/p71,b)represents the joint probability of those right-censored survival
times. The product of these two probabilities, denoted by L(b),
L(b)/p58/p80/p147
/p71/p14/p16f(t/p71,b)/p76/p147
/p71/p14/p80/p62/p16S(t/p62/p71,b)
represents the joint probability of observing t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76. A similar
interpretation applies to continuous survival. L(b)is called the likelihood
functionofb, which can also be interpreted as a measure of the likelihood of
observinga specific set of survival times t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76, given a
specific set of parameters b. The method of the MLE is to find an estimator of
bthat maximizes L(b), or in other words, which is ‘‘most likely’’ to have
produced the observed data t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76. Take the logarithm of
L(b)and denote it by l(b),
l(b)/p58logL(b)/p58/p80/p26
/p71/p14/p16log[f(t/p71,b)]/p59/p76/p26
/p71/p14/p80/p62/p16log[S(t/p62/p71,b)] (7.1.1 )
Then the MLE b/p19ofbis the set ofb/p19/p16,...,b/p19/p78that maximizes l(b):
l(b/p19)/p58max
/p0/p12/p12
b(l(b)).
It is clear that b/p19is a solution of the followingsimultaneous equations, which
are obtained by takingthe derivative of l(b)with respect to each b/p72:
/p42l(b)
/p42b/p72/p580j/p581, 2,...,p (7.1.2)
The exact forms of (7.1.2 )for the parametric survival distributions discussed in
Chapter 6 are given in Sections 7.2 to 7.7. Often, there is no closed solution forthe MLE b/p19from (7.1.2 ). To obtain the MLE b/p19, one can use a numerical
method. A commonly used numerical method is the Newton —Raphson iter-
ative procedure, which can be summarized as follows.
1. Let the initial values of b/p16,...,b/p78be zero; that is, let
b/p7/p15/p8 /p580 163
2. The changes for bat each subsequentstep, denoted by /afii9773/p7/p72/p8, is obtainedby
takingthe second derivative of the log -likelihood function:
/afii9773/p7/p72/p8/p58/p3/p57/p42/p17l(b/p7/p72/p92/p16/p8)
/p42b/p42b/p30/p4/p92/p16/p42l(b/p7/p72/p92/p16/p8)
/p42b(7.1.3 )
3. Using /afii9773/p7/p72/p8, the value of b/p7/p72/p8atjth step is
b/p7/p72/p8/p58b/p7/p72/p92/p16/p8 /p59 /afii9773/p7/p72/p8j/p581, 2,...
The iteration terminatesat, say, the mth step if /p35/afii9773/p7/p75/p8/p35/p58/afii9829, where /afii9829is a given
precision, usually a very small value, 10 /p92/p19or 10 /p92/p20. Then the MLE b/p19is defined
as
b/p19/p58b/p7/p75/p92/p16/p8 (7.1.4 )
The estimated covariance matrix of the MLE b/p19is given by
V/p19(b/p19)/p58Co/p19v(b/p19)/p58/p3/p57/p42/p17l(b/p19)
/p42b/p42b/p30/p4/p92/p16(7.1.5 )
One of the good properties of a MLE is that if b/p19is the MLE of b, theng(b/p19)is
the MLE of g(b)ifg(b)is a finite function and need not be one-to-one. The
concept of the Newton —Raphson method for p/p581 is illustrated in detail in
Appendix A.
The estimated 100 (1/p57/afii9825)% confidence interval for any parameter b/p71is
(b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71,b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71)( 7.1.6 )
wherev/p71/p71is theith diagonal element of V/p19(b/p19)andZ/p63/p30/p17is the 100 (1/p57/afii9825/2)
percentile point of the standard normal distribution [ P(Z/p57Z/p63/p30/p17)/p58/afii9825/2]. For
a finite function g(b/p71)o fb/p71, the estimated 100 (1/p57/afii9825)% confidence interval for
g(b/p71) is its respective range Ron the confidence interval (7.1.6 ), that is,
R/p58/p43g(b/p71):b/p71/p43(b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71,b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71)/p44 (7.1.7 )
In caseg(b/p71) is monotone in b/p71, the estimated 100 (1/p57/afii9825)%confidence interval
forg(b/p71)i s
[g(b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71),g(b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71)] (7.1.8 )164
7.1.2 Estimation Procedures for Data with Right-, Left-, and
Interval-Censored Observations
If the survival times t/p16,t/p17,...,t/p76observed for the npersons consist of uncen-
sored, left-, right-, and interval-censored observations, the estimation pro-ceduresaresimilar.Assumethatthesurvivaltimesfollowadistributionwiththedensityfunction f(t,b)and thesurvivorshipfunction S(t,b), wherebdenotesall
unknown parameters of the distribution. Then the log-likelihood function is
l(b)/p58logL(b)/p58/p26log[f(t/p71,b)]/p59/p26log[S(t/p71,b)]
/p59/p26log[1 /p57S(t/p71,b)]/p59/p26log[S(v/p71,b)/p57S(t/p71,b)] (7.1.9 )
where the first sum is over the uncensored observations, the second sum over
the right-censored observations, the third sum over the left-censored observa-tions, and the last sum over the interval-censored observations, with v/p71as the
lower end of a censoringinterval. The other steps for obtainingthe MLE b/p19of
bare similar to the steps shown in Section 7.1.1 by substitutingthe log -
likelihood function defined in (7.1.1 )with the log-likelihood function in (7.1.9 ).
The computation of the MLE b/p19and its estimated covariance matrix is
tedious. The following example gives the general procedure for using SAS tocarry out the computation.
Example 7.1 If the survival time observed contains uncensored, right-,
left-, and interval-censored observations, one needs to create a new data setfrom the observed data to use SAS to obtain the estimates of the parametersin the distribution. For an observed survival time t(uncensored, right-, or
left-censored ),wedefinetwovariablesLBand UBas follows:If tis uncensored,
take LB /p58UB/p58t;i ftis left-censored, LB /p58. and UB /p58t; and iftis
right-censored, then LB /p58tand UB /p58., where ‘‘.’’ means ‘‘missing’’ in SAS. If
a survival time is interval-censored, [i.e., one observed two numbers t/p16andt/p17,
t/p16/p58t/p17and the survivaltime is in theinterval (t/p16,t/p17)], let LB /p58t/p16and UB /p58t/p17.
Assume that the new data set (in terms of LB and UB )has been saved in
‘‘C:/p33EXAMPLEA.DAT’’ as a text file, which contains two columns (LB in the
first column and UB the second column )separated by a space.
As an example, the followingSAS code can be used to obtain the estimated
covariance matrix defined in (7.1.5 )and the MLE of the parameters of the
Weibull distribution for the survival data observed in the text file ‘‘C: /p33EXAM-
PLEA.DAT’’. One can replace d /p58weibull in the followingcode with the
respective distribution in Sections 7.2 to 7.6 (see the SAS code in these sections
for details )to obtain the estimate.
data w1;
infile ‘c: /p33examplea.dat’ missover;
input lb ub;
run;
proc lifereg;
model (lb,ub )/p58/covb d /p58weibull;
run; 165
7.2 EXPONENTIAL DISTRIBUTION
7.2.1 One-Parameter Exponential Distribution
The one-parameterexponential distribution has the followingdensity function;
f(t)/p58/afii9838e/p92/p72/p82 (7.2.1)
survivorship function;
S(t)/p58e/p92/p72/p82 (7.2.2)
and hazard function;
h(t)/p58/afii9838 (7.2.3)
wheret/p460,/afii9838/p570. Obviously, the exponential distribution is characterized by
one parameter, /afii9838. The estimation of /afii9838by maximum likelihood methods for
data without censoredobservationswill be given first followedby the case withcensored observations.
Estimation of /afii9838for Data without Censored Observations
Supposethat thereare npersons in thestudy and everyoneis followedto death
or failure. Let t/p16,t/p17,...,t/p76be the exact survival times of the npeople. The
likelihood function, using (7.2.1 )and (7.1.1 ),i s
L/p58/p76/p147
/p71/p14/p16/afii9838e/p92/p72/p82/p71
and the log-likelihood function is
l(/afii9838)/p58nlog/afii9838/p57/afii9838/p76/p26
/p71/p14/p16t/p71(7.2.4)
From (7.1.2 ), the MLE of /afii9838is
/afii9838/p19/p58n
/afii9814/p76/p71/p14/p16t/p71 (7.2.5)
Since the mean /afii9839of the exponential distribution is 1/ /afii9838and a MLE is invariant
under an one-to-one transformation, the MLE of /afii9839is
/afii9839/p24/p581
/afii9838/p19/p58/afii9814/p76/p71/p14/p16t/p71n/p58t/p16 (7.2.6 )166
It can be shown 2 n/afii9839/p19//afii9839has an exact chi-square distribution with 2 ndegrees of
freedom (Epstein and Sobel, 1953 ). Since /afii9838/p581//afii9839and /afii9838/p19/p581//afii9839/p24, an exact
100(1/p57/afii9825)%confidence interval for /afii9838is
/afii9838/p19 /afii9851/p17/p17/p76/p11/p16/p92/p63/p30/p172n/p58/afii9838/p58/afii9838/p19 /afii9851/p17/p17/p76/p11/p63/p30/p172n(7.2.7)
where /afii9851/p17/p17/p76/p11/p63is the 100 /afii9825percentage point of the chi-square distribution with 2 n
degrees of freedom, that is, P(/afii9851/p17/p17/p76/p57/afii9851/p17/p17/p76/p11/p63)/p58/afii9825(Table B-2 ). Whennis large
(n/p4625, say ),/afii9838/p19is approximately normally distributed with mean /afii9838and
variance /afii9838/p17/n. Thus, an approximate 100 (1/p57/afii9825)%confidence interval for /afii9838is
/afii9838/p19/p57/afii9838/p19Z/p63/p30/p17
/p40n/p58/afii9838/p58/afii9838 /p19/p59/afii9838/p19Z/p63/p30/p17
/p40n(7.2.8)
whereZ/p63/p30/p17is the 100 /afii9825/2 percentage point, P(Z/p57Z/p63/p30/p17)/p58/afii9825/2, of the standard
normal distribution (Table B-1 ).
Since 2n/afii9839/p24//afii9839has an exact chi-square distribution with 2 ndegrees of freedom,
an exact 100 (1/p57a)% confidence interval for the mean survival time is
2n/afii9839/p24
/afii9851/p17/p17/p76/p11/p63/p30/p17/p58/afii9839/p582n/afii9839/p24
/afii9851/p17/p17/p76/p11/p16/p92/p63/p30/p17(7.2.9)
The followingexample illustrates the procedures.
Example 7.2 Consider the followingremission times in weeks for 21
patients with acute leukemia: 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 8, 8, 9, 10, 10, 12, 14, 16,20, 24, and 34. Assume that remission duration follows the exponentialdistribution. Let us estimate the parameter /afii9838by usingthe formulas g iven
above.
Accordingto (7.2.5 ), the MLE of the relapse rate, /afii9838,i s
/afii9838/p19/p5821
198
/p580.106 per week
The mean remission time /afii9839is then 198/21 /p589.429 weeks. Usingthe analytical
procedures given above, confidence intervals for /afii9838and /afii9839can also be obtained.
A 95% confidence interval for the relapse rate /afii9838, following (7.2.7 ),i s
approximately
(0.106 )(24.433 )
42/p58/afii9838/p58(0.106 )(59.342 )
42
or(0.062, 0.150 ). A 95% confidence interval for the mean remission time, 167
following (7.2.9 ),i s
(42)(9.429 )
59.342/p58/afii9839/p58(42)(9.429 )
24.433
or(6.673, 16.208 ).
Once the parameter /afii9838is estimated, other estimates can be obtained. For
example,the probability of stayingin remission for at least 20 weeks, estimatedfrom (7.2.2 ),i sS/p19(20)/p58exp[/p570.106 (20)]/p580.120. Any percentile of survival
timet/p78may be estimated by equating S(t)t opand solvingfor t/p78, that is,
t/p78/p58/p57logp//afii9838/p19. For example, the median (50th percentile )survival time can be
estimated by t/p15/p13/p20/p58/p57log0.5/ /afii9838/p19/p586.539 weeks.
Estimation of
/afii9838for Data with Censored Observations
We first consider singly censored and then progressively censored data.Suppose that without loss of generality, the study or experiment begins at time0 with a total of nsubjects. Survival times are recorded and the data become
available when the subjects die one after the other in such a way that theshortest survival time comes first, the second shortest second, and so on.Suppose that the investigator has decided to terminate the study after rout of
thensubjects have died and to sacrifice the remaining n/p57rsubjects at that
time. Then the survival times for the nsubjects are
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p58t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8
wherea superscriptplus indicates a sacrificedsubject,and thus t/p62/p7/p71/p8is a censored
observation. In this case, nandrare fixed values and all of the n/p57rcensored
observations are equal.
The likelihood function, using (7.1.1 ),(7.2.1 ), and (7.2.2 ),i s
L/p58n!
(n/p57r)!
/p80/p147
/p71/p14/p16/afii9838e/p92/p72/p82/p7/p71/p8(e/p92/p72/p82/p7/p80/p8)/p76/p92/p80
and from (7.1.2 ), the MLE of /afii9838is
/afii9838/p19/p58r
/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p7/p71/p8(7.2.10)
The mean survival time /afii9839/p581//afii9838can then be estimated by
/afii9839/p24/p581
/afii9838/p19/p58/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p7/p71/p8r(7.2.11 )
It is shown by Halperin (1952 )that 2r/afii9838//afii9838/p19has a chi-square distribution with 2 r168
degrees of freedom. The mean and variance of /afii9838/p19arer/afii9838/(r/p571) and /afii9838/p17/(r/p571),
respectively. The 100 (1/p57/afii9825)% confidence interval for /afii9838is
/afii9838/p19/afii9851/p17/p17/p80/p11/p16/p92/p63/p30/p172r/p58/afii9838/p58/afii9838/p19 /afii9851/p17/p17/p80/p11/p63/p30/p172r(7.2.12)
Whennis large, the distributionof /afii9838/p19is approximatelynormal with mean /afii9838and
variance /afii9838/p17/(r/p571). An approximate 100/ (1/p57/afii9825)% confidence interval for /afii9838is
then
/afii9838/p19/p57/afii9838/p19Z/p63/p30/p17
/p40r/p571/p58/afii9838/p58/afii9838 /p19/p59/afii9838/p19Z/p63/p30/p17
/p40r/p571(7.2.13)
Epstein and Sobel (1953 )show that 2r/afii9839/p24//afii9839has a chi-square distribution with
2rdegrees of freedom. Thus a 100/ (1/p57/afii9825)%confidence interval for /afii9839(see also
Epstein, 1960b )is
2r/afii9839/p24
/afii9851/p17/p17/p80/p11/p63/p30/p17/p58/afii9839/p582r/afii9839/p24
/afii9851/p17/p17/p80/p11/p16/p92/p63/p30/p17(7.2.14)
They also develop test procedures for the hypothesis H/p15:/afii9839/p58/afii9839/p15against the
alternativeH/p16:/afii9839/p58/afii9839/p15. One of their rules of action is to accept H/p15if/afii9839/p24/p57cand
rejectH/p15if/afii9839/p24/p58c, wherec/p58(/afii9839/p15/afii9851/p17/p17/p80/p11/p63)/2rand /afii9825is the significancelevel. Or if the
estimated mean survival time calculated from (7.2.11 )is greater then c, the
hypothesisH/p15is rejected at the /afii9825level. The followingexample illustrates the
procedure.
Example 7.3 Suppose that in a laboratory experiment 10 mice are exposed
to carcinogens. The experimenter decides to terminate the study after half ofthemice are dead and to sacrifice the other halfat that time. The survival timesof the five dead mice are 4, 5, 8, 9, and 10 weeks. The survival data of the 10mice are 4, 5, 8, 9, 10, 10 /p59,1 0/p59,1 0/p59,1 0/p59, and 10 /p59. Assumingthat the
failureof these mice follows an exponentialdistribution, the survival rate /afii9838and
mean survival time /afii9839are estimated, respectively, accordingto (7.2.10 )and
(7.2.11 )by
/afii9838/p585
36/p5950
/p580.058 per week
and /afii9839/p24/p581/0.058 /p5817.241 weeks. A 95%confidence interval for /afii9838by(7.2.12 )is
(0.058 )(3.247 )
(2)(5)/p58/afii9838/p58(0.058 )(20.483 )
(2)(5) 169
or(0.019, 0.119 ). A 95% confidence interval for /afii9839following (7.2.13 )is
2(5)(17.241 )
20.483/p58/afii9839/p582(5)(17.241 )
3.247
or(8.417, 53.098 ).
The probability of survivinga g iven time for the mice can be estimated from
(7.2.2 ). For example, the probability that a mouse exposed to the same
carcinogen will survive longer than 8 weeks is
S/p19(8)/p58exp[/p570.058 (8)]/p580.629
The probability of dyingin 8 weeks is then 1 /p570.629 /p580.371.
A slightly different situation may arise in laboratory experiments. Instead of
terminatingthe study after the rth death, the experimenter may stop after a
period of time T, which may be six months or a year. If we denote the number
of deaths between 0 and Tasr, the survival data may look as follows:
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p45t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8/p58T
Mathematical derivations of the MLE of /afii9838and /afii9839are exactly the same and
(7.2.10 )can still be used. The samplingdistribution of /afii9839/p24for singly censored
data is also discussed by Bartholomew (1963 ).
Progressively censored data come more frequently from clinical studies
where patients are entered at different times and the study lasts a predeter-mined period of time. Suppose that the study begins at time 0 and terminatesat timeTand there are a total of npeople entered. Let rbe the number of
patients who die before or at time Tandn/p57rthe number of patients who are
lost to follow-up duringthe study period or remain alive at time T. The data
look as follows: t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76. Orderingthe runcensored
observations accordingto their mag nitude, we have
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p80,t/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76
The likelihood function, using (7.1.1 ),(7.2.1 ), and (7.2.2 ),i s
L/p58/p80/p147
/p71/p14/p16/afii9838e/p92/p72/p82
/p7/p71/p8/p76/p147
/p71/p14/p80/p62/p16e/p92/p72/p82/p71/p62
and the log-likelihood function is
l(/afii9838)/p58n/afii9838/p57/afii9838/p80/p26
/p71/p14/p16t/p71/p57/afii9838/p76/p26
/p71/p14/p80/p62/p16t/p62/p71(7.2.15)170
and from (7.1.2 ), the MLE of the parameter /afii9838is
/afii9838/p19/p58r
/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p71(7.2.16)
Consequently,
/afii9839/p24/p581
/afii9838/p19/p58/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p71r(7.2.17)
is the MLE of the mean survival time. The sum of all of the observations,
censored and uncensored, divided by the number of uncensored observations,gives the MLE of the mean survival time. To overcome the mathematicaldifficulties arisingwhen all of the observations are censored (r/p580), Bar-
tholomew (1957 )defines
/afii9839/p24/p58/p76/p26
/p71/p14/p16t/p62/p71(7.2.18)
In practice, this estimate has little value.
Distributions of the estimators are discussed by Bartholomew (1957 ). The
distributionof /afii9838/p19forlargenis approximatelynormalwith mean /afii9838and variance:
Var(/afii9838/p19)/p58/afii9838/p17
/p76/p26
/p71/p14/p16(1/p57e/p92/p72/p50
/p71)(7.2.19)
whereT/p71is the time that the ith person is under observation. In other words,
T/p71is computed from the time the ith person enters the study to the end of the
study. If the observation times T/p71are not known, the followingquick estimate
of Var (/afii9838/p19)can be used:
Var/p19(/afii9838/p19)/p58/afii9838/p19/p17
r(7.2.20 )
Thus an approximate 100 (1/p57/afii9838)% confidence interval for /afii9838is, by (7.1.6 ),
/afii9838/p19/p57Z/p63/p30/p17/p40Var/p19(/afii9838/p19)/p58/afii9838/p58/afii9838 /p19/p59Z/p63/p30/p17/p40Var/p19(/afii9838/p19) (7.2.21 )
The distribution of /afii9839/p24is approximately normal with mean /afii9839and variance:
Var(/afii9839/p24)/p58/afii9839/p17
/afii9814/p76/p71/p14/p16(1/p57e/p92/p72/p50/p71)(7.2.22) 171
Again, a quick estimate is
Var/p19(/afii9839/p24)/p58/afii9839/p24/p17
r(7.2.23)
An approximate 100 (1/p57/afii9825)% confidence interval for /afii9839is then, by (7.1.6 ),
/afii9839/p24/p57Z/p63/p30/p17/p40Var/p19(/afii9839/p24)/p58/afii9839/p58/afii9839 /p24/p59Z/p63/p30/p17/p40Var/p19(/afii9839/p24) (7.2.24 )
The exact distribution of /afii9839/p24derived by Bartholomew (1963 )is too cumbersome
for general use and thus is not included here.
Example 7.4 Consider the remission duration of the 21 leukemia patients
receiving6-MP in Example 3.3. The remission times in weeks were 6, 6, 6, 7,10, 13, 16, 22, 23, 6 /p59,9/p59,1 0/p59,1 1/p59,1 7/p59,1 9/p59,2 0/p59,2 5/p59,3 2/p59,3 2/p59,3 4/p59,
and 35 /p59. The hazard plot given in Figure 3.6 shows that the exponential
distribution fits the data very well. Maximum likelihood estimates of therelapse rate and the mean remission time can be obtained, respectively, from(7.2.16 )and (7.2.17 ):
/afii9838/p19/p589
109/p59250
/p580.025 per week /afii9839/p24/p581
0.025/p5840 weeks
The graphical estimate of /afii9838obtained in Example 3.3 is 0.027, which is very
close to the MLE. Thus, the remission duration of leukemia patients receiving6-MP can be described by an exponential distribution with a constant weeklyrelapse rate of 2.5% and a mean remission time of 40 weeks. The probabilityof stayingin remission for one year (or 52 weeks )or more is estimated by
S/p19(52)/p58exp[/p570.025 (52)]/p580.273
Using (7.2.20 )and (7.2.23 )for the variance of /afii9838/p19and /afii9839/p24, the 95% confidence
intervals for /afii9838and /afii9839are, respectively, (0.009, 0.041 )and (13.867, 66.133 ).
Example 7.5 The results in Examples 7.2 to 7.4 can also be obtained by
usingavailable statistical software. Let tdenote the observed survival time
(exact or censored )and CENS be an index (or dummy )variable with
CENS /p580i ftis censored and 1 otherwise. Assume that the data have been
saved in ‘‘C: /p33EXAMPLE.DAT’’ as a text file, which contains two columns (t
in the first column and CENS in second column for the same study subject ),
separated by a space.
The followingSAS code for procedure LIFEREG can be used to obtain
the estimated covariance matrix defined in (7.1.5 )and the MLE of the
parameter of the exponential distribution for the observed survival data in‘‘C:/p33EXAMPLE.DAT’’.172
data w1;
infile ‘c: /p33example.dat’ missover;
input t cens;
run;
proc lifereg;
model t*cens (0)/p58/covb d /p58exponential;
run;
The respective BMDP code for program 2L is
/input file /p58‘c:/p33example.dat’ .
variables /p582.
format /p58free.
/print level /p58brief.
cova. survival.
/variable names /p58t, cens.
/form time /p58t.
status /p58cens.
response /p581.
/regress accel /p58exponential.
/end
If SAS is used, the estimated parameter of the exponential distribution can
be obtained by
/afii9838/p19/p58exp(/p57INTERCEPT ),
where INTERCEPT is the name of output estimated parameter in SAS
procedure LIFEREG. In BMDP 2L,
/afii9838/p19/p58exp(/p57CONSTANT )
where CONSTANT is given by the program.
7.2.2 Two-Parameter Exponential Distribution
In the case where a two-parameter exponential distribution is more appropri-
ate for the data (Zelen, 1966 ), the density and survivorship functions are
defined, respectively, as
f(t)/p58/p7/afii9838e/p92/p72/p7/p82/p92/p37/p8
0t/p46G/p460, /afii9838/p460
t/p58G(7.2.25 )
and
S(t)/p58/p7e/p92/p72/p7/p82/p92/p37/p8
1t/p46G/p460, /afii9838/p460
t/p58G(7.2.26)
whereGis called the guaranteetime , the minimum survival time before which
no deaths occur. 173
Estimation of /afii9838and G for Data without Censored Observations
Ift/p16,t/p17,...,t/p76are the survival times of the npatients, using (7.1.1 ),(7.1.2 ),
(7.2.25 ), and (7.2.26 ), the MLE of /afii9838is
/afii9838/p19/p58n
/p76/p26
/p71/p14/p16(t/p71/p57G/p19)(7.2.27)
whereG/p19is an estimate of Gthat is the smallest observation in the data,
G/p19/p58min(t/p16,t/p17,...,t/p76)( 7 .2.28)
and the mean survival time is estimated by /afii9839/p24/p58G/p19/p591//afii9838/p19.
Example 7.6 Consider the survival times in months of 11 patients following
initial pulmonarymetastasis from ostenogenicsarcoma consideredby Burdetteand Gehan (1970 ). The data were 11, 13, 13, 13, 13, 13, 14, 14, 15, 15, and 17.
Suppose that the two-parameter exponential distribution is selected. Theguarantee time Gis estimated by the smallest observation (i.e.,G/p19/p5811), and
the hazard rate /afii9838/p19estimated by (7.2.27 )is
/afii9838/p19/p5811
(11/p5711)/p59(13/p5711)/p59/p37/p59(17/p5711)
/p580.367
Thus, the exponential model tells us that the minimum survival time is 11
months, and after that the chance of death per month is 0.367. Similarly, theprobability of survivinga g iven amount of time can then be estimated from(7.2.26 ). For example, the estimated probability of surviving18 months or
longer is
S/p19(18)/p58exp[/p570.367 (18/p5711)]/p580.077
Estimation of
/afii9838and G for Data with Censored Observations
We first consider singly censored data. Suppose that an experiment begins withnanimals and terminates as soon as the first rdeaths occur. For this case, we
introduce the estimation procedures derived by Epstein (1960a ).
Let the first rsurvival times be t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8and letT*be the total
survival observed between the first and the rth death:
T*/p58(n/p571)(t/p7/p17/p8/p57t/p7/p16/p8)/p59(n/p572)(t/p7/p18/p8/p57t/p7/p17/p8)/p59/p37/p59(n/p57r/p591)(t/p7/p80/p8/p57t/p7/p80/p92/p16/p8)
/p58/p57(n/p571)t/p7/p16/p8/p59t/p7/p17/p8/p59t/p7/p18/p8/p59/p37/p59t/p7/p80/p92/p16/p8/p59(n/p57r/p591)t/p7/p80/p8
/p58/p80/p26
/p71/p14/p16t/p7/p71/p8/p57nt/p7/p16/p8/p59(n/p57r)t/p7/p80/p8(7.2.29 )174
The best estimates for Gand /afii9839in the sense that they are unbiased and have
minimum variance are given by
G/p19/p58t/p7/p16/p8/p57/afii9839/p24
n(7.2.30)
and
/afii9839/p24/p58T*
r/p571(7.2.31)
Then /afii9838can then be estimated by /afii9838/p19/p581//afii9839/p24.
Confidence intervals for the mean survival time /afii9839are easy to obtain from
the fact that 2 (r/p571)/afii9839/p24//afii9839/p582T*//afii9839has a chi-square distribution with 2( r/p571)
degreesoffreedom.Thus,for r/p571,the100 (1/p57/afii9825)%confidenceintervalfor /afii9839is
2(r/p571)/afii9839/p24
/afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p63/p30/p17/p58/afii9839/p582(r/p571)/afii9839/p24
/afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p16/p92 /p63/p30/p17(7.2.32)
or
2T*
/afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p63/p30/p17/p58/afii9839/p582T*
/afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p16/p92 /p63/p30/p17(7.2.33)
To find confidence intervals for G, we use the fact that x/p16/p582n(t/p7/p16/p8/p57G)//afii9839
andx/p17/p582(r/p571)/afii9839/p24//afii9839are independent and have a chi-square distribution with
2 and 2 (r/p571) degrees of freedom, respectively. Thus the ratio
Y/p58x/p16/2
x/p17/2(r/p571)/p58n(t/p7/p16/p8/p57G)
/afii9839/p24/p58n(r/p571)(t/p7/p16/p8/p57G)
T*(7.2.34)
follows the F-distribution with 2 and 2( r/p571) degrees of freedom. Let
F/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63be the 100 /afii9825percentage point of the F/p17/p11/p17/p7/p80/p92/p16/p8distribution [i.e.,
P(Y/p46F/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63)/p58/afii9825](Table B-3 in Appendix B ), and then a 100 (1/p57/afii9825)%
confidence interval for Gis
t/p7/p16/p8/p57/afii9839/p24
nF/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63/p58G/p58t/p7/p16/p8(7.2.35 )
or
t/p7/p16/p8/p57T*
n(r/p571)F/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63/p58G/p58t/p7/p16/p8(7.2.36)
Epstein and Sobel (1953 )show that this interval is the shortest in the class of
intervalsbeingused. If for someparticularvaluesof rand /afii9825thevalueF/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63is not tabulated in the F-table, Epstein (1960a )suggests using the following 175
confidence intervals for G:
t/p7/p16/p8/p57/afii9839/p24(r/p571)
ng/p16/p92/p63/p58G/p58t/p7/p16/p8(7.2.37)
or
t/p7/p16/p8/p57T*
ng/p16/p92/p63/p58G/p58t/p7/p16/p8(7.2.38)
where
g/p16/p92/p63/p58/p11
/afii9825/p2/p16/p30/p7/p80/p92/p16/p8/p571( 7 .2.39)
is computable for any /afii9825andr. Example 7.7 illustrates the procedures.
Example 7.7 In a laboratoryexperiment20 mice areinjectedwith a tumor
inoculum. These tumor cells multiply and eventually kill the animal. Supposethat the investigator decides to terminate the experiment after 10 deaths. Thefirst occurs 30 days after the experiment starts. The total survival observedbetween the time when the first and tenth deaths occur is 600 animal days.Assumingthat the survival distribution of these mice is exponential, theshortest 95% confidence interval for Gcan be obtained by (7.2.36 ). Since
F/p17/p11/p16/p23/p11/p15/p13/p15/p20/p583.555, the interval is
30/p57600
(20)(9)
(3.555 )/p58G/p5830
or(18.150, 30 ).
The mean survival time estimated by (7.2.31 )is/afii9839/p24/p5866.667 days, and the
95%confidence interval for /afii9839computed from (7.2.33 )is
2(600)
31.526/p58/afii9839/p582(600)
8.231
or(38.064, 145.790 ).
When data are progressively censored, Gehan (1970 )derives an estimate for
Gand a modified MLE for the hazard rate /afii9838. Suppose that rout of then
individuals in the study die before the end of the study and n/p57rindividuals
are alive at the time of the last follow-up or termination. The nsurvival times
are denoted by
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8,t/p62/p7/p80/p62/p16/p8,...,t/p62/p7/p76/p8
An estimate of Gobtained by
G/p19/p58max/p1t/p7/p16/p8/p571
n/afii9838/p19,0/p2(7.2.40)176
and the variance of G/p19is
Var(G/p19)/p581
(n/afii9838/p19)/p17/p11/p591
r/p571/p2(7.2.41)
Whennis large,Gand Var (G/p19)can be estimated by
G/p19/p60t/p7/p16/p8(7.2.42)
and
Var/p19(G/p19)/p601
(n/afii9838/p19)/p17(7.2.43)
A modified MLE for /afii9838is
/afii9838/p19/p58r/p571
/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p7/p71/p8/p57nt/p7/p16/p8(7.2.44 )
with variance
Var(/afii9838/p19)/p58/afii9838/p17
r/p571(7.2.45)
Any percentile of survival time t/p78may be estimated by equating S(t)t opand
solvingfort/p19/p78; that is,t/p19/p78/p58/p57 (log/p67p)//afii9838/p19/p59G/p19.
The followingexample illustrates the procedures.
Example 7.8 Suppose that 19 patients with brain tumor are followed in a
clinical trial for a year. Their survival times in weeks are 3, 4, 6, 8, 8, 10, 12,16,17,30,33,3 /p59,8/p59,13/p59,21/p59,26/p59,3 5/p59,44/p59,and 45 /p59.Inthis case n/p5819,
r/p5811,t/p7/p16/p8/p583,/afii9814/p16/p16/p71/p14/p16t/p7/p71/p8/p58147, and /afii9814/p16/p24/p71/p14/p16/p17t/p62/p7/p71/p8/p58195. The hazard rate /afii9838per week
and its variance may be estimated by (7.2.44 )and (7.2.45 )as
/afii9838/p19/p5810
147/p59195/p5719(3)/p580.035
and
Var/p19(/afii9838/p19)/p58(0.035 )/p17
10/p580.0001
The guarantee time Gand its variance may then be estimated by (7.2.40 )and
(7.2.41 ): 177
G/p19/p58max /p13/p571
19/p590.035,0/p2/p581.496
and
Var/p19(G/p19)/p581
(19/p590.035)/p17/p11/p591
10/p2/p582.487
Thus, after a guarantee time of approximately 1.5 weeks, the chance of death
per week is 0.035. The estimated median survival time is
t/p19/p15/p13/p20/p58/p57log0.5
0.035/p591.496 /p5821.3 weeks
The probability of survivingat least six months (or 26 weeks )is estimated by
S/p19(26)/p58exp[/p570.35(26/p571.496 )]/p580.424
7.3 WEIBULL DISTRIBUTION
The Weibull distribution has the density and survivorship functions
f(t)/p58/afii9828/afii9838/p65t/p65/p92/p16exp[/p57(/afii9838t)/p65]
S(t)/p58e/p92/p7/p72/p82/p8/p65t/p460, /afii9828/p570, /afii9838/p570 (7.3.1 )
The MLE of the parameters /afii9838and /afii9828involves equations to be solved simulta-
neously. Numerical methods such as the Newton —Raphson iterative procedure
(7.1.13 )can be applied.We begin with the case whereno censoredobservations
are presented.
Lett/p16,t/p17,...,t/p76be the exact survival times of nindividuals under investiga-
tion. If their survival times follow the Weibull distribution, the log-likelihoodfunction is
l(/afii9838,/afii9828)/p58nlog/afii9828/p59n/afii9828log/afii9838/p59/p76/p26
/p71/p14/p16[(/afii9828/p571)logt/p71/p57/afii9838/p65t/p65/p71]( 7.3.2)
The MLE of /afii9838and /afii9828in(7.3.1 )can be obtained by solvingthe followingtwo
equations simultaneously:
n/p57/afii9838/p19
/afii9828/p24/p76/p26
/p71/p14/p16t/p71/afii9828/p24/p580 (7.3.3 )
n
/afii9828/p24/p59nlog/afii9838/p19/p59/p76/p26
/p71/p14/p16logt/p71/p57/afii9838/p19/afii9828/p24/p76/p26
/p71/p14/p16t/p71/afii9828/p24(log/afii9838/p19/p59logt/p71)/p580( 7.3.4)178
Next, let us consider a typical laboratory experiment in which subjects are
entered at the same time and the experiment is terminated after rof then
subjects have failed (or after a fixed period of time T). In both of these cases
the data collected are singly censored. The ordered survival data are
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p58t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8
If the time to failure follows the Weibull distribution with the density function
given in (7.3.1 ), the MLE of /afii9838and /afii9828may be obtained by solvingthe following
two equations simultaneously:
r/p57/afii9838/p19/afii9828/p24/p3/p80/p26
/p71/p14/p16t/p71/afii9828/p24/p59(n/p57r)t/afii9828/p24
/p7/p80/p8/p4/p580 (7.3.5 )
r
/afii9828/p24/p59rlog/afii9838/p19/p59/p80/p26
/p71/p14/p16logt/p71
/p59/afii9838/p19/afii9828/p24/p3/p80/p26
/p71/p14/p16t/p71/afii9828/p24(log/afii9838/p19/p59logt/p71)/p59(n/p57r)t/afii9828/p24
/p7/p80/p8(log/afii9838/p19/p59logt/p7/p80/p8)/p4/p580 (7.3.6 )
When data are progressively censored, we have
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8,t/p62/p7/p80/p62/p16/p8,...,t/p62/p7/p76/p8
If the survival distribution is Weibull defined by (7.3.1 ), the log-likelihood
function is
l(/afii9838,/afii9828)/p58rlog/afii9828/p59r/afii9828log/afii9838/p59/p80/p26
/p71/p14/p16[(/afii9828/p571)logt/p7/p71/p8/p57/afii9838/p65t/p65/p7/p71/p8]/p57/p76/p26
/p71/p14/p80/p62/p16/afii9838/p65t/p62/p65/p7/p71/p8
(7.3.7)
The MLE of /afii9838and /afii9828may be obtained by solvingthe followingtwo equations
simultaneously:
r/p57/afii9838/p19/p65/p24/p1/p80/p26
/p71/p14/p16t/p71/afii9828/p24/p59/p76/p26
/p71/p14/p80/p62/p16t/p62/p71/afii9828/p24/p2/p580( 7.3.8)
r
/afii9828/p24/p59rlog/afii9838/p19/p59/p80/p26
/p71/p14/p16logt/p71/p57/afii9838/p19/afii9828/p24/p80/p26
/p71/p14/p16t/p71/afii9828/p24(log/afii9838/p19/p59logt/p71)
/p57/afii9838/p19/afii9828/p24/p76/p26
/p71/p14/p80/p62/p16t/p71/p59/afii9828/p24(log/afii9838/p19/p59logt/p62/p71)/p580( 7.3.9)
The followingexample illustrates the use of available computer software to
obtain the MLE of /afii9838and /afii9828. 179
Example 7.9 Referringto Example 7.5, for the observed survival data in
the file ‘‘EXAMPLE.DAT’’, we can use either SAS or BMDP to obtain theestimated parameters of the Weibull distribution. The codes given in Example7.5 can be used except that d /p58exponential in the SAS code must be changed
to d /p58weibull and accel /p58exponential in BMDP code be changed to ac-
cel/p58weibull. If SAS is used, the estimated parameters of the Weibull distribu-
tion are
/afii9838/p19/p58exp(/p57INTERCEPT )and /afii9828/p24/p581
SCALE
where INTERCEPT and SCALE are produced by SAS procedure LIFEREG.
If BMDP is used,
/afii9838/p19/p58exp(/p57INTERCEPT )and /afii9828/p24/p581
SCALE
where CONSTANT and SCALE are given by procedure 2L.
7.4 LOGNORMAL DISTRIBUTION
If the survival time Tfollows the lognormal distribution with density function
f(t)/p581
t/afii9846/p402/afii9843exp/p3/p571
2/afii9846/p17(logt/p57/afii9839)/p17/p4(7.4.1)
the mean and the variance are exp (/afii9839/p59/p16/p17/afii9846/p17)and [exp (/afii9846/p17)/p571]exp (2/afii9839/p59/afii9846/p17),
respectively. Estimation of the two parameters /afii9839and /afii9846/p17has been investigated
either by using (7.4.1 )directly or by usingthe fact that Y/p58logTfollows the
normal distribution with mean /afii9839and variance /afii9846/p17. In the following, we discuss
theestimation of /afii9839and /afii9846/p17forsampleswith andwithout censoredobservations.
7.4.1 Estimation of /afii9839and/afii98462for Data without Censored Observations
Estimationsof /afii9839and /afii9846/p17for completesamples by maximumlikelihood methods
havebeen studied by many authors:forexample, Cohen (1951 )and Harterand
Moore (1966 ). But the simplest way to obtain estimates of /afii9839and /afii9846/p17with
optimum properties is by consideringthe distribution of Y/p58logT. Lett/p16,
t/p17,...,t/p76be the survival times of nsubjects. The MLE of /afii9839is the sample mean
ofYgiven by
/afii9839/p24/p581
n/p76/p26
/p71/p14/p16logt/p71(7.4.2)
The MLE of /afii9846/p17is
/afii9846/p24/p17/p581
n/p3/p76/p26
/p71/p14/p16(logt/p71)/p17/p57(/afii9814/p76/p71/p14/p16logt/p71)/p17
n /p4(7.4.3)180
The estimate /afii9839/p24is also unbiased but /afii9846/p24/p17is not. The best unbiased estimates of
/afii9839and /afii9846/p17are /afii9839/p24and the sample variance s/p17/p58 /afii9846/p24/p17[n/(n/p571)]. Ifnis moderately
large, the difference between s/p17and /afii9846/p24/p17is negligible.
One of the properties of the MLE is that if /afii9835/p21/p19is the MLE of /afii9835/p21,g(/afii9835/p21/p19)is the
MLE ofg(/afii9835/p21)ifg(/afii9835/p21)is a finite function. Therefore, the MLEs of the mean and
variance ofTare, respectively, exp (/afii9839/p24/p59/p16/p17/afii9846/p24/p17)and exp[ (/afii9846/p24/p17/p571)] exp (2/afii9839/p24/p59/afii9846/p24/p17).
It is known that /afii9839/p24/p58y/p21is normally distributed with mean /afii9839and variance
/afii9846/p17/n. Hence, if /afii9846is known, a 100 (1/p57/afii9825)% confidence interval for /afii9839is
/afii9839/p24/p60Z/p63/p30/p17/afii9846//p40n.I f /afii9846is unknown, we can use Student’s t-distribution. A
100(1/p57/afii9825)%confidence interval for /afii9839is/afii9839/p24/p60t/p63/p30/p17/p11/p7/p76/p92/p16/p8s//p40n/p571, wheret/p63/p30/p17/p11/p7/p76/p92/p16/p8is the 100 /afii9825/2 percentage point of Student’s t-distribution with n/p571 degrees of
freedom (Table B-7 ).
Confidence intervals for /afii9846/p17can be obtained by usingthe fact that n/afii9846/p24//afii9846/p17
has a chi-square distribution with n/p571 degrees of freedom. A 100 (1/p57/afii9825)%
confidence interval for /afii9846/p17is
n/afii9846/p24/p17
/afii9851/p17/p7/p76/p92/p16/p8/p11/p63/p30/p17/p58/afii9846/p17/p58n/afii9846/p24/p17
/afii9851/p17/p7/p76/p92/p16/p8/p11/p16/p92 /p63/p30/p17(7.4.4)
The followinghypothetical example illustrates the procedures.
Example 7.10 Fivemelanoma (resected )patientsreceivingimmunotherapy
BCG are followed. The remissionduration in weeks are, in order of magnitude,8, 16, 23, 27, and 28. Suppose that the remission times follow a lognormaldistribution. In this case, parameters are estimated by (7.4.2 )and (7.4.3 )as
follows:
t logt (logt)/p17
82 .079 4 .322
16 2 .773 7 .690
23 3 .135 9 .828
27 3 .296 10 .864
28 3 .332 11 .102——— ———14.615 43 .806
/afii9839/p24/p5814.615
5/p582.923
/afii9846/p24/p17/p581
5/p343.806 /p571
5(14.615 )/p17/p4/p580.217
s/p17/p585/afii9846/p24/p17
5/p571/p580.271 181
The mean remission time is exp (2.923 /p590.217/2 ), or 20.728, weeks and the
standard deviation of the remission time is /p43[exp (0.217 )/p571]
exp(5.846 /p590.217 )/p44/p16/p30/p17, or 10.204, weeks. A 95% confidence interval for /afii9839is
2.923 /p572.776 /p10.521
/p404/p2/p58/afii9839/p582.923 /p592.776 /p10.521
/p404/p2
or(2.200,3.646 ). A 95% confidence interval for /afii9846/p17, following (7.4.4 ),i s
5(0.217 )
11.1433/p58/afii9846/p17/p585(0.217 )
0.4844
or(0.097, 2.240 ).
7.4.2 Estimation of /afii9839and/afii98462for Data with Censored Observations
We first consider samples with singly censored observations. The data consist
ofrexact survival times t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8andn/p57rright-censored survival
times that are at least t/p7/p80/p8, denoted by t/p62/p7/p80/p8. Again, we use the fact that Y/p58logT
has normal distribution with mean /afii9839and variance /afii9846/p17. Estimates of /afii9839and /afii9846/p17
can be obtained from the transformed data y/p71/p58logt/p71. Many authors have
investigatedthe estimation of /afii9839and /afii9846/p17: for example, Gupta (1952 ), Sarhan and
Greenberg (1956, 1957, 1958, 1962 ), Saw (1959 ), and Cohen (1959, 1961 ).W e
shall discuss the methods of Sarhan and Greenbergand Cohen because of theavailable table that reduces computation time and efforts.
The best linear estimates of /afii9839and /afii9846proposed by Sarhan and Greenbergare
linear combinations of the logarithms of the rexact survival times:
/afii9839/p24/p58/p80/p26
/p71/p14/p16a/p71logt/p7/p71/p8(7.4.5)
and
/afii9846/p24/p58/p80/p26
/p71/p14/p16b/p71logt/p7/p71/p8(7.4.6)
where the coefficients a/p71andb/p71are calculated and tabulated by Saharan and
Greenbergfor n/p4520 and are partially reproduced in Table B-8. The variance
and covariance of /afii9839/p24and /afii9846/p24are tabulated in Table B-9.
The followingexample illustrates the procedure.
Example 7.11 Supposethatinastudyoftheefficacyofanewdrug,12mice
with tumors are given the drug. The experimenter decides to terminate thestudy after 9 mice have died. The survival times are, in weeks, 5, 8, 9, 10, 12,15, 20, 21, 25, 25 /p59,2 5/p59, and 25 /p59. Assume that the times to death of these182
mice follow the lognormal distribution. In this case n/p5812,r/p589, and
n/p57r/p583. Using (7.4.5 ),(7.4.6 ), and Table B-8, /afii9839/p24and /afii9846/p24can be calculated as
/afii9839/p24/p580.036log5 /p590.0581log8 /p590.0682log9 /p590.0759log10 /p590.0827log12
/p590.0888log15 /p590.0948log20 /p590.1006log21 /p590.3950log25
/p582.811
/afii9846/p24/p58/p570.2545log5 /p570.1487log8 /p570.1007log9 /p570.0633log10
/p570.0308log12 /p570.0007log15 /p590.0286log20 /p590.0582log21
/p590.5119log25
/p580.747
The variance of /afii9839/p24and /afii9846/p24given in Table B-9 are, respectively, 0.0926 and 0.0723
and the covariance of /afii9839/p24and /afii9846/p24is 0.0152.
Cohen’s (1959, 1961 )MLEs for the normal distribution can be used for
n/p5720. Let
y/p21/p581
r/p80/p26
/p71/p14/p16logt/p7/p71/p8(7.4.7)
and
s/p17/p581
r/p3/p26(logt/p7/p71/p8)/p17/p57(/afii9814logt/p7/p71/p8)/p17
r /p4(7.4.8)
Then the MLEs of /afii9839and /afii9846/p17are
/afii9839/p24/p58y/p21/p57/afii9838/p19(y/p21/p57logt/p7/p80/p8)( 7 .4.9)
and
/afii9846/p24/p17/p58s/p17/p59 /afii9838/p19(y/p21/p57logt/p7/p80/p8)/p17 (7.4.10)
where the value of /afii9838/p19has been tabulated by Cohen (1961 )as a function of aand
b. The proportion of censored observations, b, is calculated as
b/p58n/p57r
n
and
a/p581/p57Y(Y/p57c)
(Y/p57c)/p17
whereY/p58[b/(1/p57b)]f(c)/F(c),f(c) andF(c) beingthe density and distribu- 183
tion functions, respectively, of the standard normal distribution, evaluated at
c/p58(logt/p7/p80/p8/p57/afii9839)//afii9846. Table 7.1 gives values of /afii9838/p19forb/p580.01 to 0.90 and a/p580.00
to 1.00. For a censored sample, after computing a/p58s/p17/(y/p21/p57logt/p7/p80/p8)/p17, and
b/p58(n/p57r)/n, enter Table 7.1 with these values of aandbto obtain /afii9838/p19. For
values not tabulated, two-way linear interpolation can be used.
The asymptotic variances and covariance are the following:
Var(/afii9839/p24)/p58/afii9846/p17
nm/p16
Var(/afii9846/p24)/p58/afii9846/p17
nm/p17(7.4.11 )
Cov(/afii9839/p24,/afii9846/p24)/p58/afii9846/p17
nm/p18
wherem/p16,m/p17, andm/p18are also tabulated by Cohen (1961 ). The table is
reproduced in Table 7.2. For any censored sample, compute c/p24/p58(logt/p7/p80/p8/p57/afii9839/p24)//afii9846/p24
and then enter the appropriate columns of Table 7.2 with y/p58/p57c/p24, and
interpolateto obtain the required values of m/p71,i/p581, 2, 3, if the experiment was
terminated after a predetermined time. If the experiment was terminated aftera given proportion of animals have died, enter Table 7.2 through the percentcensored column with percentage censored /p58100band interpolate to obtain
the required value of m/p71.
To illustrate the use of Tables 7.1 and 7.2 for the computation of /afii9839/p24,/afii9846/p24/p17,
Var(/afii9839/p24), Var (/afii9846/p24), and Cov (/afii9839/p24,/afii9846/p24), consider Example 7.12, adapted from Cohen
(1961 ).
Example 7.12 Suppose that in a laboratory experiment 300 insects were
followed until 119 died within seconds, y/p21/p581,304.832 seconds, s/p17/p5812,128.250,
and logt/p7/p16/p16/p24/p8/p581,450.000 seconds. In this case n/p58300 andr/p58119. Accord-
ingly,
a/p24/p5812,128.25
(1,304.832/p571,450) /p17
/p580.575b/p58300/p57119
300/p580.603
From Table 7.1, /afii9838/p19is approximately 1.36. Using (7.4.9 )and (7.4.10 ), we obtain
/afii9839/p24/p581,304.832 /p571.36(1,304.832 /p571,450 )/p581,502.26 seconds
/afii9846/p24/p17/p5812,128.250 /p591.36(1,304.832 /p571,450 )/p17/p5840,788.55
and /afii9846/p24/p17/p58201.96 seconds.
For the variance and covariance of /afii9839/p24and /afii9846/p24, we enter Table 7.2 with
percentage censored 100 b/p5860.3 and interpolate linearly to obtain m/p16/p582.002,184
Table7.1 Estimate d Value s for /afii9839/p24and/afii98462
b
a.01 .02 .03 .04 .05 .06 .07 .08 .09 .10 .15 .20 .25 .30 .35 .40 .45 .50 .55 .60 .65 .70 .80 .90
.00 .010100 .020400 .030902 .041583 .052507 .063627 .074953 .086488 .09824 .11020 .17342 .24268 .31862 .4021 .4941 .5961 .7096 0.8368 0.9808 1.145 1 .336 1.561 2.176 3.282
.05 .010551 .021294 .032225 .043350 .054670 .066189 .077909 .089834 .10197 .11431 .17935 .25033 .32793 .4130 .5066 .6101 .7252 0.8540 0.9994 1.166 1 .358 1.585 2.203 3.314
.10 .010950 .022082 .033398 .044902 .056596 .068483 .080568 .092852 .10534 .11804 .18479 .25741 .33662 .4233 .5184 .6234 .7400 0.8703 1.017 1.185 1. 379 1.608 2.229 3.345
.15 .011310 .022798 .034466 .046318 .058356 .070586 .038009 .095629 .10845 .12148 .18985 .26405 .34480 .4330 .5296 .6361 .7542 0.8860 1.035 1.204 1. 400 1.630 2.255 3.376
.20 .011642 .023459 .035453 .047629 .059990 .072539 .085280 .098216 .11135 .12469 .19460 .27031 .35255 .4422 .5403 .6483 .7678 0.9012 1.051 1.222 1. 419 1.651 2.280 3.405
.25 .011952 .024076 .036377 .048858 .061522 .074372 .087413 .10065 .11408 .12772 .19910 .27626 .35993 .4510 .5506 .6600 .7810 0.9158 1.067 1.240 1.4 39 1.672 2.305 3.435
.30 .012243 .024658 .037249 .050018 .062969 .076106 .089433 .10295 .11667 .13059 .20338 .28193 .36700 .4595 .5604 6.713 .7937 0.9300 1.083 1.257 1.4 57 1.693 2.329 3.464
.35 .012520 .025211 .038077 .051120 .064345 .077756 .091355 .10515 .11914 .13333 .20747 .28737 .37379 .4676 .5699 .6821 .8060 0.9437 1.098 1.274 1.4 76 1.713 2.353 3.492
.40 .012784 .025738 .038866 .052173 .065660 .079332 .093193 .10725 .12150 .13595 .21139 .29260 .38033 .4755 .5791 .6927 .8179 0.9570 1.113 1.290 1.4 94 1.732 2.376 3.520
.45 .013036 .026243 .039624 .053182 .066921 .080845 .094958 .10926 .12377 .13847 .21517 .29765 .38665 .4831 .5880 .7029 .8295 0.9700 1.127 1.306 1.5 11 1.751 2.399 3.547
.50 .013279 .026728 .040352 .054153 .068135 .082301 .096657 .11121 .12595 .14090 .21882 .30253 .39276 .4904 .5967 .7129 .8408 0.9826 1.141 1.321 1.5 28 1.770 2.421 3.575
.55 .013513 .027196 .041054 .055089 .069306 .083708 .098298 .11308 .12806 .14325 .22235 .30725 .39870 .4976 .6051 .7225 .8517 0.9950 1.155 1.337 1.5 45 1.788 2.443 3.601
.60 .013739 .027649 .041733 .055995 .070439 .085068 .099887 .11490 .13011 .14552 .22578 .31184 .40447 .5045 .6133 .7320 .8625 1.007 1.169 1.351 1.56 1 1.806 2.465 3.628
.65 .013958 .028087 .042391 .056874 .071538 .086388 .10143 .11666 .13209 .14773 .22910 .31630 .41008 .5114 .6213 .7412 .8729 1.019 1.182 1.366 1.577 1.824 2.486 3.654
.70 .014171 .028513 .043030 .057726 .072605 .087670 .10292 .11837 .13402 .14987 .23234 .32065 .41555 .5180 .6291 .7502 .8832 1.030 1.195 1.380 1.593 1.841 2.507 3.679
.75 .014378 .028927 .043652 .058556 .073643 .088917 .10438 .12004 .13590 .15196 .23550 .32489 .42090 .5245 .6367 .7590 .8932 1.042 1.207 1.394 1.608 1.858 2.528 3.705
.80 .014579 .029330 .044258 .059364 .074655 .090133 .10580 .12167 .13773 .15400 .23858 .32903 .42612 .5308 .6441 .7676 .9031 1.053 1.220 1.408 1.624 1.875 2.548 3.730
.85 .014755 .029723 .044848 .060153 .075642 .901319 .10719 .12325 .13952 .15599 .24158 .33307 .43122 .5370 .6515 .7761 .9127 1.064 1.232 1.422 1.639 1.892 2.568 3.754
.90 .014967 .030107 .045425 .060923 .076606 .092477 .10854 .12480 .14126 .15793 .24452 .33703 .43622 .5430 .6586 .7844 .9222 1.074 1.244 1.435 1.653 1.908 2.588 3.779
.95 .015154 .030483 .045989 .061676 .077549 .0093611 .10987 .12632 .14297 .15983 .24740 .34091 .44112 .5490 .6656 .7925 .9314 1.085 1.255 1.448 1.66 8 1.924 2.607 3.803
1.00 .015338 .030850 .046540 .062413 .078471 .094720 .11116 .12780 .14465 .16170 .25022 .34471 .44592 .5548 .6724 .8005 .9406 1.095 1.267 1.461 1.68 2 1.940 2.626 3.827
Source:Cohen (1961 ).
/p63For all values 0 /p45a/p451,/afii9838/p580.
185
Table 7.2 Estimated Values of m1,m2. and m3for Var( /afii9839/p24), Var( /afii9846/p24),
and Cov( /afii9839/p24,/afii9846/p24)
Percentage
ym/p16m/p17m/p18Censored
/p574.0 1.00000 0.500030 0.000006 0.00
/p573.5 1.00001 0.500208 0.000052 0.02
/p573.0 1.00010 0.501180 0.000335 0.13
/p572.5 1.00056 0.505280 0.001712 0.62
/p572.4 1.00078 0.506935 0.002312 0.82
/p572.3 1.00107 0.509030 0.003099 1.07
/p572.2 1.00147 0.511658 0.004121 1.39
/p572.1 1.00200 0.514926 0.005438 1.79
/p572.0 1.00270 0.518960 0.007123 2.28
/p571.9 1.00363 0.523899 0.009266 2.87
/p571.8 1.00485 0.529899 0.011971 3.59
/p571.7 1.00645 0.537141 0.015368 4.46
/p571.6 1.00852 0.545827 0.019610 5.48
/p571.5 1.01120 0.556186 0.024884 6.68
/p571.4 1.01467 0.568417 0.031410 8.08
/p571.3 1.01914 0.582981 0.039460 9.68
/p571.2 1.02488 0.600046 0.049355 11.51
/p571.1 1.03224 0.620049 0.061491 13.57
/p571.0 1.04168 0.643438 0.076345 15.87
/p570.9 1.05376 0.670724 0.094501 18.41
/p570.8 1.06923 0.702513 0.116674 21.19
/p570.7 1.08904 0.739515 0.143744 24.20
/p570.6 1.11442 0.782574 0.176698 27.43
/p570.5 1.14696 0.832691 0.217183 30.85
/p570.4 1.18876 0.891077 0.266577 34.46
/p570.3 1.24252 0.959181 0.327080 38.21
/p570.2 1.31180 1.03877 0.401326 42.07
/p570.1 1.40127 1.13198 0.492641 46.02
0.0 1.51709 1.24145 0.605233 50.00
0.1 1.66743 1.37042 0.744459 53.980.2 1.86310 1.52288 0.917165 57.930.3 2.11857 1.70381 1.13214 61.790.4 2.45318 1.91942 1.40071 65.540.5 2.89293 2.17751 1.73757 69.15
0.6 3.47293 2.48793 2.16185 72.57
0.7 4.24075 2.86318 2.69858 75.800.8 5.2612 3.3192 3.3807 78.810.9 6.6229 3.8765 4.2517 81.591.0 8.4477 4.5614 5.3696 84.131.1 10.903 5.4082 6.8116 86.43
1.2 14.224 6.4616 8.6818 88.49
1.3 18.735 7.7804 11.121 90.321.4 24.892 9.4423 14.319 91.92186
Table7.2 Continued
Percentage
ym/p16m/p17m/p18Censored
1.5 33.339 11.550 18.539 93.32
1.6 44.986 14.243 24.139 94.521.7 61.132 17.706 31.616 95.54
1.8 83.638 22.193 41.664 96.41
1.9 115.19 28.046 55.252 97.13
2.0 159.66 35.740 63.750 97.722.1 222.74 45.930 99.100 98.212.2 312.73 59.526 134.08 98.612.3 441.92 77.810 182.68 98.932.4 628.58 102.59 250.68 99.18
2.5 899.99 136.44 346.53 99.38
Source:Cohen (1961 ).
m/p17/p581.635, andm/p18/p581.051.Substitutingthese values and /afii9846/p24/p17/p5840,788.55 into
(7.4.11 ), we obtain
Var(/afii9839/p24)/p6040,788.55 (2.022 )
300/p58274.91
var(/afii9846/p24)/p6040,788.55 (1.635 )
300/p58222.30
Cov(/afii9839/p24,/afii9846/p24)/p6040,788.55 (1.051 )
300/p58142.90
When the data are progressively censored, let t/p16,t/p17,...,t/p80be the uncensored
andt/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76be the censored observations, the likelihood function,
using (7.4.1 )and (7.1.1 ),i s
l(/afii9839,/afii9846/p17)/p58/p57rlog(2/afii9843/afii9846/p17)
2/p57/p80/p26
/p71/p14/p16/p1logt/p71/p59(logt/p71/p57/afii9839)/p17
2/afii9846/p17 /p2
/p59/p76/p26
/p71/p14/p80/p62/p16log/p7/p16/p27
/p82/p62/p711
x/p402/afii9843/afii9846/p17exp/p3/p571
2/afii9846/p17(logx/p57/afii9839)/p17/p4dx/p8
and the MLE of /afii9839and /afii9846/p17can be obtained by solvingthe followingtwo 187
equations:
/p80/p26
/p71/p14/p16logt/p71/p57/afii9839
/afii9846/p17/p59/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p62/p71logx/p57/afii9839
x/afii9846/p17/p402/afii9843/afii9846/p17exp/p3/p571
2/afii9846/p17(logx/p57/afii9839)/p17/p4dx
/p16/p27
/p82/p62/p711
x/p402/afii9843/afii9846exp/p3/p571
2/afii9846/p17(logx/p57/afii9839)/p17/p4dx/p580
/p57n
2/afii9846/p17/p59/p80/p26
/p71/p14/p16(logt/p71/p57/afii9839)/p17
2/afii9846/p19
/p59/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p62/p71(logx/p57/afii9839)/p17
x2/afii9846/p19/p402/afii9843/afii9846/p17exp/p3/p571
2/afii9846/p17(logx/p57/afii9839)/p17/p4dx
/p16/p27
/p82/p62/p711
x/p402/afii9843/afii9846/p17exp/p3/p571
2/afii9846/p17(logx/p57/afii9839)/p17/p4dx/p580
Again,this can be done by applyingthe Newton —Raphson iterative procedure.
The followingexample illustrate the use of SAS and BMDP to obtain
estimates of the lognormal parameters.
Example 7.13 Referringto Example 7.5, for the observed survival data in
the file ‘‘EXAMPLE.DAT’’, by changing d /p58exponential in SAS code to
d/p58lnormal and accel /p58exponential in BMDP code to accel /p58lnormal we
can obtain the estimated parameters of the lognormal distribution. If SAS isused, the estimated parameters of the lognormal distribution are
/afii9839/p24/p58INTERCEPT and /afii9846/p24/p17/p58SCALE /p17
where INTERCEPT and SCALE are the names of output estimated par-
ameters in SAS procedure LIFEREG. If BMDP is used,
/afii9839/p24/p58CONSTANT and /afii9846/p24/p17/p58SCALE /p17
where CONSTANT and SCALE are given by procedure 2L.
7.5 STANDARD AND GENERALIZED GAMMA DISTRIBUTIONS
The density function of the standard gamma distribution is
f(t)/p58/afii9838
/afii9772(/afii9828)(/afii9838t)/p65/p92/p16exp(/p57/afii9838t)t/p460, /afii9838,/afii9828/p570 (7.5.1 )188
where
/afii9772(/afii9828)/p58/p7/p25/p27/p15x/p65/p92/p16e/p92/p86dx
(/afii9828/p571)! if /afii9828is an integer(7.5.2 )
In this section we discuss the MLE of /afii9838and /afii9828for data with and without
censored observations.
7.5.1 Estimation of /afii9838and/afii9828for Data without Censored Observations
Suppose that the npatients under study are followed to death and their exact
survival times t/p16,t/p17,...,t/p76are known. The MLE of /afii9838and /afii9828can be obtained
by solvingsimultaneously the two equations
n/afii9828/p24
/afii9838/p19/p57/p76/p26
/p71/p14/p16t/p71/p580( 7 .5.3)
and
nlog/afii9838/p19/p57n/afii9772/p30(/afii9828/p24)
/afii9772(/afii9828/p24)/p59/p76/p26
/p71/p14/p16logt/p71/p580( 7 .5.4)
where /afii9772/p30(/afii9828)is the derivative of /afii9772(/afii9828),
/afii9772/p30(/afii9828)/p58/p16/p27
/p15x/p65/p92/p16log(x)e/p92/p86dx (7.5.5)
From (7.5.3 ), we have
/afii9838/p19/p58n/afii9828/p24
/afii9814/p76/p71/p14/p16t/p71(7.5.6)
On eliminating /afii9838, we substitute (7.5.6 )into (7.5.4 )and obtain
/afii9772/p30(/afii9828/p24)
/afii9772(/afii9828/p24)/p57log/afii9828/p24/p57log(/afii9811/p76/p71/p14/p16t/p71)/p16/p30/p76
/afii9814/p76/p71/p14/p16t/p71/n/p580( 7 .5.7)
to solve for /afii9828/p24. This can be done by usingthe Newton —Raphson iterative
procedure. Tables for the solution of (7.5.7 )for/afii9828/p24as a function of Rare given
by Greenwood and Durand (1960 ), whereRis the ratio of the geometric mean
to the arithmetic mean of the nobservations:
R/p58(/afii9811/p76/p71/p14/p16t/p71)/p16/p30/p76
/afii9814/p76/p71/p14/p16t/p71/n(7.5.8) 189
Wilket al. (1962a )show thatthe relationshipbetween /afii9828/p24and 1/ (1/p57R) is linear.
A table of /afii9828/p24values of as a function of 1/ (1/p57R)given in their paper is
reproduced in Table B-10. Thus if Rand 1/ (1/p57R) are computed from the
sample, a MLE of /afii9828can be found from Table B-10. For values not tabulated,
linear interpolation can be used. Having /afii9828/p24so determined, /afii9838/p19can be obtained
from (7.5.6 ).
In the method of moments (Fisher, 1922 ), the estimators are obtained
simply by equatingthe population mean and variance to the sample mean and
variance. The moment estimators of /afii9828and /afii9838are
/afii9838*/p58/afii9814/p76/p71/p14/p16t/p71/afii9814/p76/p71/p14/p16(t/p71/p57t/p16)/p17(7.5.9)
and
/afii9828*/p58(/afii9814/p76/p71/p14/p16t/p71)/p17
n/afii9814/p76/p71/p14/p16(t/p71/p57t/p16)/p17(7.5.10)
Both types of estimators give biased estimates. The moment estimators are
easy to calculate but are inefficient in the sense that their variances are largerthan the variance of the MLE. To reduce the bias, Lilliefors (1971 )suggests
correction factors for these two types of estimators. The corrected MLE of /afii9828
and /afii9838are, respectively,
/afii9828/p24/p65/p58/afii9828/p24
1/p593/n(7.5.11)
and
/afii9838/p19/p65/p58/afii9828/p24/p65t/p16/p11/p571
n/afii9828/p24/p65/p2(7.5.12)
The corrected moments estimators of /afii9828and /afii9838are
/afii9828*/p65/p58/afii9828*
1/p592/n/p573
n(7.5.13)
and
/afii9838*/p65/p58/afii9828*/p65t/p11/p571
n/afii9828*/p65/p2(7.5.14)
Lilliefors shows by the Monte Carlo method that the corrected MLE and
the method-of-moment estimates are approximately unbiased. In addition, aslongas /afii9828/p462, the corrected moments estimators have no more bias than the
corrected MLE and for n/p5810 have considerably less bias. For n/p5810, 20 and
/afii9828/p462, the variance is close to that of the MLE.190
Example 7.14 Ten patients with melanoma achieve remission after surgery
and therapy. They are followed to relapse. The durations of remission inmonths are recordedas follows:5, 8, 10,11, 15, 20,21, 23, 30,and 40. Assumingthat the distribution of remission duration is standard gamma, we firstcalculate the MLE of /afii9828and /afii9838accordingto Wilk et al. (1962a ). To compute R,
we obtain /afii9814/p76/p71/p14/p16t/p71/p58183 and (/afii9811/p76/p71/p14/p16t/p71)/p16/p30/p16/p15 /p5815.43. Therefore R/p580.84 and
1/(1/p57R)/p586.25. From Table B-10, /afii9828/p24/p582.89830 for 1/ (1/p57R)/p586.0 and
/afii9828/p24/p583.14984 for 1/ (1/p57R)/p586.5. By linear interpolation, for 1/ (1/p57R)/p586.25,
/afii9828/p24/p583.02407. From (7.5.6 ),/afii9838/p19/p580.16525. The corrected MLEs obtained from
(7.5.11 )and (7.5.12 )are/afii9828/p24/p65/p582.326 and /afii9838/p19/p65/p580.122. The moment estimates of /afii9828
and /afii9838following (7.5.9 )and (7.5.10 )are /afii9838*/p580.173 and /afii9828*/p583.171. With the
correction factors, /afii9838*/p65/p580.1225 and /afii9828*/p65/p582.3425, which are very close to the
corrected MLE.
7.5.2 Estimation of /afii9828and/afii9838for Data with Censored Observations
When data are singly censored, the survival times can be ordered as
t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p45t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8
whererpersons in the study have exact survival times recorded and n/p57r
others have their lives terminated after the rth death occurs. In this case, the
maximum likelihood procedure becomes much more complicated.
Let /afii9834/p58/afii9838t/p7/p80/p8,P/p58[/afii9811/p80/p71/p14/p16t/p7/p71/p8]/p16/p30/p80/t/p7/p80/p8, andS/p58/afii9814/p80/p71/p14/p16t/p7/p71/p8/rt/p7/p80/p8. The MLE of /afii9834and
/afii9828and can be obtained by solvingsimultaneously
logP/p58n/afii9772/p30(/afii9828)
r/afii9772(/afii9828)
/p57n
rlog/afii9834/p57/p1n
r/p571/p2J/p30(/afii9828,/afii9834)
J(/afii9828,/afii9834)(7.5.15)
and
S/p58/afii9828
/afii9834/p571
/afii9834/p1n
r/p571/p2e/p92/p69
J(/afii9828,/afii9834)(7.5.16)
where
J(/afii9828,/afii9834)/p58/p16/p27
/p16t/p65/p92/p16e/p92/p69/p82dt (7.5.17)
and
J/p30(/afii9828,/afii9834)/p58/p42
/p42/afii9828J(/afii9828,/afii9834)/p58/p16/p27
/p16t/p65/p92/p16logte/p92/p69/p82dt (7.5.18)
Wilk et al. (1962a )generate, for a grid of values of PandSandn/r, tables of
values of /afii9828/p24and /afii9839/p24/p58/afii9828/p24//afii9834/p24based on the solutions of (7.5.15 )and (7.5.16 ). The
tables are reproduced in Table B-11. Thus, to find /afii9828/p24and /afii9838/p19, one needs to 191
computePandSfirst. For specific values of P,S, andn/r,/afii9828/p24and /afii9839/p24may be
looked up from Table B-11. Then /afii9838/p19can be obtained from /afii9838/p19/p58/afii9828/p24/[/afii9839/p24t/p7/p80/p8].
Interpolations may be needed when any of the values of P,S, andn/rare not
tabulated.
Example 7.15, adapted from Wilk et al. (1962a ), illustrates the procedure of
calculating /afii9828/p24,/afii9839/p24, and /afii9838/p19when Table B-11 is used. This method can also be used
in the case of a complete sample (no censored observations ); that is,r/p58n.I f
it is obvious that some of the observations may be outliers (too large or too
small ), it is reasonable not to use them in estimation. In this case, ris the
number of observations used in the estimation procedure.
Example 7.15 Consider an experiment with n/p5834 animals. The following
data are the lifetimes t/p71in weeks of 34 animals: 3, 4, 5, 6, 6, 7, 8, 8, 9, 9, 9, 10,
10, 11, 11, 11, 13, 13, 13, 13, 13, 17, 17, 19, 19, 25, 29, 33, 42, 42, 52, 52 /p59,5 2/p59,
and 52 /p59. The study is terminated when 31 animals have died and the other 3
are sacrificed. In our notation, n/p5834 andr/p5831.
1. Compute n/r,P, andS:
n
r/p5834
31/p581.10
S/p58/afii9814t/p71rt/p7/p80/p8/p58487
(31)(52)/p580.30
To compute P, it is easier first to compute log P:
logP/p581
r/p26logt/p7/p71/p8/p57logt/p7/p80/p8/p581
31/p5933.90207 /p571.716/p58/p570.622385
HenceP/p580.24.
2. Consider the entries for n/r/p581.10 andP/p580.24 in Table B-11:
S/p580.28: /afii9828/p24/p581.986 /afii9839/p24/p580.365
S/p580.32: /afii9828/p24/p581.449 /afii9839/p24/p580.410
Usinglinear interpolation, approximate estimates of /afii9828and /afii9839are/afii9828/p24/p581.72
and /afii9839/p24/p580.39.
3. Finally, /afii9838/p19/p581.72/ (0.39/p5952)/p580.085.
For a more accurate two-way interpolation, the reader is referred to Wilk
et al. (1962a ).192
When the data are progressively censored, let t/p16,t/p17,...,t/p80be the uncensored
andt/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76be the censored observations; the likelihood function is
l(/afii9838,/afii9828)/p58logL(/afii9838,/afii9828)/p58n/afii9828log/afii9838/p57nlog/afii9772(/afii9828)
/p59/p80/p26
/p71/p14/p16[(/afii9828/p571)logt/p71/p57/afii9838t/p71]/p59/p76/p26
/p71/p14/p80/p62/p16log/p1/p16/p27
/p82/p62/p71x/p65/p92/p16e/p92/p72/p86dx/p2
and the MLE of /afii9838and /afii9828can be obtained by solvingthe two equations
n/afii9828
/afii9838/p57/p80/p26
/p71/p14/p16t/p71/p57/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p71/p62x/p65e/p92/p72/p86dx
/p16/p27
/p82/p71/p62x/p65/p92/p16e/p92/p72/p86dx/p580
nlog/afii9838/p57n/afii9772/p30(/afii9828)
/afii9772(/afii9828)/p59/p80/p26
/p71/p14/p16logt/p71/p59/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p71/p62x/p65/p92/p16e/p92/p72/p86log(x)dx
/p16/p27
/p82/p71/p62x/p65/p92/p16e/p92/p72/p86dx/p580
usingthe Newton —Raphson iterative procedure.
7.5.3 Estimation of /afii9838,/afii9828, and /afii9825in the Extended Generalized Gamma
Distribution for Data with or without Censored Observations
The extended generalized gamma distribution has density function defined in
(6.4.10 ),
f(t)/p58/p34/afii9825/p34/afii9828/p65/afii9838/p63/p65t/p63/p65/p92/p16exp[/p57/afii9828(/afii9838t)/p63]
/afii9772(/afii9828)t/p570, /afii9828/p570, /afii9838/p570( 7.5.19)
Let us consider the case /afii9825/p570. Lett/p16,t/p17,...,t/p80be the uncensored and
t/p62/p80/p62/p16,...,t/p62/p76the censored observations from npersons and the survival times
follow the generalized gamma distribution. Then the likelihood function is
l(/afii9838,/afii9828,/afii9825)/p58n/afii9825/afii9828log/afii9838/p59nlog/afii9825/p59n/afii9828log/afii9828/p57nlog/afii9772(/afii9828)
/p59/p80/p26
/p71/p14/p16[(/afii9825/afii9828/p571)logt/p71/p57/afii9828(/afii9838t/p71)/p63]
/p59/p76/p26
/p71/p14/p80/p62/p16log/p7/p16/p27
/p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x)/p63]dx/p8 193
and the MLE of /afii9838,/afii9828, and /afii9825can be obtained by solvingthe three equations
n/afii9825/afii9828
/afii9838/p57/afii9828/afii9825/afii9838/p63/p92/p16/p80/p26
/p71/p14/p16t/p63/p71/p57/afii9828/afii9825/afii9838/p63/p92/p16/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p71/p62x/p63/p65/p92/p16/p62/p63exp(/p57/afii9828(/afii9838x)/p63)dx
/p16/p27
/p82/p71/p62x/p63/p65/p92/p16exp(/p57/afii9828(/afii9838x)/p63)dx/p580( 7.5.20)
n/afii9825log/afii9838/p59nlog/afii9828/p59n/p57n/afii9772/p30(/afii9828)
/afii9772(/afii9828)/p59/p80/p26
/p71/p14/p16[/afii9825logt/p71/p57(/afii9838t/p71)/p63]
/p59/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x)/p63][/afii9825log(x)/p57(/afii9838x)/p63]dx
/p16/p27
/p82/p71/p62x/p63/p65/p92/p16exp(/p57/afii9828(/afii9838x)/p63)dx/p580 (7.5.21 )
n/afii9828log/afii9838/p59n
/afii9825/p59/p80/p26
/p71/p14/p16[/afii9828logt/p71/p57/afii9828(/afii9838t/p71)/p63log(/afii9838t/p71)]
/p59/p76/p26
/p71/p14/p80/p62/p16/p16/p27
/p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x)/p63][/afii9828log/p67(x)/p57/afii9828(/afii9838x)/p63log(/afii9838x)]dx
/p16/p27
/p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x]/p63)dx/p580( 7.5.22)
usingthe Newton —Raphson iterative procedure.
If all the observed survival times are uncensored, the respective equations
for the MLE of /afii9838,/afii9828, and /afii9825can be obtained simply by replacing rwithnin
(7.5.20 )—(7.5.22 ). The SAS procedure LIFEREG can be used to obtain the
MLE of /afii9838,/afii9828, and /afii9825in the extended generalized gamma distribution.
Example 7.16 Referringto Example 7.5, for the observed survival data in
the file ‘‘EXAMPLE.DAT‘, by changing d /p58exponential in the SAS code to
d/p58gamma, one can obtain the MLE of the parameters of the extended
generalized gamma distribution:
/afii9838/p19/p58exp(/p57INTERCEPT ) /afii9825/p24/p58SHAPE1
SCALE/afii9828/p24/p581
SHAPE1 /p17
where INTERCEPT, SHAPE1, and SCALE are given by the SAS LIFEREG
procedure.194
7.6 LOG-LOGISTIC DISTRIBUTION
The log-logistic distribution has the density function
f(t)/p58/afii9825/afii9828t/p65/p92/p16
(1/p59/afii9825t/p65)/p17(7.6.1)
and survivorship function
S(t)/p581
1/p59/afii9825t/p65(7.6.2)
wheret/p460,/afii9825/p570,/afii9828/p570.Lett/p16,t/p17,...,t/p80be the uncensored and t/p62/p80/p62/p16,
t/p62/p80/p62/p17,...,t/p62/p76the censored observations from npersons and the survival times
follow the log-logistic distribution. Then the MLE of /afii9825and /afii9828can be obtained
from solvingthe followingtwo simultaneous equations:
r/p57/afii9825/p12/p80/p26
/p71/p14/p16t/p65/p711/p59/afii9825t/p65/p71/p59/p76/p26
/p71/p14/p80/p62/p16t/p62/p65/p711/p59/afii9825t/p62/p65/p71/p2/p580 (7.6.3 )
r
/afii9828/p59/p80/p26
/p71/p14/p16log(t/p71)/p57/afii9825/p32/p80/p26
/p71/p14/p16t/p65/p71log(t/p71)
1/p59/afii9825t/p65/p71/p59/p76/p26
/p71/p14/p80/p62/p16t/p62/p65/p71log(t/p62/p71)
1/p59/afii9825t/p62/p65/p71/p4/p580 (7.6.4 )
usingthe Newton —Raphson iterative procedure. If all the survival times
observed are uncensored, the respective equations for the MLE of /afii9825and /afii9828can
be obtained simply by replacing rwithnin(7.6.3 )and (7.6.4 ).
Example 7.17 Referringto Example 7.5, for the observed survival data in
file ‘‘EXAMPLE.DAT‘, replacingd /p58exponential in the SAS code by
d/p58llogistic and accel /p58exponential in the BMDP code by accel /p58llogistic,
we can obtain the estimated parameters of the log-logistic distribution. If SASis used, the estimated parameters of the log-logistic distribution are
/afii9825/p24/p58exp/p1/p57INTERCEPT
SCALE /p2and /afii9828/p24/p581
SCALE
where INTERCEPT and SCALE are given by the SAS LIFEREG procedure.
If BMDP is used,
/afii9825/p24/p58exp/p1/p57CONSTANT
SCALE /p2and /afii9828/p24/p581
SCALE
where CONSTANT and SCALE are produced by the BMDP procedure 2L.- 195
Example 7.18 Assume that the tumor-free time of the 30 rats in the low-fat
diet group in Table 3.4 follows the log-logistic distribution. The estimates ofthe two parameters from either SAS or BMDP are /afii9825/p24/p580.000025484 and
/afii9828/p24/p582.01866. Therefore, from Section 6.5, the median survival time for this
group is 188.64 days, and the hazard function approach the peak at 190.37days.
7.7 OTHER PARAMETRIC SURVIVAL DISTRIBUTIONS
The Gompertz distribution (Section 6.6 )has the followingsurvivorship and
probability density functions:
S(t)/p58exp
/p3/p57e/p72
/afii9828(e/p65/p82/p571)/p4(7.7.1)
f(t)/p58exp/p3(/afii9838/p59/afii9828t)/p571
/afii9828(e/p72/p62/p65/p82/p57e/p72)/p4(7.7.2)
7.7.1 Estimation of /afii9838and/afii9828for Data with or without Censored Observations
Assumethat t/p16,t/p17,...,t/p76are the observedsurvivaltimes from nindividualsand
the survival times follow the Gompertz distribution, without loss of generality,andassumethat t/p16,t/p17,...,t/p80areuncensoredand t/p62/p80,t/p62/p80/p62/p17,...,t/p62/p76right-censored.
The MLE of /afii9838and /afii9828can be obtained by solvingthe equations
r/p59e/p72
/afii9828/p7/p80/p26
/p71/p14/p16[1/p57exp(/afii9828t/p71)]/p59/p76/p26
/p71/p14/p80/p62/p16[1/p59exp(/afii9828t/p62/p71)]/p8/p580( 7.7.3)
/p80/p26
/p71/p14/p16t/p71/p57e/p72
/afii9828/p17/p7/p80/p26
/p71/p14/p16[1/p59(/afii9828t/p71/p571)exp (/afii9828t/p71)]/p59/p76/p26
/p71/p14/p80/p62/p16[1/p59(/afii9828t/p62/p71/p571)exp (/afii9828t/p62/p71)]/p8/p580
(7.7.4 )
usingthe Newton —Raphson iterative procedure.
If allt/p16,t/p17,...,t/p76are uncensored, the MLE of /afii9838and /afii9828can be obtained
similarly by replacing rwithnin(7.7.3 )and (7.7.4 ). The MLE of the
parameters of the other models in Section 6.6 can be obtained in a similarmanner.
Bibliographical Remarks
In addition to the papers cited in this chapter, Gross and Clark (1975 )have
chapterson estimation and inferencein the exponentialdistribution and on the
estimation of parameters of three distributions, includingthe Weibull and196
gamma. Mann et al. (1974 ), Lawless (1982 ), and Nelson (1982 )also provide a
chapter on the estimation procedures for survival distributions, includingtheexponential, Weibull, gamma, and lognormal. A more recent book is by Kleinand Moeschberger (1997 ). Readers with a background in mathematical statis-
tics and an interest in mathematical treatment of estimation procedures arereferred to these books.
EXERCISES
7.1Consider the survival times given in Exercise 8.2. Assuming that they
follow the one-parameter exponential distribution, obtain:(a)The MLE of /afii9838
(b)The MLE of /afii9839
(c)The 95% confidence intervals for /afii9838and /afii9839
7.2Assumingthatthe correctentriesbetweenerrors inExercise 8.3follow the
two-parameter exponential distribution, obtain:(a)An estimate of G
(b)The MLE of /afii9838
(c)The MLE of /afii9839
(d)The probability of 100 correct entries between two errors
7.3Consider the survival data in Exercise 8.5. Obtain the MLE of the
parameter (s)and mean survival times, assuming:
(a)A one-parameter exponential distribution
(b)A Weibull distribution
7.4In a study of deep venous thrombosis, the followingblood clot lysis times
in hours were recorded from 20 patients: 2, 3, 4, 5.5, 9, 13, 16.5, 17.5, 12.5,7, 6, 17.5, 11.5, 6, 14, 25, 49, 37.5, 49, and 28. Assume that the blood clotlysis times follow the lognormal distribution.(a)Obtain MLEs of the parameters /afii9839and /afii9846/p17.
(b)Obtain 95% confidence intervals for /afii9839and /afii9846/p17.
7.5Consider the followingtumor-free times in days of 10 animals: 2, 3.5, 5,
7, 9, 10, 15, 20, 30, and 40. Assume that the tumor-free times follow thelog-logistic distribution. Estimate the parameters /afii9825and /afii9828. 197
CHAPTER 8
Graphical Methods for
Survival Distribution Fitting
The use of probabilitymodels for survival experience has play ed an increasing-lyimportant role in biomedical sciences. Survival models summarize thesurvival pattern, suggest further studies, and generate hypotheses. In thischapter we introduce three graphical methods for survival distribution fitting.
In Section 8.1 we discuss the advantages of the graphical techniques. In
Section 8.2 we discuss probabilityplotting, including how to make probabilityplots and how to estimate parameters from them. In Section 8.3 we discuss thetheoryand applications of hazard plotting for censored data. In Section 8.4 weintroduce the Cox —Snell residual method.
8.1 INTRODUCTION
Graphical methods have long been used for displayand interpretation of data
because theyare simple and effective. Often used in place of or in conjunctionwith numerical analysis, a plot of data serves a number of purposes simulta-neouslythat no numerical method can. The basic idea of the three graphicalmethods is to see if the survival time itself, or a function of it, has a linearrelationship with the distribution function and the cumulative hazard functionof a given parametric distribution, or a function of the distribution functionand the cumulative hazard function. If such a linear relationship exists, it canbe demonstrated graphicallyas a straight line. Thus, if one chooses theappropriate distribution and makes a probability, or hazard, plot, the resultwill be a straight line fit to the data. Parameters of the distribution chosen canbe estimated from the probabilityor hazard plots without tedious numericalcalculations. Such estimates maybe adequate and useful for preliminarypurposes. However, prior information is often not sufficient to choose asuitable distribution, and the plot maynot be a straight line. If the plot is nota straight line, there is no need to estimate the parameters and an alternative
198
Figure 8.1 Two curved normal probabilityplots.
Figure 8.2 Two skewed densityfunctions.distributionmaybe selected.If the Cox —Snell residual plotting method is used,
estimates of the parameters must be obtained first.
A nonlinear plot can provide insight into the data. There are several pos-
sible interpretations. First, the wrong theoretical distribution might havebeen used. Second, the sample might be from a mixture of populations. In thelatter case, it is necessaryto separate the data accordinglyand make a separateplot for each population. If one or two points are wayout of line, theymightbe the results of errors in collecting and recording the data or theymight notbe from the same population. Other reasons for peculiar looking plots andinterpretations of them are discussed byKing (1971 )and Hahn and Shapiro
(1967 ).
Consider the normal probabilityplots in Figures 8.1 aand b. The plot in
Figure 8.1 ais convex, indicating that the data have a long tail to the left and
could be from a distribution with a negativelyskewed densityfunction such asin Figure 8.2 a. On the contrary, the concave plot in Figure 8.1 bindicates that 199
the data have a long tail to the right and could be from a distribution with a
positivelyskewed densityfunction such as in Figure 8.2 b. From the discussion
in Chapter 6, we maytryto fit a lognormal or gamma distribution.
The advantages of graphical methods can be summarized as follows:
1. Theyare fast and simple to use, in contrast with numerical methods,
which maybe computationallytedious and require considerable analy ti-cal sophistication. The additional accuracyof numerical methods is
usuallynot great enough in practice to warrant the effort involved.
2. Probabilityand hazard plots provide approximate estimates of the
parameters of the distribution bysimple graphical means.
3. Theyallow one to assess whether a particular theoretical distribution
provides an adequate fit to the data.
4. Peculiar appearance of a plot or points in a plot can provide insight into
the data when the reasons for the peculiarities are determined.
5. A graph provides a visual representation of the data that is easyto grasp.
This is useful not onlyfor oneself but also in presenting data to others,since a plot allows one to assess conclusions drawn from the data bygraphical or numerical means.
8.2 PROBABILITY PLOTTING
The basic ideas in probabilityplotting are illustrated bythe following example.
Example 8.1 Consider the white blood cell counts (WBCs )of 23 pediatric
leukemia patients given in Table 8.1, ranging from 8000 to 120,000. A sample
cumulative distribution is constructed byordering the data from smallest to
largest,as shownin Table 8.1. A sample cumulativedistributioncurve can thenbemade byplottingeachWBCvalue versusthepercentage ofthe sampleequalto or less than that value. That is, the ith ordered data value in a sample of n
values is plotted against the percentage 100 i/n. Note that for tied observations,
we compute and plot the sample distribution onlyfor the one with the largestivalue. This gives a conservative estimate of the survivorship function. For
example, the third value of WBC, 10, is plotted against a percentage of100/p593/23/p5813%.
A plot of the cumulative distribution function for most large populations
contains manycloselyspaced values and can be well approximated byasmooth curve drawn though the points. In contrast, a sample cumulativedistribution function has a relativelysmall number of points and thus some-what ragged appearance. To approximate the population cumulative distribu-tion function, one draws a smooth curve through the data points, obtaining abest fit byey e. Such a curve from the WBC data is given in Figure 8.3. It is an200
Table 8.1 Ordered WBCsData and Sample Cumulative
Distribution for Example 8.1
Sample
Distribution
WBC Order,
(10/p18) ii / 23 ( i/p570.5)/23 /afii9818/p92/p16(F)/p63
81
82 0 .087 0 .065 /p571.512
10 3 0 .130 0 .109 /p571.233
15 4 0 .174 0 .152 /p571.027
20 5 0 .217 0 .196 /p570.857
30 6 0 .261 0 .239 /p570.709
50 7
50 8
50 950 1050 11 0 .478 0 .457 /p570.109
60 1260 13 0 .565 0 .543 0 .109
75 14
75 15 0 .652 0 .630 0 .333
80 1680 17 0 .739 0 .717 0 .575
90 1890 19
90 20 0 .870 0 .848 1 .027
100 21 0 .913 0 .891 1 .233
110 22 0 .957 0 .935 1 .512
120 23 1 .000 0 .978 2 .019
/p63/afii9818/p92/p16(·) denotes the inverse of the standard normal distribu-
tion function.
estimate of the cumulative distribution function of the population and is used
to obtain estimates and other information about the population.
An estimate of the population median (50th percentile )is obtained by
enteringtheplot onthepercentagescale at50%goinghorizontallyto thefittedline and then verticallydown to the data scale to read the estimate of themedian. For the WBC data, an estimate of the population median is 65,000.The median is a representative of nominal value for the population since halfof the population values are above it and half below. An estimate of anyotherpercentile can be obtained similarlybyentering the plot at the appropriatepoint on the percentage scale going horizontallyto the fitted line and thenverticallydown to the data scale where the estimate is read. For example, anestimate for the 25th percentile is 40,000. 201
Figure 8.3 Sample cumulative distribution curve of the WBC data.
One can obtain an estimate of the proportion of the population that has a
WBC below a specific value in a similar way. For example, to find theproportion of the population with a WBC of 10,000 or less, you enter the ploton the horizontal axis at the given value, 10, go verticallyup to the line fittedto the data, and then horizontallyto the probabilityscale, where the estimateof the population proportion is read, 8%. An estimate of the proportion of apopulation between two given values is obtained byfirst getting an estimate ofthe proportion below each value and then taking the difference. For example,the estimate of the population proportion with WBC between 10,000 and65,000 is 50 /p578/p5842%.
As mentioned above, a smooth curve can be fitted byey e to a sample
cumulative distribution function to obtain an estimate of the populationdistribution function. Also, one can fit data with a theoretical cumulativedistribution function byusing a probabilityplot and then use this plot to
estimatethe parametersin thetheoreticalcumulativedistributionfunction.Thedistribution maybe the normal, lognormal, exponential, Weibull, gamma, orlog-logistic. To make a probabilityplot, one generallyuses (i/p570.5)/nor
i/(n/p591) to estimate the sample cumulative distribution function at the ith
ordered value of the nobservations in the sample. The (i/p570.5)/nfor the WBC
data are given in Table 8.1.
The probabilityplot is so constructed that if the theoretical distribution is
adequate for the data, the graph of a function of t(used as the y-axis )versus
a function of the sample cumulative distribution function (used as the x-axis )
will be close to a straight line. The parameters of the theoretical distributioncan then be estimated from a fitted line. This is carried out as follows.
Step 1. A theoretical distribution for the survival time thas to be selected.
Step 2.The sample cumulative distribution function is estimated byusing
(i/p570.5)/nori/(n/p591),i/p581, 2, ...,n, for the ith ordered tvalue. For tied202
Figure 8.4 Normal probabilityplot of the WBC data in Example 8.1.observations have the same value, the sample cumulative distribution function
is plotted against onlythe twith the largest ivalue.
Step 3.Plot tor a function of it versus the estimated sample cumulative
distribution or a function of it.
Step 4.Fit a straight line through the points byey e. The position of the
straight line should be chosen to provide a fit to the bulk of the data and mayignore outliers or data points of doubtful validity.
Figure8.4 givesa normal probabilityplot of the WBCversus /afii9818/p92/p16(F), where
/afii9818/p92/p16(·)is the inverse of the standard normal distribution function. The values
of/afii9818/p92/p16(F/p19(WBC/p7/p71/p8))are shown in Table 8.1. The plot is reasonablylinear. The
straight line fitted byey e in a probabilityplot can be used to estimatepercentiles and proportions within given limits in the same manner as for thesample cumulative distribution curve. In addition, a probabilityplot providesestimates of the parameters of the theoretical distribution chosen. The mean(or median )WBC estimated from the normal probabilityplot in Figure 8.4 is
56,000 [at /afii9818/p92/p16(F)/p580,F/p580.5 and WBC /p5856,000]. At /afii9818/p92/p16(F)/p581,
WBC /p5891,000, which corresponds to the mean plus 1 standard deviation.
Thus, the standard deviation is estimated as 35,000.
We now discuss probabilityplots of the exponential, Weibull, lognormal,
and log-logistic distributions. 203
Table 8.2 Probability Plotting for Example 8.2
Order, F,
ti (i/p570.5)/21 log[1/ (1/p57F)]
11
1 2 0.071 0.074
23
2 4 0.167 0.1823 5 0.214 0.241464 7 0.310 0.37058
5 9 0.405 0.519
6 10 0.452 0.60281 18 12 0.548 0.7939 13 0.595 0.904
10 14
10 15 0.690 1.173
12 16 0.738 1.34014 17 0.786 1.54016 18 0.833 1.79220 19 0.881 2.12824 20 0.929 2.639
34 21 0.976 3.738Exponential Distribution
The exponential cumulative distribution function is
F(t)/p581/p57exp[/p57(/afii9838t)] t/p570( 8 .2.1)
The probabilityplot for the exponential distribution is based on the relation-
ship between tand F(t), from (8.2.1 ),
t/p581
/afii9838log1
1/p57F(t)(8.2.2)
This relationship is linear between tand the function log[1/ (1/p57F(t))]. Thus,
an exponential probabilityplot is made byplotting the ith ordered observed
survival time t/p7/p71/p8versus log[1/ (1/p57F/p19(t/p7/p71/p8))], where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8),
for example, (i/p570.5)/n, for i/p581,...,n.
From (8.2.2 ), at log /p431/[1/p57F(t)]/p44/p581,t/p581//afii9838. This fact can be used to
estimate 1/ /afii9838and thus /afii9838from the fitted straight line. That is, the value t204
Figure 8.5 Exponential probabilityplot of the data in Example 8.2.corresponding to log /p431/[1/p57F(t)]/p44/p581 is an estimate of the mean 1/ /afii9838and its
reciprocal is an estimate of the hazard rate /afii9838.
Example 8.2 Suppose that 21 patients with acute leukemia have the
following remission times in months: 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 8, 8, 9, 10, 10, 12,14, 16, 20, 24, and 34. We would like to know if the remission time follows theexponential distribution. The ordered remission times t/p7/p71/p8and the log /p431/
[1/p57F(t)]/p44are given in Table 8.2. The exponential probabilityplot is shown
in Figure 8.5. A straight line is fitted to the points byey e,and the plot indicatesthat the exponential distribution fits the data verywell. At the point log[1/(1/p57F(t))]/p581.0, the corresponding t, approximately9.0 months, is an esti-
mateof the mean1/ /afii9838andthus an estimateof thehazardrateis /afii9838/p19/p581/9/p580.111
per month. An alternative is to use (7.2.5 )to estimate /afii9838,/afii9838/p19/p5821/198 /p580.107,
which is veryclose to the graphical estimate.
Weibull Distribution
The Weibull cumulative distribution function is
F(t)/p581/p57exp[/p57(/afii9838t)/p65] t/p570,/afii9828/p570,/afii9838/p570( 8 .2.3)
The probabilityplot for the Weibull distribution is based on the relationship
logt/p58log1
/afii9838
/p591
/afii9828log/p3log1
1/p57F(t)/p4(8.2.4) 205
between tand the cumulative distributionfunction Foftobtained from (8.2.3 ).
This relationship is linear between log tand the function log (log/p431/[1/p57F(t)]/p44).
Thus, a Weibull probabilityplot is a graph of log( t/p7/p71/p8) and log (log/p431/
[1/p57F/p19(t/p7/p71/p8)]/p44), where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8), for example, (i/p570.5)/n, for
i/p581,...,n.
The shape parameter /afii9828is estimatedgraphicallyas the reciprocal of the slope
of the straight line fitted to the graph. If the fitted line is appropriate, then atlog(log/p431/[1/p57F(t)]/p44)/p580, the corresponding log (t)is an estimate of log (1//afii9838)
from (8.2.4 ). This fact can be used to estimate 1/ /afii9838and thus /afii9838graphicallyfrom
a Weibull probabilityplot. At log (log/p431/[1/p57F(t)]/p44)/p580.5,(8.2.4 )reduces to
logt/p58log(1//afii9838)/p590.5//afii9828. This equation can be used to estimate /afii9828.
Estimatesoftheparameterscan alsobeobtainedfromthemethoddescribed
in Chapter 7 if the Weibull distribution appears to be a good fit graphically.The following hypothetical example illustrates the use of the Weibull probabil-ityplot. The small number of observations used in the example is onlyforillustrative purposes. In practice, manymore observations are needed toidentifyan appropriate theoretical model for the data.
Example 8.3 Six mice with brain tumors have survival times, in months of
3, 4, 5, 6, 8, and 10. Log( t/p7/p71/p8) plotted against log (log/p431/[1/p57(i/p570.5)/6]/p44) for
i/p581,...,6 is shown in Figure 8.6. A straight line is fitted to the data point by
eye. From the fitted line, at log (log/p431/[1/p57F(t)]/p44)/p580, the corresponding
log(t)/p581.9, and thus an estimate of 1/ /afii9838is approximately6.69 [ /p58exp(1.9)]
months and an estimate of /afii9838is 0.150. At log (log/p431/[1/p57F(t)]/p44)/p580.5, the
corresponding log (t)/p582.09, and thus an estimate of /afii9828/p580.5/(2.09 —1.9)/p582.63.
The maximum likelihood estimates of /afii9828and /afii9838obtained from the SAS
procedure LIFEREG are 2.75 and 0.148, respectively. The graphical estimatesof/afii9828and/afii9838are close to the MLE.
Lognormal Distribution
If the survival time tfollows a lognormal distribution with parameters /afii9839and
/afii9846/p17, log tfollows the normal distribution with mean /afii9839and variance /afii9846/p17.
Consequently, (logt/p57/afii9839)//afii9846has the standard normal distribution. Thus, the
lognormal distribution function can be written as
F(t)/p58/afii9818
/p1logt/p57/afii9839
/afii9846/p2t/p570( 8 .2.5)
where /afii9818(·)is the standard normal distribution function and /afii9839and /afii9846are,
respectively, the mean and standard deviation of log t.
A probabilityplot for the lognormal distribution is based on the following
relationship obtained from (8.2.5 ):
logt/p58/afii9839/p59/afii9846/afii9818/p92/p16(F(t)) (8 .2.6)206
Figure 8.6 Weibull probabilityplot of the data in Example 8.3.
The function /afii9818/p92/p16(·)is the inverse of the standard normal distribution func-
tion or its 100 Fpercentile. This relationship is linear between the value
logtand the function /afii9818/p92/p16(F(t)). Thus, a log-normal probabilityplot is a
graph of log( t/p7/p71/p8) versus /afii9818/p92/p16(F/p19(t/p7/p71/p8)), where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8).
From (8.2.6 ),a t/afii9818/p92/p16(F(t))/p580, log t/p58/afii9839; and at, /afii9818/p92/p16(F(t))/p581,/afii9846/p58logt/p57/afii9839.
These facts can be used to estimate /afii9839and/afii9846from a straight line fitted to the
graph.
Example 8.4 In a studyof a new insecticide, 20 insects are exposed.
Survival times in seconds are 3, 5, 6, 7, 8, 9, 10, 10, 12, 15, 15, 18, 19, 20, 22,25, 28, 30, 40, and 60. Suppose that prior experience indicates that the survivaltime follows a lognormal distribution; that is, some insects might react to theinsecticide veryslowlyand not die for a long time. The log( t/p7/p71/p8) versus
/afii9818/p92/p16[(i/p570.5)/20], i/p581,...,20, are plotted in Figure 8.7. The plot shows a
reasonablystraight line. From the fitted line, at /afii9818/p92/p16(F(t))/p580, log tis an
estimate of /afii9839, which is equal to 2.64, and at /afii9818/p92/p16(F(t))/p581, log t/p583.4 and thus
/afii9846/p583.4/p572.64/p580.76. /afii9818/p92/p16(F(t)) can be obtained byapply ing Microsoft Excel
function NORMSINV. 207
Figure 8.7 Lognormal probabilityplot of the data in Example 8.4.
Log-Logistic Distribution
The log-logistic distribution function is
F(t)/p58/afii9825t/p65
1/p59/afii9825t/p65t/p570,/afii9828/p570,/afii9825/p570( 8 .2.7)
A probabilityplot for the log-logistic distribution is based on the following
relationship obtained from (8.2.7 ):
logt/p581
/afii9828log/p31
1/p57F(t)/p571/p4/p571
/afii9828log/afii9825 (8.2.8 )
Thus, a log-logistic probabilityplot is a graph of log( t/p7/p71/p8) versus log (/p431/
[1/p57F/p19(t/p7/p71/p8)]/p44/p571), where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8), for example, (i/p570.5)/n,
fori/p581,...,n. From (8.2.8 ), at log /p43[1/(1/p57F)]/p571/p44/p580, log t/p58/p57(1//afii9828) log/afii9825;
and at log /p43[1/(1/p57F)]/p571/p44/p581, log t/p58(1//afii9828)(1/p57log/afii9825). These facts can be
used to estimate /afii9828and/afii9825. The following example illustrates the log-logistic
probabilityplot.
Example 8.5 Consider the following survival times of 10 experimental rats
in days: 8, 15, 25, 30, 50, 90, 95, 100, 150, and 300. Figure 8.8 plots log( t/p7/p71/p8)208
Figure 8.8 Log-logistic probabilityplot of the data in Example 8.5.
against log (/p431/[1/p57(i/p570.5)/10]/p44/p571) for i/p581,...,10. To estimate /afii9828and/afii9825,
from the fitted line, at log (/p431/[1/p57F(t)]/p44/p571)/p580, log t/p584.0; and at log (/p431/
[1/p57F(t)]/p44/p571)/p581, log t/p584.6. Thus, we have two equations:
4.0/p58/p571
/afii9828log/afii9825and 4.6 /p581
/afii9828(1/p57log/afii9825)
From these two equations, /afii9828/p24/p581.667 and /afii9825/p24/p580.0013.
8.3 HAZARD PLOTTING
Hazard plotting (Nelson 1972, 1982 )is analogous to probabilityplotting, the
principal difference being that the survival time (or a function of it )is plotted
against the cumulative hazard function (or a function of it )rather than the
distribution function. Hazard plotting is designed to handle censored data.Similar to probabilityplotting, estimates of parameters in the distribution canbe determined from the hazard plot with little computational effort.
To determine if a set of survival time with censored observation is from a
given theoretical distribution, we construct a hazard plot byplotting thesurvival time (or a function of it )versus an estimation cumulative hazard (or 209
a function of it ). The cumulativehazard function can be estimated byfollowing
the steps below.
Step 1.Orderthe nobservationsinthe samplefromsmallesttolargest without
regard to whether theyare censored. If some uncensored and censoredobservations have the same value, theyshould be listed in random order. Inthe list of ordered values, the censored data are each marked with a plus.
Step 2.Number the ordered observations in reverse order, with nassigned to
the smallest data value, n/p571 to the second smallest, and so on. The numbers
so obtained are called K values orreverse-order numbers . For the uncensored
observation, Kis the number of subjects still at risk at that time.
Step 3.Obtain the corresponding hazard value for each uncensored observa-
tion. Censored observations do not have a hazard value. The hazard value foran uncensored observationis 1 /K. This is the fraction of the Kindividuals who
survived that length of time and then failed. It is an observed conditionalfailure probabilityfor an uncensored observation.
Step 4.For each uncensored observation, calculate the cumulative hazard
value. This is the sum of the hazard values of the uncensored observation andof all preceding uncensored observations. For tied uncensored observations,the cumulative hazard is evaluated onlyat the smallest Kamong the uncen-
sored observations.
The table in the following example illustrates the procedure.Example 8.6 Consider the remission data of the 21 leukemia patients
receiving 6-MP in Example 3.3. Table 8.3 illustrates the procedure for estima-ting the cumulative hazard function.
We now discuss the basic idea underlying hazard plotting for the exponen-
tial, Weibull, lognormal, and log-logistic distributions.
Exponential Distribution
The exponential distribution has constant hazard function h(t)/p58/afii9838. Thus, the
cumulative hazard function is
H(t)/p58/afii9838t (8.3.1 )
From (8.3.1 ), the time can be written as a linear function of the cumulative
hazard H,
t/p581
/afii9838
H(t)( 8 .3.2)
Thus, tplots as a straight-line function of H. The slope of the fitted line is the210
Table 8.3 Estimation of Cumulative Hazard
Reversed Cumulative
Order, Hazard, Hazard,
tK 1/K H /p19(t)
62 10 .048
6/p59 20
61 90 .053
61 80 .056 0 .156
71 70 .059 0 .215
9/p59 16
10 15 0 .067 0 .281
10/p59 14
11/p59 13
13 12 0 .083 0 .365
16 11 0 .091 0 .456
17/p59 10
19/p59 9
20/p59 8
22 7 0 .143 0 .598
23 6 0 .167 0 .765
25/p59 5
32/p59 4
32/p59 3
34/p59 2
35/p59 1
mean survival time 1/ /afii9838of the distribution. More simply, 1/ /afii9838is the value of t
when H(t)/p581. This fact is used to estimate 1/ /afii9838from an exponential hazard
plot.
Example 8.7 Using the estimated cumulative hazard values H/p19(t) in Table
8.3, we construct the exponential hazard plot in Figure 3.5 byplotting eachexact time tagainst its corresponding H/p19(t). The configuration appears to be
reasonablylinear, suggesting that the exponential distribution provides areasonable fit. In Chapter 3 we see that the Weibull distribution gives a betterfit than the exponential. We use the data here just to demonstrate how theparameter can be estimated.
To find an estimate for the mean remission time of the leukemia patients,
we can use H(t)/p580.5 since the time for which H/p581 is out of the range of
the horizontal axis. At H(t)/p580.5,t/p5816.9, from (8.3.2 ), an estimate of
/afii9838is 0.5/16.9 /p580.0296. Thus, an estimate of the mean remission time is 34
weeks. 211
Figure 8.9 Cumulativehazard functionsofthe Weibulldistributionwith /afii9828/p580.5, 1, 2,4.Weibull Distribution
The Weibull distribution has the hazard function
h(t)/p58/afii9838/afii9828(/afii9838t)/p65/p92/p16 t/p570
The cumulative hazard function is
H(t)/p58(/afii9838t)/p65 t/p570( 8 .3.3)
and is plotted in Figure 8.9 for four different values of /afii9828: 0.5, 1, 2, and 4. From
(8.3.3 ), the time tcan be written as a function of the cumulative hazard
function, that is,
t/p581
/afii9838[H(t)]/p16/p30/p65 (8.3.4)
Taking the logarithm of (8.3.4 ), we obtain
logt/p58log1
/afii9838/p591
/afii9828logH(t)( 8 .3.5)
Since log tis a linear function of log H(t), a plot of log tagainst log H(t)i sa
straight line. For log H(t)/p580o r H(t)/p581,(8.3.5 )reduces to log t/p58log(1//afii9838),
and thus the corresponding time tequals 1/ /afii9838. This fact is used to estimate 1/ /afii9838
and consequently, /afii9838. The slope of the fitted straight line is 1/ /afii9828,o ra t
logH(t)/p581,(8.3.5 )can be written as /afii9828/p581/(logt/p59log/afii9838). This equation can
be used to estimate /afii9828.212
Figure 8.10 Weibull hazard plot of the data in Example 8.8.
Example 8.8 Consider the following survival times in months of 14
patients: 15, 25, 38, 40 /p59, 50, 55, 65, 80 /p59, 90, 140, 150 /p59, 155, 250 /p59, 252.
Figure 8.10 is the hazard plot with log tversus log H(t) of the data. From the
fitted line, at log H(t)/p580, log t/p584.8.Thus, t/p58121.5 and the estimate of /afii9838is
/afii9838/p19/p581/t/p580.0082. Similarly, at, log H(t)/p581, log t/p585.6, and thus /afii9828/p24/p581/
(5.6/p574.8)/p581.25.
Lognormal Distribution
The densityfunction of a lognormal distribution is
f(t)/p581
t/afii9846/p402/afii9843exp/p3/p571
2/afii9846/p17(logt/p57/afii9839)/p17/p4
/p581
t/afii9846g/p1logt/p57/afii9839
/afii9846/p2t/p570 (8.3.6 )
where g(x) is the standard normal densityfunction. The lognormal cumulative
distribution function is
F(t)/p58/afii9818/p1logt/p57/afii9839
/afii9846/p2t/p570( 8 .3.7) 213
Figure 8.11 Cumulative hazard functions of the lognormal distribution with /afii9846/p580.1,
0.5, 1.0.where /afii9818(·)is the standard normal distribution function. Thus, by (2.10), the
hazard function can be written as
h(t)/p581
t/afii9846g/p1logt/p57/afii9839
/afii9846/p2
1/p57/afii9818/p1logt/p57/afii9839
/afii9846/p2(8.3.8)
The cumulative hazard function, plotted in Figure 8.11 for three values of /afii9846,i s
H(t)/p58/p57log/p31/p57/afii9818/p1logt/p57/afii9839
/afii9846/p2/p4(8.3.9)
From (8.3.9 ), the logarithm of the survival time tas a function of the
cumulative hazard His
logt/p58/afii9839/p59/afii9846/afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]( 8 .3.10)
where /afii9818/p92/p16(·)is the inverse of the standard normal distribution function.
Thus, log tis a linear function of /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]. The log-normal hazard
plot is a graph of log tversus /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]. From (8.3.10 ),a t
/afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p580, log t/p58/afii9839; and at /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p581, log t/p58/afii9839/p59/afii9846.
These facts can be used to estimate /afii9839and/afii9846.
Example 8.9 Consider the following remission times in months of 18
cancer patients: 4, 5, 6, 7, 8, 9 /p59, 12, 12 /p59, 13, 15, 18, 20, 25, 26 /p59,2 8/p59, 35,
35/p59, 56. Figure 8.12 gives the log-normal hazard plot. From the fitted line by
eye, at /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p580, log t/p582.8; and at /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p581,214
Figure 8.12 Lognormal hazard plot of the data in Example 8.9.logt/p583.76. Thus, the estimate of /afii9839is 2.8 and the estimate of /afii9846is
3.76/p572.8/p580.96.
Log-Logistic Distribution
The cumulative hazard function of the log-logistic distribution is
H(t)/p58log(1/p59/afii9825t/p65)
This equation can be written as
logt/p581
/afii9828log/p43exp[ H(t)]/p571/p44/p571
/afii9828log/afii9825 (8.3.11 )
Thus, log tis a linearfunction of log /p43exp[ H(t)]/p571/p44. A log-logistic hazardplot
is a graph of log tversus log /p43exp[ H(t)]/p571/p44. From (8.3.11 ),a t
log/p43exp[ H(t)]/p571/p44/p580, log t/p58/p57(1//afii9828) log/afii9825; and at log /p43exp[ H(t)]/p571/p44/p581,
logt/p58(1//afii9828)/p57(1//afii9828) log/afii9825. These facts can be used to estimate /afii9828and/afii9825.
8.4 COX--SNELL RESIDUAL METHOD
The Cox —Snell (1968 )residual method can be applied to anyparametric
model. The Cox —Snell residual r/p71for the ith individual with observed survival
time t/p71, uncensored or censored, is defined as
r/p71/p58/p57logS/p19(t/p71) i/p581, 2, ...,n (8.4.1)— 215
where S/p19(t) is the estimated survival function based on the MLE of the
parameters.If the observed t/p71is censored, the corresponding r/p71is also censored.
Since the cumulative hazard function H(t)/p58/p57logS(t), the Cox —Snell residual
r/p71is an estimated cumulated hazard value at t/p71. The important propertyof the
Cox —Snell residual is that if the model selected fits the data, r/p71’s follow the unit
exponential distribution with densityfunction f/p48(r)/p58e/p92/p80.
Let S/p48(r) denote the survival function of the Cox —Snell residual r/p71. Then
S/p48/p58/p25/p27/p80f/p48(x)dx/p58/p25/p27/p80e/p92/p86dx/p58e/p92/p80, and
/p57logS/p48(r)/p58/p57log(e/p92/p80)/p58r (8.4.2 )
Let S/p19/p48(r) denote the Kaplan —Meier estimate of S/p48(r).It is clear from (8.4.2 )
that the plot of r/p71versus /p57logS/p19/p48(r/p71)should be a straight line with unit slope
and zero intercept if the fitted survival distribution is appropriate, regardlessof the form of the distribution.
The procedure for using Cox —Snell residuals can be summarized as follows.
1. Use the methods shown in Sections 7.1 to 7.7 to find the MLE of the
parameters of the selected theoretical distribution.
2. Calculate Cox —Snell residuals r/p71/p58/p57logS/p19(t/p71),i/p581, 2, ...,n, where S/p19(t/p71)
is the estimated survival function with the MLE of the parameters.
3. Applythe Kaplan —Meier method to estimate the survival function S/p48(r)
of the Cox —Snell residuals r/p71’s obtained in step 2, then using the estimate
S/p19/p48(r), calculate /p57logS/p19/p48(r/p71),i/p581, 2, ...,n.
4.Plot r/p71versus /p57logS/p19/p48(r/p71),i/p581, 2, ...,n. If the plot is closed to a straight
line with unit slope and zero intercept, the fitted distribution is appropri-ate.
From (8.4.1 ), if an individual survival time is right-censored, say, t/p62/p71and
the fitted model is correct, the corresponding Cox —Snell residual
/p57logS(t/p62/p71)/p58H(t/p62/p71) is smaller than the residual evaluated at an uncensored
observationwiththesame value t/p71since H(t) is a monotone-increasingfunction
oft. To take this into account, two modified Cox —Snell residuals have been
proposed for censored observations (Crowleyand Hu, 1977 ). One is based on
the mean, and the other is based on the median (/p58log2 /p580.693 )of the unit
exponential distribution byassuming that difference between H(t/p71)and H(t/p62/p71)
also follows the unit exponential distribution. For a censored observation t/p62/p71,
the modified residual r/p62/p71is defined as
r/p62/p71/p58r/p71/p591( 8 .4.3)
or
r/p62/p71/p58r/p71/p590.693 where r/p71/p58/p57logS/p19(t/p71)( 8 .4.4)
Example 8.10 Consider the tumor-free time data observed from rats fed
with saturated diets in Table 3.4. We select the lognormal distribution for this216
Figure 8.13 Cox —Snell residual plot for the fitted lognormal model on the tumor-free
time data for rats fed with saturated diets.set of data for illustrative purposes. Using methods discussed in Chapter 7, the
MLE of the parameters obtained are /afii9839/p584.76458 and /afii9846/p580.56053. We then
calculate the Cox —Snell residuals r/p71/p58/p57logS(t/p71)/p58/p57log[1 /p57F(t/p71)], where
F(t) is the distribution function of the lognormal distribution. An easywayto
compute r/p71for the lognormal distributionis to use the relationshipbetween the
normal and lognormal distributions, i.e., the distribution function of thelognormaldistribution, F(t), is equivalent to /afii9818[(logt/p57/afii9839)//afii9846], where /afii9818()is the
distribution function of the standard normal distribution. We can use Micro-soft Excel function NORMSDIST to calculate /afii9818(t). Thus, for the lognormal
distribution,
S(t/p71)/p581/p57/afii9818([log( t/p71)/p574.76458] /0.56053)
Using the specific notation of NORMSDIST, ln for log,
r/p71/p58/p57ln(1/p57normsdist /p43[ln(t/p71)/p574.76458] /0.56053 /p44)
The r/p71’s so obtained are given in Table 8.4. The next step is to obtain the
Kaplan —Meier estimate of the survival function S(r/p71), and compute /p57logS(r/p71).
These values are also given in Table 8.4.
Figure 8.13 gives the graph of r/p71versus /p57logS/p19/p48(r/p71),i/p581,...,22. The graph
is close to a straight line with unit slope and zero intercept. Therefore, a— 217
Table 8.4 Kaplan--Meier Estimate of Survivorship
Function for the Cox--Snell Residuals from the Fitted
Lognormal Model on Tumor-Free Time Data for Rats
Fed with Saturated Diets
tr /p63 S/p19/p48(r)/p64 /p57logS/p19/p48(r)
0.000 1 .000 0 .000
43 0 .037 0 .967 0 .034
46 0 .049 0 .933 0 .069
56 0 .098 0 .900 0 .105
58 0 .110 0 .867 0 .143
68 0 .181 0 .833 0 .182
75 0 .239 0 .800 0 .223
79 0 .275 0 .767 0 .266
81 0 .294 0 .733 0 .310
86 0 .342 0 .667 0 .405
86 0 .342 0 .667 0 .405
89 0 .373 0 .633 0 .457
96 0 .447 0 .600 0 .511
98 0 .469 0 .567 0 .568
105 0 .548 0 .533 0 .629
107 0 .571 0 .500 0 .693
110 0 .606 0 .467 0 .762
117 0 .690 0 .433 0 .836
124 0 .776 0 .400 0 .916
126 0 .800 0 .367 1 .003
133 0 .889 0 .333 1 .099
142 1 .004 0 .267 1 .322
142 1 .004 0 .267 1 .322
165 1 .305 0 .233 1 .455
170/p59 1.371/p59
200/p59 1.769/p59
200/p59 1.769/p59
200/p59 1.769/p59
200/p59 1.769/p59
200/p59 1.769/p59
200/p59 1.769/p59
/p63r, ordered Cox —Snell residuals from the fitted lognormal model.
/p64S/p48(r), Kaplan —Meier estimate of survivorship function for the
Cox —Snell residuals.
lognormal model maybe appropriate for the tumor-free times observed. In
Chapter 9 (Example 9.2 )we will see that the lognormal model was not rejected
based on a goodness-of-fit test. Thus the result is consistent with thoseobtainedbyusing the analy ticalmethod.A weakness of theCox —Snell residual
method is that the plot does not indicate the kind of departure the data havefrom the model selected if the configuration is not linear.218
Bibliographical Remarks
Probabilityplotting has been widelyused since Daniel’s (1959 )classical work
on the use of half-normal plot. A quite complete and excellent treatment ofprobabilityplotting is given byKing (1971 ). Although examples given are
applications to industrial reliability, its interpretation of probability plots ofmanydistributions, such as the uniform, lognormal, Weibull, and gamma, areapplicable to biomedical research. Recent applications of probabilityplotting
include Leitner et al. (1986 ), Horner (1987 ), Waters et al. (1991 ), and
Tsumagari et al. (2000 ).
Hazard plotting was developed byNelson (1972, 1982 ). Applications in-
cluded Gore (1983 )and Wurpel et al. (1986 ).
EXERCISES
8.1Show that the Cox —Snell residuals defined in (8.4.1 )follow the unit
exponential distribution with densityfunction f(r)/p58exp(/p57r).
8.2Consider the following survival times of 16 patients in weeks: 4, 20, 22,
25, 38, 38, 40, 44, 56, 83, 89, 98, 110, 138, 145, and 27.
(a)Does the exponential distribution provide a reasonable fit to the
survival data? Use the probabilityplotting technique.
(b)Estimate graphicallythe parameter /afii9838of the exponential distribution
and consequently, the mean survival time.
8.3To computerize patients’ records, a data clerk is hired to transcribe
medical data from the patients’ charts to computer coding forms. Thenumber of correct entries between errors is listed in chronological orderof occurrence over a period of five days as follows: 73, 12, 40, 65, 100,15, 70, 40, 110, 64, 200, 6, 90, 102, 20, 102, 90, 34. The assumption is thatthe data clerk, during the five days, would not change her error rateappreciably. Use the technique of probability plotting to evaluate theassumption above. What is your conclusion?
8.4Twenty-five rats were injected with a give tumor inoculum. Their times,
in days, to the development of a tumor of a certain size are given below.
30 53 77 91 118
38 54 78 95 12045 58 81 101 12546 66 84 108 13450 69 85 115 135
Which of the distributions discussedin this chapter providea reasonable
fit to the data? Estimate graphicallythe parameters of the distribution
chosen. 219
8.5In a clinical study, 28 patients with cancer of the head and neck did not
respondtochemotherapy.Theirsurvivaltimesinweeksaregivenbelow.
1.7 8.3 14.0 22.7 6.0 /p5913.1/p59
5.1 9.6 15.9 33.0 7.4 /p5913.4/p59
5.3 11.3 16.7 3.7 /p598.0/p5916.1/p59
6.0 12.1 17.0 5.0 /p598.3/p59
8.3 12.3 21.0 5.9 /p599.1/p59
(a)Make a hazard plot for each of the following distributions:exponen-
tial, Weibull, lognormal, and log-logistic.
(b)Which distribution provides a reasonable fit to the data? Estimate
graphicallythe parameters of the distribution chosen.
8.6Thirty-one patients with advanced melanoma treated with combined
chemotherapy, immunotherapy, and hormonal therapy have survivaltimes as given below.
26.3/p5916.1 24.0 4.3 31.3 /p59
94.0 49.6 77.9 97.6 /p5917.6/p59
9.1 27.3 16.6 /p597.3 16.3
34.6/p5961.9/p593.4 75.6 /p59
9.4 46.6 /p5910.9 14.3
25.7 22.4 /p5913.0 56.4
88.7 7.1 64.4 /p599.1
(a)Make a hazard plot for each of the following distributions:exponen-
tial, Weibull, lognormal, and log-logistic.
(b)Which distribution provides a reasonable fit to the data? Estimate
the parameters of the distribution chosen.
8.7Consider the survival times of the hypernephroma patients in Exercise
Table 3.1 (see Exercise 4.5 ). Make a hazard plot for the distribution you
chose in Exercise 6.8. Did you make a good selection? If not, try twoother distributions.
8.8Consider the following survival times in weeks of 10 mice with injection
of tumor cells: 5, 16, 18 /p59, 20, 22 /p59,2 4/p59, 25, 30 /p59, 35, 40 /p59. Make an
exponential hazard plot. Does the exponential distribution provide areasonable fit? If not, is the lognormal distribution better?
8.9Consider the following survival times in months of 25 patients with
cancer of the prostate. Use a graphical method to see if the survival timeof prostate cancer patients follows the exponential distribution with/afii9838/p580.01: 2, 19, 19, 25, 30, 35, 40, 45, 45, 48, 60, 62, 69, 89, 90, 110, 145,
160, 9 /p59,1 0/p59,2 0/p59,4 0/p59,5 0/p59, 110/p59, 130/p59.
8.10Make a log-logistic hazard plot of the following data and estimate the
two parameters: 20, 30, 32 /p59, 40, 60, 100, 150, 200 /p59, 300.220
CHAPTER 9
Tests of Goodness of Fit
and Distribution Selection
In Chapter 8 we discuss three graphical methods for checking if a parametricdistribution fits the observed data. Parametric distributions can be groupedintofamilies.First, anygivendistributionwithdifferentparametervaluesformsa family. Second, if a distribution includes other distributions as its specialcases, this distribution is a nesting (larger )family of these distributions. For
example, the distributions introduced in Chapter 6 belongto more than onenested family. First, the Weibull distribution reduces to the exponential when/afii9828/p581. Therefore, the exponential distribution is a special case of the Weibull
and the two distributions are said to belongto one family, the Weibull family.Second, consider the standard gamma distribution; when /afii9828/p581, it reduces to
the exponential, and when /afii9838/p58/p16/p17
and/afii9828/p58/p16/p17/afii9840, it becomes the chi-square
distribution with /afii9840degrees of freedom. Thus, the gamma distribution includes
the exponential and chi-square as a family. Now let us consider the generalizedgamma distribution. It reduces to the exponential if /afii9825/p58/afii9828/p581, the Weibull if
/afii9828/p581, the lognormal if /afii9828/p59/p45, and the gamma if /afii9825/p581. Thus, the generalized
gamma distribution includes these four distributions and represents a largefamily of distributions. The relationship of the generalized gamma distributionto the exponential, Weibull, lognormal, and gamma distributions allows us toevaluate the appropriateness of these distributions relative to each other andto a more general distribution. It is known that the generalized gammadistribution is a special case of the generalized F-distribution and therefore
belongs to the generalized Ffamily (Kalbfleisch and Prentice, 1980 )Because
of its complexity, we do not cover the generalized Ffamily.
In this chapter we discuss several analytical procedures for comparing
parametric distributions and assessingg oodness of fit. In Section 9.1 weintroduce several widely used statistics for testingthe appropriateness of adistribution. Readers who are not familiar with linear algebra or are notinterested in the mathematical details may skip this section without loss ofcontinuity.In Section 9.2 we discuss statistics for testingwhether a distribution
221
is appropriate by comparingit with other distributions in the same family or
a more general family. Section 9.3 covers the selection of a distribution basedon Baysian information criteria. Section 9.4 covers the statistics for testingwhethera given distributionwith knownparameters is appropriate. Allthe teststatistics discussed in Sections 9.1 to 9.4 are based on asymptotic likelihoodinferences. In Section 9.5 we introduce the test statistic of Hollander andProschan (1979 )for testingwhether a distribution with g iven parameters is
appropriate. Computer codes for BMDP or SAS that can be used to carry out
the test procedures are provided.
9.1 GOODNESS-OF-FIT TEST STATISTICS BASED ON
ASYMPTOTIC LIKELIHOOD INFERENCES
We take the exponential distribution as an example to see how to construct
statistics to test whether it is appropriate for the observed survival times. Asnoted in Chapter 6, the Weibull family with /afii9828/p581, the gamma family with
/afii9828/p581, and the generalized gamma family with /afii9825/p58/afii9828/p581 reduce to the
exponential distribution. Therefore, to test if the exponential distribution isappropriate for the observed survival time, we can first fit a Weibull distribu-tion and test if /afii9828/p581, or fit a gamma distribution, then test if /afii9828/p581, or fit a
generalized gamma distribution, then test if /afii9825/p58/afii9828/p581. Similarly, to test
whether the family of Weibull distributions, or the gamma distributions, or thelognormal distributions is appropriate for the survival data observed, we canfit a generalized gamma distribution (their nestingdistribution )and then test
if/afii9828/p581, or/afii9825/p581, or with /afii9828/p59/p45,respectively.Thus,testingtheappropriateness
of a family of distributions is equivalent to testingwhether a subset of theparameters in its nestingdistribution equal to some specific values. If the datacan be assumed to follow a certain distribution but the values of its parametersare uncertain, we need to test only that the parameters are equal to certainvalues. In the following, we separately introduce test statistics for testingwhether some of the parameters in a distribution are equal to certain valuesand whether all parameters in a distribution are equal to certain values.Readers who are interested in a detailed discussion of these statistics arereferred to Kalbfleisch and Prentice (1980 ).
9.1.1 Testing a Subset of Parameters in a Distribution
Letb/p58(b/p16,b/p17)denote all the parameters in a parametric distribution, where
b/p16andb/p17are subsets of parameters, and let the hypothesis be
H/p15:b/p17/p58b/p15(9.1.1 )
whereb/p15is a vector of specific numbers. Let b/p19be the MLE of b,b/p19/p16(b/p15)the
MLE of b/p16givenb/p17/p58b/p15, andV/p19/p17(b/p19)the submatrix of the covariance matrix in222
(7.1.5 ),V/p19(b/p19), correspondingto b/p17. Under H/p15and some mild assumptions, both
of the followingtwo statistics have an asymptotic chi-square distribution withdegrees of freedom equal to the dimension of (or the number of parameters in )
b/p17.
Log-likelihood ratio statistic:
X/p42/p582[l(b/p19)/p57l(b/p19/p16(b/p15),b/p15)] (9.1.2 )
Wald statistic:
X/p53/p58(b/p19/p17/p57b/p15)/p30V/p19/p92/p16/p17(b/p19)(b/p19/p17/p57b/p15)( 9.1.3 )
If the number of parameters in b/p17is equal to q, for a given significant level
/afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p79/p11/p63when the likelihood ratio statistic is used; or if
X/p53/p57/afii9851/p17/p79/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p79/p11/p16/p92/p63/p30/p17,(two-sided test )orX/p53/p57/afii9851/p17/p79/p11/p63(one-sided test )
when the Wald’s statistic is used, where /afii9851/p17/p79/p11/p63,/afii9851/p17/p79/p11/p63/p30/p17and/afii9851/p17/p79/p11/p16/p92/p63/p30/p17are the
100(1/p57/afii9825), 100 (1/p57/afii9825/2), and 100 /afii9825/2 percentile points of the chi-square dis-
tribution with qdegrees of freedom; that is,
P(/afii9851/p17/p79/p57/afii9851/p17/p79/p11/p63)/p58/afii9825andP(/afii9851/p17/p79/p57/afii9851/p17/p79/p11/p63/p30/p17)/p58P(/afii9851/p17/p79/p58/afii9851/p17/p79/p11/p16/p92/p63/p30/p17)/p58/afii9825
2
Example 9.1 Suppose that we wish to test whether the observed data are
from an exponential distribution. We can use a Weibull distribution and testwhether its shape parameter, /afii9828, is equal to 1. The Weibull distribution has two
parameters, /afii9838and/afii9828; thusb/p58(/afii9838,/afii9828)and the null and alternativehypotheses are:
H/p15:/afii9828/p581(the underlyingdistribution is an exponential distribution )
(9.1.4 )
H/p16:/afii9828/p341(the underlyingdistribution is a Weibull distribution )
Letb/p19/p58(/afii9838/p19,/afii9828/p19)be the MLE of b,l/p53(b/p19)/p58l/p53(/afii9838/p19,/afii9828/p19) andl/p35(/afii9838/p19)be the log-likelihood
of the Weibull and exponential distributions, respectively, l/p35(/afii9838/p19)/p89l/p53(/afii9838/p19(1),1),
where /afii9838/p19(1)is the MLE of /afii9838in the Weibull distribution given /afii9828/p581. The
log-likelihoodratio andWald statisticsdefined in (9.1.2 )and(9.1.3 )in this case
become
X/p42/p582[l/p53(/afii9838/p19,/afii9828/p19)/p57l/p53(/afii9838/p19(1),1)] (9.1.5 )
and
X/p53/p58(/afii9828/p24/p571)V/p19/p92/p16/p17(/afii9838/p19,/afii9828/p24)(/afii9828/p24/p571) (9 .1.6) -- 223
respectively, where V/p19/p17(/afii9838/p19,/afii9828/p24) is the second diagonal element of the covariance
matrix
V/p19(/afii9838/p19,/afii9828/p24)/p58/p57/p1/p42/p17l/p53(/afii9838/p19,/afii9828/p24)
/p42/afii9838/p17/p42/p17l/p53(/afii9838/p19,/afii9828/p24)
/p42/afii9838/p42/afii9828
/p42/p17l/p53(/afii9838/p19,/afii9828/p24)
/p42/afii9828/p42/afii9838/p42/p17l/p53(/afii9838/p19,/afii9828/p24)
/p42/afii9828/p17/p2/p92/p16
(9.1.7)
and
V/p19/p92/p16/p17(/afii9838/p19,/afii9828/p24)/p58/p57[/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9838/p17][/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9828/p17]/p57(/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9838 /p42/afii9828)/p17
/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9838/p17(9.1.8)
For a given significant-level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63, when the likelihood
ratio statistic is used; or if X/p53/p57/afii9851/p17/p16/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p16/p11/p16/p92/p63/p30/p17, when the Wald
statistic is used.
It must be pointed out that failure to reject H/p15in(9.1.4 )does not imply that
anexponentialdistributionprovidesthe bestfit to thedata. On the other hand,rejection of H/p15does not indicate that a Weibull distribution is the choice
either. Further testingof other distributions is needed. The details andexamples are given in Section 9.2.
Since the gamma and generalized gamma distribution also include the
exponential as a special case, similar test statistics can be constructed to testthe null hypothesisthat the data are from the exponentialdistributionby usingthe gamma, the generalized gamma, or the extended generalized gammadistribution.
9.1.2 Testing All Parameters in a Distribution
To test whether all of the parameters in bequal a given set of known values
b/p15, the null hypothesis is
H/p15:b/p58b/p15(9.1.9 )
and the followingthree test statistics can be used.Log-likelihood ratio statistic:
X/p42/p582[l(b/p19)/p57l(b/p15)] (9.1.10 )
Wald statistic:
X/p53/p58/p57(b/p19/p57b/p15)/p30/p42/p17l(b/p15)
/p42b/p42b/p30
(b/p19/p57b/p15) /p3or/p58/p57(b/p19/p57b/p15)/p30/p42/p17l(b/p19)
/p42b/p42b/p30(b/p19/p57b/p15)/p4
(9.1.11 )224
Score statistic:
X/p49/p58/p3/p42l(b/p15)
/p42b/p4/p30/p3/p57/p42/p17l(b/p15)
/p42b/p42b/p30/p4/p92/p16/p42l(b/p15)
/p42b /p1or/p58/p3/p42l(b/p15)
/p42b/p4/p30V/p19(b/p19)/p42l(b/p15)
/p42b/p2
(9.1.12 )
whereV/p19(b/p19)is the estimated covariance matrix in (7.1.5 ). Under H/p15and the
assumption that b/p19has approximately multinormal distribution, each of the
three statistics has an asymptotic chi-square distribution with p(the dimension
ofbor the number of parameters in b)degrees of freedom.
For a given significant-level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p78/p11/p63, when the
likelihood ratio statistic is used; or if X/p53/p57/afii9851/p17/p78/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p78/p11/p16/p92/p63/p30/p17, when the
Wald statistic is used;or if X/p49/p57/afii9851/p17/p78/p11/p63/p30/p17orX/p49/p58/afii9851/p17/p78/p11/p16/p92/p63/p30/p17, when the score statistic
is used.
It must be pointed out that rejection of H/p15in(9.1.9 )means only that the
given distribution with the known parameters b/p15, not the family of distribu-
tions to which the given distribution belongs, is not appropriate for theobserved data. It is possible that a distribution with different b/p15in the family
may be appropriate.
9.2 TESTS FOR APPROPRIATENESS OF A FAMILY OF
DISTRIBUTIONS
The usual method for testingwhether a distribution is appropriate for the
observed data is to compare the distribution with a larger or more generalfamily that includes the distribution of interest as a special case (Hagar and
Bain, 1970 ).
Letl/p35(/afii9838),l/p53(/afii9838,/afii9828),l/p37(/afii9838,/afii9828),l/p42/p44(/afii9839,/afii9846/p17), andl/p37/p37(/afii9825,/afii9838,/afii9828) denote, respectively,
the log-likelihoodfunction definedin (7.1.1 )based on the exponential,Weibull,
gamma, lognormal, and extended generalized gamma distribution, and l/p35(/afii9838/p19),
l/p53(/afii9838/p19,/afii9828/p24),l/p37(/afii9838/p19,/afii9828/p24),l/p42/p44(/afii9839/p24,/afii9846/p24/p17), andl/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)denote the respective log-likelihood
values where /afii9838/p19,(/afii9838/p19,/afii9828/p24),(/afii9838/p19,/afii9828/p24),(/afii9839/p24,/afii9846/p24/p17), and (/afii9825/p24,/afii9838/p19,/afii9828/p24)are the MLE. For example, the
log-likelihood of the exponential distribution can be obtained from
l/p35(/afii9838/p19)/p58/p80/p26
/p71/p14/p16log(/afii9838/p19e
/p57/afii9838/p19t/p71)/p59/p76/p26
/p71/p14/p80/p62/p16log(e/p57/afii9838t/p62/p71)/p58rlog/afii9838/p19/p57/afii9838/p19/p80/p26
/p71/p14/p16t/p71/p57/afii9838/p19/p76/p26
/p71/p14/p80/p62/p16t/p62/p71
for a set of observed survival times t/p16,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76. The log-likelihood
value and the estimated covariance matrix in (7.1.5 )and parameters for each
of the distributions discussed in Sections 7.2 to 7.6 can be obtained from SASorBMDP.Theresults canbe used toconstructthelog-likelihoodratiostatisticand the Wald statistic defined in (9.1.2 )and (9.1.3 ). In the following, we 225
introduceseveral testsfor the appropriatenessofa familyofdistributionsbased
on the log-likelihoods. Construction of the respective Wald statistics is left tothe reader as exercises.
1.Testing the hypothesis that the underlying distribution is exponential . The
null hypothesis is
H/p15:The underlyingdistribution is an exponential distribution
If the Weibull distribution is used, testingthe null hypothesis above is
equivalent to testingthe followingnull and alternative hypotheses:
H/p15:/afii9828/p581(the underlyingdistribution is an exponential distribution )
H/p16:/afii9828/p341(the underlyingdistribution is a Weibull distribution )
Let/afii9838/p19(1)be the MLE of /afii9838in the Weibull distribution given /afii9828/p581, the
log-likelihood ratio statistic is
X/p42/p582[l/p53(/afii9838/p19,/afii9828/p24)/p57l/p53(/afii9838/p19(1),1)] (9.2.1 )
which has an asymptotic chi-square distribution with 1 degree of freedom. For
a given level of significance /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63. Note that
l/p53(/afii9838/p19(1),1)/p89l/p35(/afii9838/p19).
Similarly, a log-likelihood ratio statistic can be constructed by using the
gamma or the extended generalized gamma distribution. These will be left tothe reader as exercises.
2.Testing the hypothesis that the underlying distribution is Weibull . The null
hypothesis is
H/p15:The underlyingdistribution is a Weibull distribution
We can use the extended generalized gamma distribution and test whether its
parameter /afii9828equals 1. Thus the null and alternativehypotheses can be stated as
H/p15:/afii9828/p581(the underlyingdistribution is a Weibull distribution )
H/p16:/afii9828/p341(the underlyingdistribution is an extended g eneralized
gamma distribution )
Let/afii9838/p19(1)and/afii9825/p24(1)be the MLE of /afii9838and/afii9825in the extended generalized gamma
distribution given /afii9828/p581. Accordingto Section 6.4, an extended g eneralized226
gamma distribution with /afii9828/p581 is a Weibull distribution. The likelihood ratio
statistic is
X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p37/p37(/afii9825/p24(1),/afii9838/p19(1), 1)] (9 .2.2)
which follows asymptotically the chi-square distribution with 1 degree of
freedom. H/p15is rejected at a significance level of /afii9825ifX/p42/p57/afii9851/p17/p16/p11/p63. Note that
l/p37/p37(/afii9825/p24(1),/afii9838/p24(1), 1) /p89l/p53(/afii9838/p19,/afii9828/p24).
3.Testing the hypothesis that the underlying distribution is standard gamma .
The null hypothesis is
H/p15:The underlyingdistribution is a g amma distribution
Followingthe same log ic in Section 6.4, the null hypothesis above is equivalent
to the following if the extended generalized gamma distribution is used.
H/p15:/afii9825/p581(the underlyingdistribution is a standard g amma distribution )
H/p16:/afii9825/p341(the underlying distribution is a generalized gamma distribution ).
The likelihood test statistic is
X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p37/p37(1,/afii9838/p19(1),/afii9828/p24(1))] (9 .2.3)
where /afii9828/p24(1)and/afii9838/p19(1)are the MLE of /afii9828and/afii9838given /afii9825/p581, which has an
asymptotic chi-square distribution with 1 degree of freedom under H/p15. The
rejection rule is the same as that for the exponential or Weibull distribution.Note that l/p37/p37(1,/afii9838/p19(1),/afii9828/p24(1))/p89l/p37(/afii9838/p19,/afii9828/p24).
4. Testing the hypothesis that the underlying distribution is lognormal . The
null hypothesis is
H/p15:the underlyingdistribution is a log normal distribution
The log-likelihood test statistic is
X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p42/p44(/afii9839/p24,/afii9846/p24/p17)]
which has an asymptotic chi-square distribution with 1 degree of freedom
underH/p15. The rejection rule is the same as that for the exponential or Weibull
distribution.
For the log-logistic and extended generalized gamma distributions, it can be
shown that a generalized F-distribution (Kalbfleisch and Prentice, 1980 )
includes the exponential, Weibull, lognormal, gamma, generalized gamma, 227
Table 9.1 Summary of Goodness-of-Fit Tests for
Testing Whether a Family of Models Is Appropriate forthe Observed Data /p63
Hypothesized
Model LL X/p42df
Generalized gamma l/p37/p37Lognormal l/p42/p442(l/p37/p37/p57l/p42/p44)1
Gamma l/p372(l/p37/p37/p57l/p37)1
Weibull l/p532(l/p37/p37/p57l/p53)1
Exponential l/p352(l/p37/p37/p57l/p35)2
Exponential l/p352(l/p37/p57l/p35)1
Exponential l/p352(l/p53/p57l/p35)1
/p63LL, log-likelihood; X/p42, likelihood ratio chi-square statistic; df,
degrees of freedom.
extended generalized gamma, and log-logistic distributions as special cases.
Therefore, one can follow the same logic to construct either the log-likelihoodratio or the Wald statistic to test the appropriateness of a family of generalizedgamma or log-logistic distributions. However, methods for testing the appro-priateness of a generalized F-distribution remain unknown. Unless we can find
a more general distribution that includes the generalized F-distribution as a
special case, there is no formal way to check whether the generalized F-
distribution is appropriate. However, the generalized gamma distribution is arich family and includes a considerable number of distributions. It should besufficient for most applications. All the tests introduced in this section aresummarized in Table 9.1.
As pointed out in Section 9.1, when usingany of the testingprocedures
above, failure to reject H/p15does not imply that the hypothesized distribution
provides a perfect fit to the data. On the other hand, rejection of H/p15does not
mean that the distribution under the alternative hypothesis is the best choiceeither. In practice, with the help of available computer software, it is easy to fitseveral distributions simultaneously and then select the most appropriate one,usually the simplest one, as the final choice for the data. The followingexamples illustrate the procedure.
Example 9.2 Consider the tumor-free times of the 30 rats that are fed with
a saturated diet in Table 3.4. UsingSAS, we obtainthe MLE of the parametersand the log-likelihoods for the exponential, Weibull, lognormal, and generaliz-ed gamma distributions. The results are given in Table 9.2. For example, theMLE of /afii9838in the exponential distribution is 5.054 and the corresponding
log-likelihood is /p5735.359, and the MLE of the two parameters in the Weibull
distribution are /afii9838/p19/p585.002 and /afii9828/p24/p580.500 and the correspondinglog -likelihood228
Table 9.2 Goodness-of-Fit Tests Based on Asymptotic Likelihood Inference for the Tumor-Free Time Data from Rats
F e dw i t hS a t ur a t e dD i e ti nT a b l e3 . 4 /p63
Estimated Parameters
Model A /p64 B/p65 C/p66 LL X/p42pValue BIC AIC
Exponential 5.054 — — /p5735.359 11.922 /p67 /p580.001 — —
Exponential 5.054 — — /p5735.359 19.762 /p68 /p580.001 /p5737.060 /p5737.359
Weibull 5.002 0.500 — /p5729.398 7.840 /p680.005 /p5732.800 /p5733.398
Lognormal 4.765 0.561 — /p5726.641 2.326 /p680.127 /p5730.042 /p5730.641
Extended generalized 4.495 0.527 /p571.088 /p5725.478 — — /p5730.580 /p5731.478
gamma
Log-logistic 4.739 0.332 — /p5726.867 — — /p5730.268 /p5730.867
/p63LL, log-likelihood; X/p42,l i k e l i h o o dr a t i os t a t i s t i c ; pvalue,P(X/p17/p57X/p42).
/p64A/p58/p57log/afii9838for the exponential and the extended generalized gamma, /p58/p57 (1//afii9828)log/afii9838for the Weibull, /p58/afii9839for the lognormal, and
/p58/p57 (1//afii9828)log/afii9825for the log-logistic distribution.
/p65B/p581//afii9828for the Weibull and log-logistic, /p58/afii9846for the lognormal, and /p581/(/afii9825/afii9828/p15/p13/p20)for the extended generalized gamma distribution.
/p66C/p581//afii9828for the extended generalized gamma distribution.
/p67Relative to the Weibull.
/p68Relativew to the extended generalized gamma.
229
is/p5729.398. To test the null hypothesis that the underlyingdistribution is an
exponential distribution versus the alternative hypothesis that the underly-ingdistribution is Weibull (or extended generalized gamma ), the likelihood
ratio test statistic X/p42/p582(35.359 /p5729.398 )/p5811.922 [or 2 (35.359 /p5725.478 )/p58
19.762]. The probability of observingsuch a chi-square value is /p580.001;
therefore, the exponential distribution is rejected and the Weibull or thegeneralized gamma is preferred.
However, the Weibull distribution is also rejected at the 0.001 level relative
to the extended generalized gamma distribution (X/p42/p587.840,p/p580.001). This
implies that the extended generalized gamma distribution may be better.However, the extended generalized gamma distribution is not significantlybetter than the lognormal distribution (X/p42/p582.326,p/p580.127 ). Thus, among
these distributions, the lognormal and extended generalized gamma distribu-tions are our choices. Because of its simplicity, we may select the lognormaldistribution as the choice for this set of data.
Example 9.3 Table 9.3 contains a set of remission times from 137 cancer
patients. These remission times are a subset of the data from a bladder cancerstudy and are used here only for illustrative purposes. The results of goodnessof fit tests based on asymptotic likelihood inferences are shown in Table 9.4.From this table, we see that the exponential distribution is not rejected relativeto the Weibull distribution based on the statistic defined in (9.2.1 )(X/p42/p580.638,
p/p580.425 ). The hypothesis that the underlyingdistribution is exponential
versus the alternative hypothesis that the distribution is the extended general-ized gamma is rejected (X/p42/p586.772,p/p580.034 ). Furthermore, the Weibull and
lognormal distributions are also rejected in favor of the extended generalizedgamma (X/p42/p586.135, 8.120,p/p580.013 and 0.004, respectively ). This implies that
the exponential distribution may not be an appropriate distribution since theWeibull distribution (its nestingdistribution )is rejected. Therefore, we may
accept the extended generalized gamma as our final choice of distribution forthe data.
9.3 SELECTION OF A DISTRIBUTION USING BIC OR AIC
PROCEDURES
The test procedures discussed in Section 9.2 require knowledge of the distribu-
tion family to which the distribution of interest belongs. In this section weintroduce a simpler selection procedure called the Baysian information criterion
(BIC; Schwarz, 1978 ). This criterion is based on the log-likelihood l(b/p19), the
number of parameters in the distribution (p), and the total number of
observations (n). For each candidate distribution, compute
r/p58l(b/p19)/p57p
2
logn (9.3.1)230
Table 9.3 Remission Times (Months) of 137 Cancer
Patients
tt t t
4.50 32.15 3.88 13.80
19.13 4.87 3.02 /p59 5.85
14.24 5.71 19.36 /p59 7.09
7.87 7.59 20.28 5.32
5.49 3.02 46.12 4.33 /p59
2.02 4.51 5.17 2.839.22 1.05 0.20 8.373.82 9.47 36.66 14.77
26.31 79.05 10.06 8.53
4.65/p59 2.02 4.98 11.98
2.62 4.26 5.06 1.76
0.90 11.25 16.62 4.40
21.73 10.34 12.07 34.26
0.87/p59 10.66 6.97 2.07
0.51 12.03 0.08 17.123.36 2.64 1.40 12.63
43.01 14.76 2.75 7.66
0.81 1.19 7.32 4.183.36 8.66 1.26 13.291.46 14.83 6.76 23.63
24.80 /p59 5.62 8.60 /p59 3.25
10.86 /p59 18.10 7.62 7.63
17.14 25.74 3.52 2.87
15.96 17.36 9.74 3.31
7.28 1.35 0.40 2.264.33 9.02 5.41 2.69
22.69 6.94 2.54 11.79
2.46 7.26 2.69 5.34
3.48 4.70 /p59 8.26 6.93
4.23 3.70 0.50 10.756.54 3.64 5.32 13.118.65 3.57 5.09 7.395.41 11.64 2.092.23 6.25 7.93
4.34 25.82 12.02
whereb/p19denotes the MLE of all the parameters in the distribution.
The candidate distribution with the largest rvalue is the distribution that
fits the data the best. It has been shown that for some distribution familiesand under mild assumptions, for sufficiently large n, the distribution select-
ed by the BIC procedure approaches the true underlyingdistribution, ifit exists. 231
Table 9.4 Goodness-of-Fit Tests Based on Asymptotic Likelihood Inference for Data in Table 9.3 /p63
Estimated Parameters
Model A /p64 B/p65 C/p66 LL X/p42pValue BIC AIC
Exponential 2.303 — — /p57198.234 0.638 /p670.425 — —
Exponential 2.303 — — /p57198.234 6.772 /p680.034 /p57200.694 /p57200.234
Weibull 2.322 0.949 — /p57197.915 6.134 /p680.013 /p57202.835 /p57201.915
Lognormal 1.821 1.090 — /p57198.908 8.120 /p680.004 /p57203.828 /p57202.908
Extended generalized 2.087 0.999 0.520 /p57194.848 — — /p57202.228 /p57200.848
gamma
Log-logistic 1.866 0.591 — /p57195.344 — — /p57200.264 /p57199.344
/p63LL, log-likelihood; X/p42,l i k e l i h o o dr a t i os t a t i s t i c ; pvalue,P(X/p17/p57X/p42).
/p64/p11/p65/p11/p66See the footnotes in Table 9.2.
/p67Relative to the Weibull.
/p68Relative to the extended generalized gamma.
232
In general,thelargerthenumberofparameters pinadistribution,thelarger
thelog-likelihood l(b/p19)in(9.3.1 ).Thus thefirst term representsthe gain by using
a distribution with more parameters. But the larger the p, the larger the second
term in (9.3.1 )is, which represents a penalty by havingmore parameters in the
distribution. Therefore, the BIC provides a balance between the gain and thepenalty.
Another widely used criterion is called an information criterion (AIC;
Akaike, 1969 ), in which ris defined as
r/p58l(b/p19)/p572p (9.3.2)
Example 9.4 The values of the BIC and AIC for the various distributions
considered in Examples 9.2 and 9.3 are listed in the last two columns in Tables9.2 and 9.4. Based on Table 9.2, the lognormal distribution would be selectedby either the BIC or AIC procedure, which is consistent with the resultsobtained in Example 9.2. The results in Table 9.4 show that the log-logisticdistribution, rather than the extended generalized gamma distribution, shouldbe selected based on either the BIC or AIC procedure.
9.4 TESTS FOR A SPECIFIC DISTRIBUTION WITH KNOWN
PARAMETERS
In this section we introduce the likelihood ratio statistic for testingif the
survival data observed follow a given distribution with known parameters. Weuse the same notations as in Section 9.2. In addition to the exponential,Weibull,lognormal,gamma,generalizedgamma distributions,wealso considerthe log-logistic distribution. Let l/p42/p42(/afii9825,/afii9828) andl/p42/p42(/afii9825/p24,/afii9828/p24) denote its log-likelihood
function and the log-likelihood with (/afii9825/p24,/afii9828/p24), the MLE of (/afii9825,/afii9828).
1.Testing the hypothesis that the underlying distribution is exponential with
known parameter /afii9838/p15. The null hypothesis is
H/p15:the underlyingdistribution is the exponential distribution with /afii9838/p58/afii9838/p15
The likelihood ratio test statistic based on (9.1.10 )is
X/p42/p582[l/p35(/afii9838/p19)/p57l/p35(/afii9838/p15)] (9 .4.1)
X/p42has an asymptotic chi-square distribution with 1 degree of freedom under
H/p15.H/p15is rejected if X/p42/p57X/p17/p16/p11/p63, where /afii9825is the significance level. Similarly, the
Wald test statistic and the score statistic can be derived by following (9.1.11 )
and (9.1.12 ). This is left to the reader as exercises.
Example 9.5 Consider the followingsurvival times in weeks of 10 mice
with a given tumor: 1, 3, 5, 8, 10 /p59, 15, 18, 19, 22, 25 /p59. We test the following 233
null hypothesis:
H/p15:the underlyingdistribution of the observed data is
exponential with /afii9838/p580.06
In this case, n/p5810,r/p588,/afii9814/p80/p71/p14/p16t/p71/p5891, and /afii9814/p76/p71/p14/p80/p62/p16t/p62/p71/p5835. The MLE of /afii9838
based on (7.2.16 )is
/afii9838/p19/p588
91/p5935/p580.0635
andl/p35(/afii9838/p19)/p588(log0.0635 )/p570.0635 (91)/p570.0635 (35)/p58/p5730.055. Under H/p15,
l/p35(/afii9838/p15)/p588(log0.06 )/p570.06(91)/p570.06(35)/p58/p5730.067. Thus, following (9.4.1 ),
X/p42/p582[/p5730.055/p57(/p5730.067)] /p580.024.X/p17/p16/p11/p15/p13/p15/p20/p583.84; therefore, we cannot
reject the null hypothesis that the data are from the exponential distributionwith/afii9838/p580.06.
2.Testing the hypothesis that the underlying distribution is Weibull with
known parameters /afii9838/p15and/afii9828/p15. The null hypothesis is
H/p15: the underlyingdistribution is Weibull with known parameters /afii9838/p58/afii9838/p15and/afii9828/p58/afii9828/p15
Based on (9.1.10 ), the likelihood ratio test statistic is
X/p42/p582[l/p53(/afii9838/p19,/afii9828/p24)/p57l/p53(/afii9838/p15,/afii9828/p15)] (9 .4.2)
UnderH/p15,X/p42has an asymptotic chi-square distribution with 2 degrees of
freedom. H/p15is rejected if X/p42/p57X/p17/p17/p11/p63where /afii9825is the significance level.
3.Testing the hypothesis that the underlying distribution is lognormal with
known parameters /afii9839/p15and/afii9846/p17/p15. Similar to the procedures above, the likelihood
ratio test statistic is
X/p42/p582[l/p42/p44(/afii9839/p24,/afii9846/p24/p17)/p57l/p42/p44(/afii9839/p15,/afii9846/p17/p15)] (9 .4.3)
UnderH/p15,X/p42has an asymptotic chi-square distribution with 2 degrees of
freedom.
4.Testing the hypothesis that the underlying distribution is standard gamma
with known parameters /afii9838/p15and/afii9828/p15. The likelihood ratio test statistic is
X/p42/p582[l/p37(/afii9838/p19,/afii9828/p24)/p57l/p37(/afii9838/p15,/afii9828/p15)] (9 .4.4)
Under H/p15,X/p42is asymptotically chi-square distributed with 2 degrees of
freedom.234
5.Testing the hypothesis that the underlying distribution is generalized gamma
with known parameters /afii9825/p15,/afii9838/p15,and/afii9828/p15. The likelihood ratio test statistic is
X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p37/p37(/afii9825/p15,/afii9838/p15,/afii9828/p15)] (9 .4.5)
Under H/p15,X/p42is asymptotically chi-square distributed with 3 degrees of
freedom.
6.Testing the hypothesis that the underlying distribution is log-logisticwith
known parameters /afii9825/p15and/afii9828/p15. The likelihood ratio test statistic is
X/p42/p582[l/p42/p42(/afii9825/p24,/afii9828/p24)/p57l/p42/p42(/afii9825/p15,/afii9828/p15)] (9 .4.6)
UnderH/p15,X/p42has an asymptotic chi-square distribution with 2 degrees of
freedom.
Note that the respective Wald and score statistics can be constructed for
these tests by following (9.1.11 )and (9.1.12 ). These are left to the reader as
exercises. As noted in Section 9.3, the log-likelihood and estimated covariancematrix and parameters in (9.1.10 )—(9.1.12 )for each of the distributions dis-
cussed in Sections 7.2 to 7.6 can be obtained from SAS or BMDP. The otherterms in these test statistics can also be obtained by usingSAS or BMDP. Thefollowingexample illustrates the use of SAS and BMDP.
Example 9.6 To use the likelihood ratio statistic (9.4.1 )to test the null
hypothesis in Example 9.5, H/p15:/afii9838/p58/afii9838/p15/p580.06, we need to calculate the
log-likelihood l/p35(/afii9838/p19) andl/p35(/afii9838/p15).l/p35(/afii9838/p19) can be obtained by applyingeither the
SASorBMDPcodes inExample7.5.We nowshowhowto useSAS or BMDPto calculate l/p35(/afii9838/p15). Suppose that the survival data of the 10 mice in Example
9.5 are saved in the file ‘‘C: /p33EXAMPLE.DAT’’. If SAS is used, we specify that
the distribution is exponential by using
D/p58EXPONENTIAL in the ‘‘model’’
statement and lettingINTERCEPT /p582.813 [ /p58/p57log/afii9838/p15/p58
/p57log(0.06)]. If BMDP is used, we specify the distribution by lettingAC-
CEL /p58EXPONENTIAL and CONSTANT /p582.813. The followingSAS or
BMDP codes can be used to obtain the l/p35(/afii9838/p15)and the terms needed for the
Wald and score statistics in (9.1.11 )and (9.1.12 ).
SAS code:
data w1;
infile ‘c: /p33example.dat’ missover;
input t cens;
run;
proc lifereg;
model t*cens (0)/p58/maxit /p580 covb itprint d /p58exponential intercept /p582.813;
run; 235
BMDP code:
/input file /p58‘c:/p33example.dat’ .
variables /p582.
format /p58free.
/print level /p58brief.
cova. iterations.
/variable names /p58t, cens.
/form time /p58t.
status /p58cens.
response /p581.
/regress iteration /p580.
accel /p58exponential.
constant /p582.813.
/end
Similarly, to obtain the log-likelihood ratio statistic, the Wald and the score
statistics in (9.1.10 )—(9.1.12 )for testingnull hypotheses about the parameters
of other distributions, we can follow the same procedure but change the D /p58
and ACCEL /p58statements to reflect the distribution under the null hypothesis.
We also need to provide values for the input variables INTERCEPT andSCALE, for Weibull, lognormal, and log-logistic distributions, if SAS is used.For the extended generalized gamma distribution, we need to provide a valuefor SHAPE1. BMDP does not have a procedure for the gamma distribution.For the Weibull, lognormal, and log-logistic distributions, we need to providevalues for CONSTANT and SCALE. All of these input variables are based onthe distribution and their relationship to the parameters under the nullhypothesis (see notes at the end of each of the SAS or BMDP codes in Section
7.2 to 7.6 ).
9.5 HOLLANDER AND PROSCHAN’S TEST FOR
APPROPRIATENESS OF A GIVEN DISTRIBUTION WITH KNOWNPARAMETERS
Another test for the appropriateness of a parametric distribution with known
parameters was proposed by Hollander and Proschan (1979 ). Let
0/p58t/p7/p15/p8/p58t/p7/p16/p8/p58t/p7/p17/p8/p58/p37/p58t/p7/p76/p8be a set of distinct ordered survival times and
some of the t/p7/p71/p8’s may be censored. If censored observations are tied with
uncensoredobservations,treat the censoredobservationsof tieas beingg reaterthan the uncensored of the tie. Let S(t) be the underlyingsurvivorship function
andS/p15(t) the survivorship function of the specific distribution. The null
hypothesis is
H/p15:S(t)/p58S/p15(t)236
Usingthe Kaplan —Meier product-limit method, S(t) is estimated as
S/p19(t)/p58/p7/p73/p92/p16/p147
/p72/p14/p16/p1n/p57j
n/p57j/p591/p2/p66/p7/p72/p8t/p7/p73/p92/p16/p8/p58t/p45t/p7/p73/p8,k/p581,...,n
0 t/p57t/p7/p76/p8(9.5.1)
where /afii9829/p7/p72/p8/p581i ft/p7/p72/p8is uncensored and /afii9829/p7/p72/p8/p580i ft/p7/p72/p8is censored. Hollander and
Proschan’s test statistic for the null hypothesis that the data are from adistribution with survivorship function S(t)i s
C/p58 /p26
/p0/p12/p12 /p21/p14/p2/p4/p14/p19/p15/p18/p4/p3/p15/p1/p19/p4/p18/p22/p0/p20/p9/p15/p14/p19S/p15(t/p7/p71/p8)f/p19(t/p7/p71/p8)( 9 .5.2)
where f/p19(t/p7/p71/p8) is the jump of the Kaplan —Meier estimates at consecutive
uncensored observation and at the largest observation, uncensored or not,
f/p19(t/p7/p71/p8)/p581
n/p71/p92/p16/p147
/p72/p14/p16/p1n/p57j/p591
n/p57j/p2/p16/p92/p66/p7/p72/p8. (9.5.3 )
Under the null hypothesis,
C*/p58/p40n(C/p570.5)
/afii9846/p24(9.5.4)
follows approximatelythe standardnormal distribution, where /afii9846/p24is an estimate
of the standard deviation of Cand
/afii9846/p24/p17/p581
16/p76/p26
/p71/p14/p16n
n/p57i/p591[S/p19/p15(t/p7/p71/p92/p16/p8)/p57S/p19/p15(t/p7/p71/p8)] (9 .5.5)
To test H/p15:S/p58S/p15versusH/p16:S/p57S/p15, we reject H/p15ifC*/p58/p57Z/p63; to test H/p15versusH/p16:S/p58S/p15, we reject H/p15ifC*/p57Z/p63; and to test H/p15versusH/p16:S/p34S/p15,
we reject H/p15ifC*/p58/p57Z/p63/p30/p17orC*/p57Z/p63/p30/p17, where Z/p63is the upper /afii9825percentile
point of the standard normal distribution.
The procedure for the calculation of C*can be summarized as follows.
1. Compute the Kaplan —Meier estimate S/p19(t) for each uncensored observa-
tion.
2. Compute the jump of the Kaplan —Meier distribution at each t/p7/p71/p8uncen-
sored, that is, f/p19(t/p7/p71/p8), which is the difference of F/p19(t)/p581/p57S/p19(t) at two
consecutive uncensored observations.
3. Compute S/p15(t/p7/p71/p8) for each observation. ’ 237
4. Multiply S/p15(t/p7/p71/p8)b yf/p19(t/p7/p71/p8) and sum over all uncensored t/p7/p71/p8’s to obtain C.
5. Compute /afii9846/p24/p17accordingto (9.5.5 )and consequently, C*accordingto
(9.5.4 ).
Example 9.7 Consider the survival times in weeks of 10 mice in Example
9.5: 8, 5, 10 /p59, 1, 3, 18, 22, 15, 25 /p59, and 19. We wish to test that the survival
time followsan exponentialdistributionwith /afii9838/p580.06.The null and alternative
hypotheses are
H/p15:S(t)/p58S/p15(t)
H/p16:S(t)/p34S/p15(t)
whereS/p15(t)/p58exp(/p570.06t).
Followingtheprocedureoutlinedabove,wefirstarrangetheobservationsin
ascendingorderand computetheKaplan —Meierestimatesasshownincolumn
(d)of Table 9.5. The jumps are given in column (e). For example, the first jump
is between S/p19(0) and S/p19(1) or 1 /p570.9/p580.1. Column (f)gives the survival
function under the null hypothesis, for example, S/p15(3)/p58exp(/p570.06/p593)/p58
0.835. Following (9.5.2 ), column (g)gives the value of C/p580.4808. The last
three columns are for calculation of the estimated variance of C. Thus,
/afii9846/p24/p17/p581
16(1.3166 )/p580.0823
and
C*/p58/p4010(0.4808 /p570.5)
/p400.0823/p58/p570.2116
For/afii9825/p580.05,Z/p63/p30/p17/p581.96,C*does not fall in the rejection region. From Table
B-1 we obtain that the pvalue correspondingto C*/p58/p570.2116 is approxi-
mately 0.84. Therefore, we conclude that there is insufficient evidence to saythat the data are not from an exponential distribution with /afii9838/p580.06. Figure
9.1, which plots the Kaplan —Meier estimates and the hypothesized theoretical
distribution S/p15(t)/p58exp(/p570.06t), demonstrates a close agreement between the
two. The result is consistent with that obtained in Example 9.5, where thelikelihood ratio test is used.
Bibliographical Remarks
Readers with a background in mathematical statistics and an interest in
mathematical details about asymptotic likelihood theory, likelihood ratio,
Wald’s, and score statistics are referred to Cox (1961, 1962 a), Atkinson (1970 ),
Hagar and Bain (1970 ), Cox and Hinkley (1974 ), and Kalbfleisch and Prentice238
Table 9.5 Calculation of Test Statistic C*for Data in Example 9.7
(a)( b)( c)( d)( e)( f)( g)( h)( i)( j)
t/p7/p71/p8in/p57i
n/p57i/p591S/p19(t)f/p19(t/p7/p71/p8)S/p15(t/p7/p71/p8) (e)/p59(f)S/p19/p15(t/p7/p71/p8)S/p19/p15(t/p7/p71/p92/p16/p8)/p57S/p19/p15(t/p7/p71/p8)n
n/p57i/p591/p59(i)
1 1 0.900 0.900 0.100 0.941 0.0941 0.7841 0.2159 /p63 0.2159
3 2 0.889 0.800 0.100 0.835 0.0835 0.4861 0.2980 0.3311
5 3 0.875 0.700 0.100 0.741 0.0741 0.3015 0.1846 0.2308
8 4 0.857 0.600 0.100 0.619 0.0619 0.1468 0.1547 0.2210
10/p59 5 — — 0 0.549 0 0.0908 0.0560 0.0933
15 6 0.800 0.480 0.120 0.407 0.0560 0.0274 0.0634 0.126818 7 0.750 0.360 0.120 0.340 0.0408 0.0134 0.0140 0.035019 8 0.667 0.240 0.120 0.320 0.0384 0.0105 0.0029 0.0097
22 9 0.500 0.120 0.120 0.267 0.0320 0.0051 0.0054 0.0270
25/p5910 — — 0 0.223 0 0.0025 0.0026 0.0260——— ———0.4808 1.3166
/p63S/p19/p15(t/p7/p15/p8)/p58S/p19/p15(0)/p581.
239
Figure 9.1 Kaplan—Meier estimator S/p19(t) and the hypothesized survival function
S/p15(t)/p58exp(/p570.06t).
(1980 ). There have been many papers about the AIC and BIC criteria in the
literature since the introduction of AIC by Akaike (1969 ). The asymptotic
properties of the two criteria and their relationships with other criteria werediscussed by Akaike (1974 ), Parzen (1974 ), Schwarz (1978 ), Hannan (1979 ),
Shibata (1980 ),Wang (1984,1989 ), Rissanen (1986 ), and Wei (1992 ). Interested
readers are referred to these papers for details.
When there are no censored observations, the chi-square goodness of fit test
introduced by Karl Pearson in 1900 can be used to test any distributionalassumption. In addition, tests for the exponential and lognormal (Shapiro and
Wilk, 1965a, b )are available.
EXERCISES
9.1Derive the likelihood ratio and Wald test statistics following (9.1.2 )and
(9.1.3 )for the followingnull hypothesis:
H/p15: The underlyingdistribution is exponential
versus the alternatives
(a)H/p16: The underlyingdistribution is g amma
(b)H/p16: The underlying distribution is generalized gamma240
9.2Derive the Wald test statisticsfollowing (9.1.3 )for the followingnull and
alternative hypotheses
H/p15: The underlyingdistribution is Weibull
H/p16: The underlying distribution is generalized gamma
9.3Derive the respective Wald test statistics by following (9.1.3 )for the null
hypothesis that the distribution is the standard gamma versus thealternative hypothesis that the distribution is the generalized gamma.
9.4Derive the Wald and score statistics for testingthe null hypotheses in
Section 9.4 by following (9.1.11 )and (9.1.12 ).
9.5Consider the survival time of 28 cancer patients in Exercise 8.5.
(a)Obtain the log-likelihoods for the exponential, Weibull, lognormal,
and generalized gamma distributions. Perform the likelihood ratiotest and select the best distribution amongthese four distributions.
(b)Use the BIC and AIC procedures to select the best distribution
amongthefourdistributionsinpart (a)plusthelog-logisticdistribu-
tion.
(c)Comparetheresults obtainedin parts (a)and(b)and those obtained
in Exercise 8.5.
9.6Consider the survival time of 31 patients with advanced melanoma in
Exercise 8.6.(a)Select the best distribution usingthe likelihood ratio, Wald, and
score statistics among the exponential, Weibull, gamma, lognormal,and generalized gamma distributions.
(b)Use the BIC and AIC procedures to do the same as in part (a)with
the addition of the log-logistic distribution.
(c)Comparetheresults obtainedin parts (a)and(b)and those obtained
in Exercise 8.6.
(d)Compare the MLE of the parameters with those estimates obtained
by usingthe g raphical methods in Exercise 8.6.
9.7Do the same as in Exercise 9.5 for the data in Exercise Table 3.1.
9.8Consider the followingsurvival time in weeks of 10 mice with injection
of tumor cells: 5, 16, 18 /p59, 20, 22 /p59,2 4/p59, 25, 30 /p59, 35, 40 /p59. Do the data
follow the exponential distribution with /afii9838/p580.02?
(a)Use the likelihood ratio test.
(b)Use Hollander and Proschan’s test statistic.
(c)Plot the Kaplan —Meier estimator of S(t) and the hypothesized
distribution. 241
9.9Consider the followingsurvival time in months of 25 patients with
cancer of the prostate. Test the hypothesis that the survival time ofprostate cancer patients follows the exponential distribution with/afii9838/p580.01: 2, 19, 19, 25, 30, 35, 40, 45, 45, 48, 60, 62, 69, 89, 90, 110, 145,
160, 9 /p59,1 0/p59,2 0/p59,4 0/p59,5 0/p59, 110 /p59, 130 /p59.
9.10The Gompertz distribution belongs to the Gompertz —Makeham dis-
tribution family and the Gompertz —Makeham distribution (Makeham,
1860 )has the followinghazard function:
h(t)/p58/afii9825/p59exp(/afii9838/p59/afii9828t)/afii9825/p460
Construct alikelihood ratiostatistic and Wald’s statisticto test theappro-
priateness of a Gompertz distribution.242
CHAPTER 10
Parametric Methods for Comparing
Two Survival Distributions
In Chapter 5 we discussed several nonparametric tests for comparing twosurvival distributions. If the distributions follow a known model, parametrictests are more powerful than nonparametric tests, but their computation ismoretedious.Inthischapterwe firstdiscussthelikelihoodratio testingeneralfor comparing two survival distributions in Section 10.1. Readers who are notfamiliarwith linear algebramay skip this section without loss of continuity. InSections 10.2 to 10.4 we present either the likelihood ratio test or other testsfor the comparison of two survival patterns that follow the exponential,Weibull, and gamma distributions.
10.1 LIKELIHOOD RATIO TEST FOR COMPARING TWO
SURVIVAL DISTRIBUTIONS
Letx/p16,...,x/p76/p129, and y/p16,...,y/p76/p130betheobservedexactorcensoredsurvivaltimes
ofn/p16and n/p17subjectsfromtwogroups.Assumethatthesurvivaltimesfromthe
two groups follow the same distribution with different parameters. We use thegeneral notation b/p58(b/p16,b/p17,...,b/p78)to denote the set of parameters of the
distribution, p/p461.Let l/p71(b/p71),i/p581,2,denotethelog-likelihoodfunctionforthe
observed survival times from each group, where b/p71/p58(b/p71/p16,...,b/p71/p73,
b/p71/p73/p62/p16,...,b/p71/p78)/p58(b/p71/p16,b/p71/p17), andb/p71/p16/p58(b/p71/p16,...,b/p71/p73)andb/p71/p17/p58(b/p71/p73/p62/p16,...,b/p71/p78)are
twosubsetsofthe pparameters, i/p581,2.Thenthejointlog-likelihoodfunction
for the two groups is l(b/p16,b/p17)/p58l/p16(b/p16)/p59l/p17(b/p17). Letb/p19/p71denote the MLE of b/p71,
andb/p19/p71/p17(b/p15)denote the MLE of b/p71/p17givenb/p71/p16/p58b/p15, whereb/p15is known.
For example, if the survival time of the two groups follows the Weibull
distribution with a scale parameter /afii9838and a shape parameter /afii9828. Then p/p582,
b/p16/p58(/afii9838/p16,/afii9828/p16), andb/p17/p58(/afii9838/p17,/afii9828/p17), where /afii9838/p16,/afii9828/p16and/afii9838/p17,/afii9828/p17are the respective
parameters in the two Weibull distributions. Let l/p16(/afii9838/p16,/afii9828/p16) and l/p17(/afii9838/p17,/afii9828/p17) denote
the log-likelihood functions of the observed survival times from two groups;
243
then the joint log-likelihood function for the two groups is
l(/afii9838/p16,/afii9828/p16,/afii9838/p17,/afii9828/p17)/p58l/p16(/afii9838/p16,/afii9828/p16)/p59l/p17(/afii9838/p17,/afii9828/p17)
In this case, b/p71/p16may be a singleton /afii9838/p71, and similarly, b/p71/p17may be /afii9828/p71.
The followingtests are widelyused in comparingtwosurvival distributions.
Case 1. All parameters are unknown. When b/p71,i/p581, 2, are unknown, we
test the hypothesis
H/p15:b/p16/p58b/p17/p58b (10.1.1 )
that is, that the two groups have the same survival distribution with equal but
unknown parameters b. The log-likelihood ratio test statistic
X/p42/p58/p572[l(b/p19,b/p19)/p57l(b/p19/p16,b/p19/p17)]
/p582[l/p16(b/p19/p16)/p59l/p17(b/p19/p17)/p57l(b/p19,b/p19)] (10.1.2 )
has an asymptotic chi-square distribution with pdegrees of freedom. For a
given significance level /afii9825,H/p15is rejected if
X/p42/p57/afii9851/p17/p78/p11/p63(10.1.3)
or equivalently, if
P(/afii9851/p17/p78/p57X/p42)/p58/afii9825 (10.1.4 )
where /afii9851/p17/p78denotes the chi-square random variable with pdegrees of freedom,
and/afii9851/p17/p78/p11/p63is its 100 (1/p57/afii9825)percentile points, P(/afii9851/p17/p78/p57/afii9851/p17/p78/p11/p63)/p58/afii9825.
In the case of comparing two Weibull distributions, it reduces to
H/p15:/afii9838/p16/p58/afii9838/p17/p58/afii9838and/afii9828/p16/p58/afii9828/p17/p58/afii9828
where /afii9838and/afii9828are unknown,
X/p42/p582[l/p16(/afii9838/p19/p16,/afii9828/p24/p16)/p59l/p17(/afii9838/p19/p17,/afii9828/p24/p17)/p57l(/afii9838/p19,/afii9828/p24,/afii9838/p19,/afii9828/p24)]
and H/p15is rejected if X/p42/p57/afii9851/p17/p17/p11/p63, or equivalently, if P(/afii9851/p17/p17/p57X/p42)/p58/afii9825.
Case 2. A subset of the parameters of the two survival distributions are
known and equal, say, b/p16/p16/p58b/p17/p16/p58b/p15, where the values of b/p15are known. The
null hypothesis is the equality of the remaining parameters, or
H/p15:b/p16/p17/p58b/p17/p17/p58b/p28/p17(10.1.5 )244
whereb/p28/p17is unknown. The log-likelihood ratio statistic is
X/p42/p582[l/p16(b/p15,b/p19/p16/p17(b/p15))/p59l/p17(b/p15,b/p19/p17/p17(b/p15))/p57l((b/p15,b/p19/p28/p17(b/p15)),(b/p15,b/p19/p28/p17(b/p15))]
(10.1.6 )
whereb/p19/p16/p17(b/p15)is the MLE of b/p16/p17givenb/p16/p16/p58b/p15, and so are the others. X/p42has
an asymptotic chi-square distribution with degrees of freedom equal to thenumber of parameters in b/p16/p17(orb/p17/p17).
In the case of comparing two Weibull distributions, we may assume that
/afii9838/p16/p58/afii9838/p17/p58/afii9838/p15(or/afii9828/p16/p58/afii9828/p17/p58/afii9828/p15), where the value of /afii9838/p15(or/afii9828/p15)is known,and test
the null hypothesis
H/p15:/afii9828/p16/p58/afii9828/p17/p58/afii9828(or/afii9838/p16/p58/afii9838/p17/p58/afii9838)
Then
X/p42/p582[l/p16(/afii9838/p15,/afii9828/p24/p16(/afii9838/p15))/p59l/p17(/afii9838/p15,/afii9828/p24/p17(/afii9838/p15))/p57l(/afii9838/p15,/afii9828/p24(/afii9838/p15),/afii9838/p15,/afii9828/p24(/afii9838/p15))]
(orX/p42/p582[l/p16(/afii9838/p19/p16(/afii9828/p15),/afii9828/p15)/p59l/p17(/afii9838/p19/p17(/afii9828/p15),/afii9828/p15)/p57l(/afii9838/p19(/afii9828/p15),/afii9828/p15,/afii9838/p19(/afii9828/p15),/afii9828/p15)]
and H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63, equivalently, if P(/afii9851/p17/p16/p57X/p42)/p58/afii9825.
Case 3. A subset of the parameters of the two survival distributions are
equal but unknown, say, if b/p16/p16/p58b/p17/p16/p58b/p28/p16and the values of b/p28/p16are unknown.
The null hypothesis is the equality of the remaining parameters, or
H/p15:b/p16/p17/p58b/p17/p17/p58b/p28/p17(10.1.7 )
whereb/p28/p17is unknown and needs to be estimated.In addition, b/p28/p16also needs to
be estimated. The log-likelihood ratio statistic,
X/p42/p582[l((b/p19/p28/p16,b/p19/p16/p17),(b/p19/p28/p16,b/p19/p17/p17))/p57l(b/p19,b/p19)] (10.1.8 )
has an asymptoticchi-squaredistributionwith degrees of freedomequal to the
number of parameters in b/p16/p17(orb/p17/p17).
For the case of comparing two Weibull distributions, the derivation of X/p42in(10.1.8 )is left to the reader as an exercise.
Case 4. A subset of the parameters of the two survival distributions are
known but not equal, say, if b/p16/p16/p58b/p16/p15,b/p17/p16/p58b/p17/p15, andb/p16/p15andb/p17/p15are known 245
butb/p16/p15/p34b/p17/p15. The null hypothesis is the equality of the remaining parameters,
or
H/p15:b/p16/p17/p58b/p17/p17/p58b/p28/p17(10.1.9 )
The log-likelihood ratio statistic
X/p42/p582[l/p16(b/p16/p15,b/p19/p16/p17(b/p16/p15))/p59l/p17(b/p17/p15,b/p19/p17/p17(b/p17/p15))
/p57l((b/p16/p15,b/p19/p28/p17(b/p16/p15,b/p17/p15)),(b/p17/p15,b/p19/p28/p17(b/p16/p15,b/p17/p15)))] (10.1.10 )
has an asymptoticchi-squaredistributionwith degrees of freedomequal to the
number of parameters in b/p16/p17orb/p17/p17.
For the case of comparing two Weibull distributions, the derivation of X/p42in(10.1.10 )is left to the reader as an exercise.
10.2 COMPARISON OF TWO EXPONENTIAL DISTRIBUTIONS
Suppose that two survival distributions follow the exponential model with
parameters /afii9838/p16and/afii9838/p17, respectively. Two tests can compare the distributions:
thelikelihoodratiotestand an F-testsuggestedby Cox (1953 ).Thesetwotests
can test the hypothesis that the two exponential distributions are equalwhether or not the samples include censored observations.
10.2.1 Likelihood Ratio Test
Suppose that there are n/p16and n/p17individuals in groups 1 and 2, respectively,
x/p16,...,x/p80/p129uncensored and x/p62/p80/p129/p62/p16,..., x/p62/p76/p129censored in group 1, and y/p16,...,y/p80/p130uncensored and y/p62/p80/p130/p62/p16,..., y/p62/p76/p130censored in group 2. Thus, in group 1, there are
r/p16uncensored and n/p16/p57r/p16censored observations. In group 2, there are r/p17uncensored and n/p17/p57r/p17censored observations. If it is known that the survival
times of the two groups follow the exponential distribution with densityfunction f/p71(t)/p58/afii9838/p71e/p92/p72
/p71/p82,i/p581,2,testingtheequalityoftwoexponentialdistribu-
tionsisequivalenttotestingthehypothesis H/p15:/afii9838/p16/p58/afii9838/p17.Thisisbecausethetwo
exponential distributions are characterized by the two parameters /afii9838/p16and/afii9838/p17.
Thus, the null hypothesis is H/p15:/afii9838/p16/p58/afii9838/p17/p58/afii9838and the alternative hypothesis is
H/p16:/afii9838/p16/p34/afii9838/p17. According to (10.1.2 ), the test statistic for the likelihood ratio test
is
X/p42/p58/p572logL(/afii9838/p19,/afii9838/p19)
L(/afii9838/p19/p16,/afii9838/p19/p17)(10.2.1)246
wherethedenominatoristhelikelihoodfunctionforthetwogroupscombined,
L(/afii9838/p19/p16,/afii9838/p19/p17)/p58/afii9838/p19/p80/p16/p16/afii9838/p19/p80/p17/p17exp/p3/p57/afii9838/p19/p16/p1/p80/p16/p26
/p71/p14/p16x/p71/p59/p76/p16/p26
/p71/p14/p80/p16/p62/p16x/p62/p71/p2
/p57/afii9838/p19/p17/p1/p80/p17/p26
/p71/p14/p16y/p71/p59/p76/p17/p26
/p71/p14/p80/p17/p62/p16y/p62/p71/p2/p4(10.2.2)
and/afii9838/p19/p16and/afii9838/p19/p17are the MLE of /afii9838/p16and/afii9838/p17, respectively, from groups 1 and 2.
From Section 7.2,
/afii9838/p19/p16/p58r/p16/p26/p80/p16/p71/p14/p16x/p71/p59/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71/afii9838/p19/p17/p58r/p17/p26/p80/p17/p71/p14/p16y/p71/p59/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71(10.2.3)
The numerator in (10.2.1 )is the likelihood function for the combined sample
under the null hypothesis, that is, /afii9838/p16/p58/afii9838/p17/p58/afii9838,
L(/afii9838/p19,/afii9838/p19)/p58/afii9838/p19/p80/p16/p62/p80/p17exp/p3/p57/afii9838/p19/p1/p80/p16/p26
/p71/p14/p16x/p71/p59/p76/p16/p26
/p71/p14/p80/p16/p62/p16x/p62/p71/p59/p80/p17/p26
/p71/p14/p16y/p71/p59/p76/p17/p26
/p71/p14/p80/p17/p62/p16/p2/p4
(10.2.4)
where /afii9838/p19is the MLE of /afii9838obtained from the combined sample,
/afii9838/p19/p58r/p16/p59r/p17/p26/p80/p16/p71/p14/p16x/p71/p59/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71/p59/p26/p80/p17/p71/p14/p16y/p71/p59/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71(10.2.5)
From Section 10.1, X/p42has an approximate chi-square distribution with 1
degree of freedom for samples of at least 25 ( n/p16/p59n/p17/p4625) under the null
hypothesis. For a given significance level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63,o r
equivalently, if P(/afii9851/p17/p16/p57X/p42)/p58/afii9825.
The test procedure can be summarized as follows:
1. Compute /afii9838/p19/p16and/afii9838/p19/p17following (10.2.3 ).
2. Compute L(/afii9838/p19/p16,/afii9838/p19/p17)in(10.2.2 )usingthegivendataand /afii9838/p19/p16and/afii9838/p19/p17obtained
in step 1.
3. Compute /afii9838/p19following (10.2.5 ).
4. Compute L(/afii9838/p19,/afii9838/p19)i n(10.2.4 ).
5. Compute X/p42in(10.2.1 ).I fX/p42/p57/afii9851/p17/p16/p11/p63(Table B-2 ), reject H/p15and conclude
that the two exponential survival distributions are not equal. Otherwise,the data do not provide enough evidence to reject the null hypothesis.
If there are no censored observations in the data, (10.2.1 )—(10.2.5 )are also
applicable simply be letting n/p16/p58r/p16,n/p17/p58r/p17and omitting the terms involving 247
x/p62/p71and y/p62/p71. The likelihood ratio test is primarily for two-sided tests and is
difficulttoapplyto a one-sidedtest.It isapproximateandshouldbe used withcaution when the sample size is small. The power of the test, similar to that ofother likelihood tests, is not high. That is, if the likelihood ratio test is usedregularly, one is more likely not to reject the null hypothesis when the twosurvival distributions are not equal.
Example 10.1 Consider the remission data of the two treatment groups
given in Example 5.1. The remission times in months are as follows:
CMF: 23, 16 /p59,1 8/p59,2 0/p59,2 4/p59
Control: 15, 18, 19, 19, 20
Assume that the two distributions are exponential with parameters /afii9838/p16and/afii9838/p17,
respectively. Using the likelihood ratio test, we test the following null hypoth-esis
H/p15:/afii9838/p16/p58/afii9838/p17/p58/afii9838(the two treatments are equally effective )
against
H/p16:/afii9838/p16/p34/afii9838/p17(the two treatments are not equally effective )
Following the above, we proceed as follows:
1. Compute /afii9838/p19/p16and/afii9838/p19/p17in(10.2.3 ): In this case, n/p16/p58n/p17/p585,r/p16/p581,r/p17/p585,
/p26/p80
/p16/p71/p14/p16x/p71/p5823,/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71/p5878,/p26/p80/p17/p71/p14/p16y/p71/p5891,/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71/p580,
/afii9838/p19/p16/p581
23/p5978/p581
101/p580.0099
/afii9838/p19/p17/p585
91/p580.0549
2. Compute L(/afii9838/p19/p16,/afii9838/p19/p17)in(10.2.2 ):
L(/afii9838/p19/p16,/afii9838/p19/p17)/p58(0.0099)(0 .0549) /p20exp[/p570.0099 (101)/p570.0549 (91)]
/p581.2290 (10)/p92/p16/p16
3. Compute /afii9838/p19in(10.2.5 ):
/afii9838/p19/p581/p595
23/p5978/p5991/p586
192/p580.0313248
4. Compute L(/afii9838/p19/p16,/afii9838/p19/p17)i n(10.2.4 ):
L(/afii9838/p19/p16,/afii9838/p19/p17)/p58(0.0313) /p21exp[/p570.0313 (192)]/p582.3085 (10)/p92/p16/p17
5. Compute X/p42in(10.2.1 ):
X/p42/p58/p572log2.3085 (10)/p92/p16/p17
1.2290 (10)/p92/p16/p16/p58/p572log (0.1878 )/p583.344
From Table B-2 we obtain /afii9851/p17/p16/p11/p15/p13/p15/p20/p583.84. Thus we cannot reject H/p15at the
0.05 level. Recall that in Chapter 5 the null hypothesis was rejected at the 0.05level by using the four nonparametric tests.
10.2.2 Cox’s F-Test for Exponential Distributions
If the times to failure can be assumed to follow the exponential distribution in
both treatment groups, an F-test suggested by Cox (1953 )can be used to test
for treatment differences whether or not censored observations are present.Suppose that we wish to test the hypothesis H/p15:/afii9838/p16/p58/afii9838/p17against either the
one-sided alternative H/p16:/afii9838/p16/p58/afii9838/p17(orH/p17:/afii9838/p16/p57/afii9838/p17)or the two-sided alternative
H/p18:/afii9838/p16/p34/afii9838/p17. An efficient test is to take t/p16/p16/t/p16/p17as having an F-distribution with
(2r/p16,2r/p17)degrees of freedom, where
t/p16/p16/p58/p26/p80
/p16/p71/p14/p16x/p71/p59/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71r/p16(10.2.6)
t/p16/p17/p58/p26/p80/p17/p71/p14/p16y/p71/p59/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71r/p17(10.2.7)
The test procedures are (1)forH/p16, reject H/p15ift/p16/p16/t/p16/p17/p57F/p17/p80/p16/p11/p17/p80/p17/p11/p63;(2)forH/p17,
reject H/p15ift/p16/p16/t/p16/p17/p58F/p17/p80/p16/p11/p17/p80/p17/p11/p16/p92/p63; and (3)forH/p18, reject H/p15ift/p16/p16/t/p16/p17/p57F/p17/p80/p16/p11/p17/p80/p17/p11/p63/p30/p17or
t/p16/p16/t/p16/p17/p58F/p17/p80/p16/p11/p17/p80/p17/p11/p16/p92/p63/p30/p17, where /afii9825is the significance level and F/p17/p80/p16/p11/p17/p80/p17/p11/p63is the upper
100/afii9825percentage point of the F-distribution with (2 r/p16,2r/p17) degrees of freedom.
Similarly, the hypothesis that /afii9838/p16//afii9838/p17/p58kcan be tested by referring kt/p16/p16/t/p16/p17to the
table of the F-distribution.
When there are no censored observations, that is, n/p16/p58r/p16,n/p17/p58r/p17, the
second terms of the numerators in (10.2.6 )and (10.2.7 )are zero. Then the test
statistic t/p16/p16/t/p16/p17has an F-distribution with (2 n/p16,2n/p17)degrees of freedom.
Confidence intervals for the ratio /afii9838/p16//afii9838/p17can be obtained from the fact that
/afii9838/p16t/p16/p16//afii9838/p17t/p16/p17has the F-distribution with (2 n/p16,2n/p17) degrees of freedom. It follows
thata100 (1/p57/afii9825)%confidenceintervalfortheratiooftwohazardrates /afii9838/p16//afii9838/p17is
t/p16/p17t/p16/p16F/p17/p80/p16/p11/p17/p80/p17/p11/p16/p92/p63/p30/p17/p58/afii9838/p16/afii9838/p17/p58t/p16/p17t/p16/p16F/p17/p80/p16/p11/p17/p80/p17/p11/p63/p30/p17(10.2.8) 249
Example 10.2 Thirty-six patients with glioblastoma multiforme were
divided into two groups; the experimental group contained 21 patients whohad surgery and chemotherapy, and the control group contained 15 patientswhohadsurgeryonly.Thesurvivaltimesinweeksareavailableaboutoneyearafter the start of the study (Burdette and Gerhan, 1970 ):
Experimental: 1, 2, 2, 2, 6, 8, 8, 9, 13, 16, 17, 29, 34, 2 /p59,9/p59,1 3/p59,2 2/p59,
25/p59,3 6/p59,4 3/p59,4 5/p59
Control: 0, 2, 5, 7, 12, 42, 46, 54, 7 /p59,1 1/p59,1 9/p59,2 2/p59,3 0/p59,3 5/p59,
39/p59
The hypotheses are
H/p15:/afii9838/p16/p58/afii9838/p17(no difference in survival between experimental
and control groups )
H/p16:/afii9838/p16/p58/afii9838/p17(difference in survival favoring experimental group )
In this case, n/p16/p5821, n/p17/p5815, r/p16/p5813, r/p17/p588,/afii9814x/p71/p58147,/afii9814x/p62/p71/p58195,
/afii9814y/p71/p58168, and /afii9814y/p62/p71/p58163. Hence
t/p16/p16/p58147/p59195
13
/p5826.308 t/p16/p17/p58168/p59163
8/p5841.375
and t/p16/p16/t/p16/p17/p580.636 with (26,16 )degrees of freedom. For /afii9825/p580.05, F/p17/p21/p11/p16/p21/p11/p15/p13/p15/p20is
approximately 2.23; hence the hypothesis H/p15is not rejected and the data do
not provide enough evidence that the survival time is longer in the experimen-tal group. A 95%confidence interval for the ratio /afii9838/p16//afii9838/p17is
41.375
26.308(0.419 )/p58/afii9838/p16/afii9838/p17/p5841.375
26.308(2.625 )
or(0.659, 4.128 ). The estimate of /afii9838/p19/p16//afii9838/p19/p17according to (10.2.3 )is 1.58. Hence the
data show that the death rate per week of the experimental group is close tothat of the control group.
In Example 10.2, the guarantee time in both groups is zero. In the case
where a group has a nonzero guarantee time, it can be subtracted from everyobservation in the group and the test then applied. Monte Carlo studies(Gehan and Thomas, 1969; Lee et al., 1975 )show that when samples are from
exponential distributions, with or without censoring, the F-test is the most
powerful test among the parametric or nonparametric tests discussed in thischapter and Chapter 5.250
10.3 COMPARISON OF TWO WEIBULL DISTRIBUTIONS
It is well known that if the survival time Thas a Weibull distribution with
shape parameter /afii9828, then T/p65has an exponential distribution. Thus, if /afii9828/p16and/afii9828/p17for the two groups are known, the most powerful Cox’s F-test described in
Section 10.2 can be applied to the transformed observations. However, inpractice, /afii9828/p16and/afii9828/p17are probably unknown, and so are the scale parameters, /afii9838/p16and/afii9838/p17. In this case, the likelihood ratio tests described in Section 10.1 can be
applied to test whether the observed survival times from the two groups havethe same Weibull distribution. To test the equality of two Weibull distribu-tions,itsufficestotest /afii9828/p16/p58/afii9828/p17and/afii9838/p16/p58/afii9838/p17.Ifthehypothesis /afii9828/p16/p58/afii9828/p17isrejected,
we need not test the hypothesis /afii9838/p16/p58/afii9838/p17. If the hypothesis /afii9828/p16/p58/afii9828/p17is not
rejected, we do need to test /afii9838/p16/p58/afii9838/p17. In the following, we introduce an
additional two-sample test proposed by Thoman and Bain (1969 )for uncen-
sored samples.
Assume that independent random samples of equal size ( n/p16/p58n/p17/p58n) are
obtained from Weibull distributions f/p16(t) and f/p17(t), where
f/p71(t)/p58/afii9838/p71/afii9828/p71(/afii9838/p71t)/p65
/p71/p92/p16exp[/p57(/afii9838/p71t)/p65/p71] i/p581, 2 (10 .3.1)
To test /afii9828/p16/p58/afii9828/p17, we use the property of the maximum likelihood estimator /afii9828/p24
(Thoman et al., 1969; Thoman and Bain, 1969 ). To test the null hypothesis
H/p15:/afii9828/p16/p58/afii9828/p17against H/p16:/afii9828/p16/p57/afii9828/p17, we use the fact that (/afii9828/p24/p16//afii9828/p16)/(/afii9828/p24/p17//afii9828/p17)/p58/afii9828/p24/p16//afii9828/p24/p17under H/p15. Thepercentagepoints of /afii9828/p24/p16//afii9828/p24/p17are givenin TableB-12.We compute
theMLEof /afii9828/p16and/afii9828/p17,thatis, /afii9828/p24/p16and/afii9828/p24/p17,andcompare /afii9828/p24/p16//afii9828/p17withthepercentage
points for a given /afii9825in Table B-12. Reject H/p15at/afii9825level if /afii9828/p24/p16//afii9828/p24/p17/p57l/p63. For
example,if n/p16/p58n/p17/p58n/p5810,acomputed /afii9828/p24/p16//afii9828/p24/p17/p571.897wouldleadtorejection
ofH/p15at a significance level of 0.05. For /afii9825/p460.50, percentage points l/p63can be
calculated by using the relationship l/p63/p581/l/p16/p92/p63.
The procedure described above can be generalized to test H/p15:/afii9828/p16/p58k/afii9828/p17against H/p16:/afii9828/p16/p58k/p30/afii9828/p17. For the case when k/p58k/p30, the rejection region becomes
/afii9828/p24/p16//afii9828/p24/p17/p57kl/p63, where /afii9825is the significance level.
If the hypothesis H/p15:/afii9828/p16/p58/afii9828/p17is rejected, the two Weibull distributions are
not the same. However, if the hypothesis is not rejected, we need to test theequality of the two scale parameters /afii9838/p16and/afii9838/p17. A test of H/p15:/afii9838/p16/p58/afii9838/p17against
H/p16:/afii9838/p16/p58/afii9838/p17suggested by Thoman and Bain (1969 )rejects H/p15if
G/p58/p16/p17(/afii9828/p24/p16/p59/afii9828/p24/p17)(log/afii9838/p19/p17/p57log/afii9838/p19/p16)/p57z/p63(10.3.2)
where z/p63issuchthat P(G/p58z/p63/p34H/p15)/p581/p57/afii9825and/afii9828/p24/p16,/afii9828/p24/p17,/afii9838/p19/p16,and/afii9838/p19/p17aretheMLEs
of/afii9828/p16,/afii9828/p17,/afii9838/p16, and /afii9838/p17respectively. The percentage points z/p63are given in Table
B-13.Forexample,ifthe commonsamplesizeis 10,thehypothesis H/p15:/afii9838/p16/p58/afii9838/p17is rejected if G/p460.918 at significance level 0.05. A test of H/p15:/afii9838/p16/p58/afii9838/p17against
H/p16:/afii9838/p16/p57/afii9838/p17can be constructed in a similar fashion. The critical points z/p63can
be obtained from Table B-13 by using the fact that z/p63/p58/p57 z/p16/p92/p63. 251
Table 10.1 Survival Times of Patients in Two Treatment Groups
Treatment 1 Treatment 2
5, 10, 17, 32, 32, 33, 34, 36, 43, 44, 44, 48,
48, 61, 64, 65, 65, 66, 67, 68, 82, 85,90, 92, 92, 102, 103, 106, 107, 114, 114,116,117,124,139,142,143,151,158,19520.9, 32.2, 33.2, 39.4, 40.0, 46.8, 57.3,
58.0,59.7,61.1,61.4,54.3,66.0,66.3,67.4,68.5, 69.9, 72.4, 73.0, 73.2, 88.7, 89.3,91.6, 93.1 94.2, 97.7, 101.6, 101.9, 107.6,
108.0, 109.7, 110.8, 114.1, 117.5, 119.2,
120.3, 133.0, 133.8, 163.3, 165.1
Example 10.3 illustrates the test procedures. The data are adapted and
modified from Harter and Moore (1965 ). Forty observations are generated
from a Weibull distribution with /afii9838/p16/p580.01 and /afii9828/p16/p582 and another 40 from a
Weibull distribution with /afii9838/p17/p580.01 and /afii9828/p17/p583. The resulting data are shown
in Table 10.1. For illustrative purposes, we consider the two samples as twotreatment groups.
Example 10.3 Consider the survival times of the patients in the two
treatmentgroupsinTable10.1.Thenullhypothesisisthatthetwopopulationshave the same shape parameter; that is, H/p15:/afii9828/p16/p58/afii9828/p17against H/p16:/afii9828/p16/p58/afii9828/p17.
The MLEs are /afii9828/p24/p16/p581.945, /afii9828/p24/p17/p582.715, and hence /afii9828/p24/p17//afii9828/p24/p16/p581.396, which is
significant at the 0.05 level ( l/p15/p13/p15/p20/p581.342 for n/p5840) but not significant at the
0.02 level ( l/p15/p13/p15/p17/p581.453 for n/p5840). If we choose /afii9825/p580.05 and reject H/p15, the
decision is correct. An error of not rejecting H/p15would be committedif an /afii9825of
0.02 or 0.01 is chosen. This is because the two shape parameters are very close(/afii9828/p16/p582,/afii9828/p17/p583).
To illustrate the procedure of testing the equality of the scale parameters,
letusassumethatthehypothesis H/p15:/afii9828/p16/p58/afii9828/p17isnotrejected.Totest H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p57/afii9838/p17,weneedtheMLEsof /afii9838/p16and/afii9838/p17. HarterandMoore obtain
/afii9838/p19/p16/p580.010776 and /afii9838/p19/p17/p580.010471. From (10.3.2 )we obtain
G/p58/p16/p17
(1.945 /p592.715 )(4.559 /p574.530 )/p580.068
From Table B-13, the critical region for n/p5840 is G/p570.404. Hence we do not
reject H/p15. This decision is correct since /afii9838/p16/p58/afii9838/p17/p580.01. Note that the MLEs of
/afii9828/p16,/afii9828/p17,/afii9838/p16, and /afii9838/p17are very close to their real values.
10.4 COMPARISON OF TWO GAMMA DISTRIBUTIONS
Suppose that x/p16,...,x/p76, and y/p16,...,y/p76are the survival times of patients
receivingtwo different treatmentsand that they followthe gamma distribution
with the density function given in (6.4.1 ). Let /afii9838/p16and/afii9828/p16be the parameters of252
Table 10.2 Survival Times of 40 Patients Receiving
Two Different Treatments
Treatment 1 (x) Treatment 2( y)
17, 28, 49, 98, 119 26, 34, 47, 59, 101,
133, 145, 146, 158, 160, 112, 114, 136, 154, 154,174, 211, 220, 231, 252, 161, 186, 197, 226, 226,256, 267, 322, 323, 327 243, 253, 269, 308, 465thexpopulation and /afii9838/p17and/afii9828/p17be those of the ypopulation. The likelihood
ratio tests introduced in Section 10.1 can be used to test whether the survivaltimes observed from the xpopulation and the ypopulation have different
gamma distributions. The estimation of the parameters is quite complicatedbut can be obtained using commercially available computer programs. In thefollowing we introduce an F-test for testing the null hypothesis H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p34/afii9838/p17, under the assumptions that the x/p71’s and y/p71’s are exact
(uncensored )survival times, and that /afii9828/p16and/afii9828/p17are known (usually assumed
equal ).
Letx/p21and y/p21be the sample mean survival times of the two groups. The test
is based on the fact that x/p21/y/p21has the F-distribution with 2 n/afii9828/p16and 2 n/afii9828/p17degrees
of freedom (Rao, 1952 ). Thus the test procedure is to reject H/p15at the /afii9825level if
x/p21/y/p21exceeds F/p17/p76/p65
/p16/p11/p17/p76/p65/p17/p11/p63/p30/p17, the 100 (/afii9825/2)percentage point of the F-distribution
with (2n/afii9828/p16,2n/afii9828/p17)degrees of freedom. Since the F-table gives percentage points
for integer degrees of freedom only, interpolations (linear or bilinear )are
necessary when either 2 n/afii9828/p16or 2n/afii9828/p17is not an integer.
The following example illustrates the test procedure. The data are adapted
andmodifiedfromHarterandMoore (1965 ).Theysimulated40survivaltimes
from the gamma distribution with parameters /afii9828/p16/p58/afii9828/p17/p58/afii9828/p582,/afii9838/p580.01. The
40 individuals are divided randomly into two groups for illustrative purposes.
Example 10.4 Consider the survival time of the two treatment groups in
Table 10.2. The two populations follow the gamma distributions with acommon shape parameter /afii9828/p582. To test the hypothesis H/p15:/afii9838/p16/p58/afii9838/p17against
H/p16:/afii9838/p16/p34/afii9838/p17, we compute x/p21/p58181.80,y/p21/p58173.55, and x/p21/y/p21/p581.048. Under the
nullhypothesis, x/p21/y/p21hasthe F-distributionwith (80,80 )degreesoffreedom.Use
/afii9825/p580.05, F/p23/p15/p11/p23/p15/p11/p15/p13/p15/p17/p20/p601.45. Hence, we do not reject H/p15at the 0.05 level of
significance. The result is what we would expect since the two samples aresimulated from the same overall sample of 40 with /afii9838/p580.01.
To test the equality of two lognormal distributions, we use the fact that the
logarithmic transformation of the observed survival times follows the normaldistributions, and thus we can use the standard tests based on the normaldistribution. In general, for other distributions, such as log-logistic and thegeneralized gamma, the log-likelihood ratio statistics defined in Section 10.1 253
can be applied to test whether the survival times observed from two groups
have the same distribution. Readers can follow Example 10.2.1 in Section 10.2and use the respective likelihood functions derived in Chapter 7 to constructthe needed tests.
Bibliographical Remarks
In addition to the papers cited in this chapter, readers are referred to Mann et
al.(1974 ), Gross and Clark (1975 ), Lawless (1982 ), and Nelson (1982 ).
EXERCISES
10.1Derive the likelihood ratio tests in (10.1.8 )and (10.1.10 )for testing the
equality of two Weibull distributions.
10.2Derive the likelihood ratio test in (10.1.2 )for testing the equality of two
log-logistic distributions with unknown parameters.
10.3Consider the remission data of the leukemia patients in Example 3.3.
Assume that the remission times of the two treatment groups follow theexponentialdistribution.Testthehypothesisthatthetwotreatmentsareequally effective using:
(a)The likelihood ratio test
(b)Cox’s F-test
Obtain a 95%confidence interval for the ratio of the two hazard rates.
10.4For the same data in Exercise 10.3, test the hypothesis that /afii9838/p17/p585/afii9838/p16.
10.5Suppose that the survival time of two groups of lung cancer patients
follows the Weibull distribution. A sample of 30 patients (15 from each
group )was studied. Maximun likelihood estimates obtained from the
two groups are, respectively, /afii9828/p24/p16/p583,/afii9838/p19/p16/p581.2 and /afii9828/p24/p17/p582,/afii9838/p19/p17/p580.5. Test
the hypothesis that the two groups are from the same Weibull distribu-tion.
10.6Divide the lifetimes of 100 strips (delete the last one )of aluminum
coupon in Table 6.4 randomly into two equal groups. This can be donebyassigningtheobservationsalternatelytothetwogroups.Assumethatthe two groups follow a gamma distribution with shape parameter/afii9828/p5812. Test the hypothesis that the two scale parameters are equal.
10.7Twelve brain tumor patients are randomized to receive radiation ther-
apy or radiation therapy plus chemotherapy (BCNU )in a one-year
clinical trial. The following survival times in weeks are recorded:254
1. Radiation /p59BCNU: 24, 30, 42, 15 /p59,4 0/p59,4 2/p59
2. Radiation: 10, 26, 28, 30, 41, 12 /p59
Assumingthatthesurvivaltimefollowsthe exponentialdistribution,use
Cox’s F-test for exponential distributions to test the null hypothesis
H/p15:/afii9838/p16/p58/afii9838/p17versus the alternative H/p16:/afii9838/p16/p58/afii9838/p17.
10.8Use one of the nonparametric tests discussed in Chapter 5 to test the
equalityof survivaldistributionsofthe experimentaland controlgroupsin Example 10.2. Compare your result with that obtained in Example10.2. 255
CHAPTER11
ParametricMethodsforRegression
ModelFittingandIdentificationofPrognosticFactors
Prognosis,thepredictionofthefutureofanindividualpatientwithrespecttoduration,course,andoutcomeofadiseaseplaysanimportantroleinmedicalpractice.Beforeaphysiciancanmakeaprognosisanddecideonthetreatment,amedicalhistoryaswellaspathologic,clinical,andlaboratorydataareoftenneeded.Therefore,manymedicalchartscontainalargenumberofpatient (or
individual )characteristics (also calledconcomitantvariables ,independentvari-
ables,covariates,prognosticfactors ,o rriskfactors ), and it is often difficult to
sort outwhich onesare mostcloselyrelatedto prognosis.The physiciancanusuallydecidewhichcharacteristicsare irrelevant,buta statisticalanalysisisusuallyneededtoprepareacompactsummaryofthedatathatcanrevealtheirrelationship. One way to achieve this purpose is to search for a theoreticalmodel (or distribution ), that fits the observed data and identify the most
important factors. These models, usually regression models, extend themethods discussed in previous chapters to include covariates. In this chapterwe focus on parametric regression models (i.e., we assume that the survival
time follows a theoretical distribution ). If an appropriate model can be
assumed, the probability of surviving a given time when covariates areincorporatedcanbeestimated.
InSection11.1wediscussbrieflypossibletypesofresponseandprognostic
variables and things that can be done in a preliminary screening before aformal regression analysis. This section applies to methods discussed in thenext four chapters. In Section 11.2 we introduce the general structure of acommonly used parametric regression model, the accelerated failure time(AFT )model.Sections11.3to11.7coverseveralspecialcasesofAFTmodels.
Fittingthesemodelsofteninvolvescomplicatedandtediouscomputationsandrequirescomputersoftware.Fortunately,mostoftheproceduresareavailableinsoftwarepackagessuchasSASandBMDP.TheSASandBMDPcodethat
256
canbeusedtofitthemodelsaregivenattheendoftheexamples.Readersmay
findthesecodeshelpful.Section11.8introducestwoothermodels.InSection11.9wediscussthemodelselectionmethodsandgoodnessoffittests.
11.1 PRELIMINARY EXAMINATION OF DATA
Informationconcerningpossibleprognosticfactorscanbeobtainedeitherfrom
clinical studiesdesignedmainly to identifythem, sometimescalled prognostic
studies,orfromongoingclinicaltrialsthatcomparetreatmentsasasubsidiary
aspect. The dependent variable (also called the response variable ), or the
outcome of prediction, may be dichotomous, polychotomous, or continuous.Examples of dichotomous dependent variables are response or nonresponse,life or death, and presence or absence of a given disease. Polychotomousdependentvariablesincludedifferentgradesofsymptoms (e.g.,noevidenceof
disease,minorsymptom,majorsymptom )and scoresof psychiatricreactions
(e.g.,feelingwell,tolerable,depressed,orverydepressed ).Continuousdepend-
ent variables may be length of survival from start of treatment or length ofremission,bothmeasuredonanumericalscalebyacontinuousrangeofvalues.Of these dependent variables, response to a given treatment (yes or no ),
developmentof a specific disease (yes or no ), lengthof remission,and length
ofsurvivalare particularlycommonin practice.In this chapterwe focusourattention on continuous dependent variables such as survival time and re-missionduration.Dichotomousandmultiple-responsedependentvariablesarediscussedinChapter14.
Aprognostic variable (or independent variable )or factor may be either
numerical or nonnumerical. Numerical prognostic variables may be discrete,suchasthenumberofpreviousstrokesornumberoflymphnodemetastases,or continuous, such as age or blood pressure. Continuous variables can bemadediscretebygroupingpatientsintosubcategories (e.g.,fouragesubgroups:
/p5820, 20—39, 40—59, and /p4660). Nonnumerical prognostic variables may be
unordered (e.g.,raceordiagnosis )orordered (e.g.,severityofdiseasemaybe
primary,local,ormetastatic ).Theycanalsobedichotomous (e.g.,alivereither
is or is not enlarged ).Usually, the collectionof prognosticvariables includes
someofeachtype.
Before a statistical calculation is done, the data have to be examined
carefully. If some of the variables are significantly correlated, one of thecorrelated variables is likely to be a predictor as good as all of them.Correlation coefficients between variables can be computed to detect signifi-cantly correlated variables. In deleting any highly correlated variables, infor-mationfromotherstudieshastobeincorporated.Ifotherstudiesshowthatagivenvariablehasprognosticvalue,itshouldberetained.
Inthenexteight sectionswediscussmultivariateorregressiontechniques,
which are useful in identifying prognostic factors. The regression techniquesinvolveafunctionoftheindependentvariablesorpossibleprognosticfactors. 257
Thevariablesmust bequantitative,with particularnumericalvaluesforeach
patient. This raises no problem when the prognostic variables are naturallyquantitative (e.g.,age )andcanbeusedintheequationdirectly.However,ifa
particular prognostic variable is qualitative (e.g., a histological classification
into one of three cell types A, B, or C ), something needs to be done. This
situation can be covered by the use of two dummy variables, e.g., x/p16, taking
thevalue1forcelltypeAand0otherwise,and x/p17,takingthevalue1forcell
typeBand0otherwise.Clearly,ifthereareonlytwocategories (e.g.,sex ),only
onedummyvariableisneeded: x/p16is1foramale,0forafemale.Also,abetter
descriptionofthedatamightbeobtainedby usingtransformedvaluesof theprognosticvariables (e.g.,squaresorlogarithms )orbyincludingproductssuch
asx/p16x/p17(representing an interaction between x/p16andx/p17). Transforming the
dependent variable (e.g., taking the logarithm of a response time )can also
improvethefit.
Inpractice,thereareusuallyalargernumberofpossibleprognosticfactors
associatedwiththeoutcomes.Onewaytoreducethenumberoffactorsbeforeamultivariateanalysisisattemptedistoexaminetherelationshipbetweeneachindividual factor and the dependent variable (e.g., survival time ). From the
univariate analysis, factors that have little or no effect on the dependentvariablecanbeexcludedfromthemultivariateanalysis.However,itwouldbedesirabletoincludefactorsthathavebeenreportedtohaveprognosticvaluesbyotherinvestigatorsandfactorsthatareconsideredimportantfrombiomedi-calviewpoints.Itisoftenusefultoconsidermodelselectionmethodstochoosethosesignificantfactorsamongallpossiblefactorsanddetermineanadequatemodel with as few variables as possible. Very often, a variable of significantprognosticvalueinonestudyisunimportantinanother.Therefore,confirma-tioninalaterstudyisveryimportantinidentifyingprognosticfactors.
Another frequent problem in regression analysis is missing data. Three
distinctionsaboutmissingdatacanbemade: (1)dependentversusindependent
variables, (2)manyversusfewmissingdata,and (3)randomversusnonrandom
loss of data. If the value of the dependent variable (e.g., survival time )is
unknown,thereislittletodobutdropthatindividualfromanalysisandreducethe sample size. The problem of missing data is of different magnitudedependingonhowlargeaproportionofdata,eitherforthedependentvariableor for the independent variables, is missing. This problem is obviously lesscritical if 1%of data for one independent variable is missing than if 40%ofdataforseveralindependentvariablesismissing.Whenasubstantialpropor-tion of subjects has missing data for a variable, we may simply opt to dropthemandperformtheanalysisontheremainderofthesample.Itisdifficulttospecify‘‘howlarge’’and‘‘howsmall,’’butdropping10or15casesoutofseveralhundredwould raise no serious practicalobjection.However,if missing dataoccurinalargeproportionofpersonsandthesamplesizeisnotcomfortablylarge,aquestionofrandomnessmayberaised.Ifpeoplewithmissingdatadonotshowsignificantdifferencesinthedependentvariable,the problemis notserious.Ifthedataarenotmissingrandomly,resultsobtainedfromdropping258
subjects will be misleading. Thus, dropping cases is not always an adequate
solutiontothemissingdataproblem.
If the independentvariableis measuredon anominal or categoricalscale,
analternativemethodistotreatindividualsinagroupwithmissinginforma-tion as another group. For quantitatively measured variables (e.g., age ), the
meanofthevaluesavailablecanbeusedforamissingvalue.Thisprinciplecanalso be applied to nominal data. It does not mean that the mean is a goodestimateforthemissingvalue,butitdoesprovideconvenienceforanalysis.
A more detailed discussion on missing data can be found in Cohen and
Cohen (1975,Chap.7 ),LittleandRubin (1987 ),Efron (1994 ),Crawfordetal.
(1995 ),Heitjan (1997 ),andSchafer (1999 ).
11.2 GENERAL STRUCTURE OF PARAMETRIC REGRESSION
MODELS AND THEIR ASYMPTOTIC LIKELIHOOD INFERENCE
When covariates are considered, we assume that the survival time, or a
function of it, has an explicit relationship with the covariates. Furthermore,whena parametricmodelisconsidered,weassumethat thesurvivaltime (or
afunctionofit )followsagiventheoreticaldistribution (ormodel )andhasan
explicit relationship with the covariates. As an example, let us consider theWeibulldistributioninSection6.2.Let x/p58(x/p16,...,x/p78)denotethepcovariates
considered. If the parameter /afii9838in the Weibull distribution is related to xas
follows:
/afii9838/p58e
/p57(a/p15/p59/afii9814/p78/p71/p14/p16a/p71x/p71)/p58exp[/p57(a/p15/p59a/p30x)]
wherea/p58(a/p16,...,a/p78)denotethecoefficientsfor x,thenthehazardfunctionof
theWeibulldistributionin (6.2.4 )canbeextendedtoincludethecovariatesas
follows:
h(t,x)/p58/afii9838/p65/afii9828t/p65/p92/p16 /p58 /afii9828t/p65/p92/p16e/p57(a/p15/p59/afii9814/p78/p71/p14/p16a/p71x/p71)/afii9828/p58/afii9828t/p65/p92/p16exp[/p57(a/p15/p59a/p30x)/afii9828] (11.2.1 )
Thesurvivorshipfunctionin (6.2.3 )becomes
S(t,x)/p58(e/p92/p82/p65)exp(/p57/afii9828(a/p15/p59a/p30x))(11.2.2 )
or
log[/p57logS(t,x)]/p58/p57/afii9828(a/p15/p59a/p30x)/p59/afii9828logt(11.2.3)
whichpresentsalinearrelationshipbetweenlog[ /p57logS(t,x)]andlogtandthe
covariates. In Sections 11.2 to 11.7 we introduce a special model called the
acceleratedfailuretimemodel.
Analogous to conventional regression methods, survival time can also be
analyzedbyusingthe acceleratedfailuretime (AFT )model.TheAFTmodel 259
forsurvivaltimeassumesthattherelationshipoflogarithmofsurvivaltime T
andthecovariatesislinearandcanbewrittenas
logT/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p59/afii9846/afii9830 (11.2.4 )
wherex/p72,j/p581,...,p, are the covariates, a/p72,j/p580, 1,...,pthe coefficients, /afii9846
(/afii9846/p570)is an unknown scale parameter, and /afii9830, the error term, is a random
variablewithknownformsofdensityfunction g(/afii9830,d)andsurvivorshipfunction
G(/afii9830,d)butunknownparameters d. Thismeansthat thesurvivalisdependent
onboththecovariateandanunderlyingdistribution g.
Considerasimplecasewherethereisonlyonecovariate xwithvalues0and
1.Then (11.2.4 )becomes
logT/p58a/p15/p59a/p16x/p59/afii9846/afii9830
LetT/p15andT/p16denote the survival times for two individuals with x/p580 and
x/p581, respectively. Then, T/p15/p58exp(a/p15/p59/afii9846/afii9830), andT/p16/p58exp(a/p15/p59a/p16/p59/afii9846/afii9830)/p58
T/p15exp(a/p16).Thus,T/p16/p57T/p15ifa/p16/p570andT/p16/p58T/p15ifa/p16/p580.Thismeansthatthe
covariatexeither ‘‘accelerates’’ or ‘‘decelerates’’ the survival time or time to
failure—thusthename acceleratedfailuretimemodels forthisfamilyofmodels.
In the following we discuss the general form of the likelihood function of
AFTmodels,theestimationproceduresoftheregressionparameters (a/p15,a,/afii9846,
andd)in(11.2.4 )andtestsofsignificanceofthecovariatesonthesurvivaltime.
The calculations of these procedures can be carried out using availablesoftwarepackagessuchasSASandBMDP.ReaderswhoarenotinterestedinthemathematicaldetailsmayskiptheremainingpartofthissectionandmoveontoSection11.3withoutlossofcontinuity.
Lett/p16,t/p17,...,t/p76betheobservedsurvivaltimesfrom nindividuals,including
exact, left-, right-, and interval-censored observations. Assume that the logsurvival time can be modeled by (11.2.4 )and let a/p30/p58(a/p16,a/p17,...,a/p78), and
b/p30/p58(a/p30,d/p30,a/p15,/afii9846).Similarto (7.1.1 ),thelog-likelihoodfunctionintermsofthe
densityfunction g(/afii9830) andsurvivorshipfunction G(/afii9830)o f/afii9830is
l(b)/p58logL(b)/p58/p26log[g(/afii9830/p71)]/p59/p26log[G(/afii9830/p71)]
/p26log[1 /p57G(/afii9830/p71)]/p59/p26log[G(/afii9834/p71)/p57G(/afii9830/p71)] (11.2.5 )
where
/afii9830/p71/p58logt/p71/p57a/p15/p57/p26/p78/p72/p14/p16a/p72x/p72/p71/afii9846
(11.2.6)
/afii9834/p71/p58log/afii9840/p71/p57a/p15/p57/p26/p78/p72/p14/p16a/p72x/p72/p71/afii9846(11.2.7)260
The first term in the log-likelihood function sums over uncensored observa-
tions, the second term over right-censored observations, and the third termover left-censored observations, and the last term over interval-censoredobservationswith /afii9840/p71asthelowerendofacensoringinterval.Notethatthelast
two summations in (11.2.5 )do not exist if there are no left- and interval-
censoreddata.
Alternatively,let
/afii9839/p71/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71i/p581,2,...,n (11.2.8 )
Then (11.2.4 )becomes
logT/p58/afii9839/p59/afii9846/afii9830 (11.2.9)
The respective alternative log-likelihood function in terms of the density
functionf(t,b)andsurvivorshipfunction S(t,b)ofTis
l(b)/p58logL(b)/p58/p26log[f(t/p71,b)]/p59/p26log[S(t/p71,b)]
/p59/p26log[1 /p57S(t/p71,b)]/p59/p26log[S(/afii9840/p71,b)/p57S(t/p71,b)] (11.2.10 )
wheref(t,b)canbederivedfrom (11.2.4 )throughthedensityfunction g(/afii9830)b y
applyingthedensitytransformationrule
f(t,b)/p58g((logt/p57/afii9839)//afii9846)
/afii9846t
(11.2.11)
andS(t,b)isthecorrespondingsurvivorshipfunction.Thevector bin(11.2.10 )
and (11.2.11 )includes the regression coefficients and other parameters of the
underlyingdistribution.
Either (11.2.5 )or(11.2.10 )can be used to derive the maximum likelihood
estimates (MLEs )of parameters in the model. For a given log-likelihood
functionl(b),theMLE b/p19isasolutionofthefollowingsimultaneousequations:
/p42(l(b))
/p42b/p71/p580 foralli (11.2.12)
Usually, there is no closed solution for the MLE b/p19from (11.2.12 )and the
Newton—RaphsoniterativeprocedureinSection7.1mustbeappliedtoobtain
b/p19. By replacing the parameters bwith its MLE b/p19inS(t/p71,b), we have an
estimated survivorship function S(t,b/p19), which takes into consideration the
covariates.
All of the hypothesis tests and the ways to construct confidence intervals
showninSection7.1canbeappliedhere.Inaddition,wecanusethefollowingteststotestlinearrelationshipsamongtheregressioncoefficients a/p16,a/p17,...,a/p78. 261
To test a linear relationship among x/p16,...,x/p78is equivalent to testing the
nullhypothesisthatthereisalinearrelationshipamong a/p16,a/p17,...,a/p78.H/p15can
bewritteningeneralas
H/p15:La/p58c (11.2.13 )
whereLisamatrixorvectorofconstantsforthelinearhypothesisand cisa
knowncolumnvectorofconstants.ThefollowingWald’sstatisticscanbeused:
X/p53/p58(La/p24/p57c)/p30[LV/p19/p63(a/p24)L/p30]/p92/p16(La/p24/p57c)( 11.2.14 )
whereV/p19/p63(a/p24)isthesubmatrixofthecovariancematrix V/p19(b/p19)correspondingto a.
UndertheH/p15andsomemildassumptions, X/p53hasanasymptoticchi-square
distributionwith /afii9840degrees of freedom,where /afii9840is the rank of L. For a given
significancelevel /afii9825,H/p15isrejectedifX/p53/p57/afii9851/p17/p74/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p74,1/p57/afii9825/2.
Forexample,if p/p583andwewishtotestif x/p16andx/p17haveequaleffectson
thesurvivaltime,thenullhypothesisis H/p15:a/p16/p58a/p17(ora/p16/p57a/p17/p580).Itiseasy
toseethatforthishypothesis,thecorresponding L/p58(1,/p571,0)and c/p580since
La/p58(1,/p571, 0)(a/p16,a/p17,a/p18)/p30/p58a/p16/p57a/p17
Letthe (i,j)elementofV/p19/p63(a/p24)be/afii9840/p71/p72;thentheX/p53definedin (11.2.14 )becomes
X/p53/p58(a/p24/p16/p57a/p24/p17)/p3(1,/p571, 0)/p1/afii9840/p16/p16/afii9840/p16/p17/afii9840/p16/p18
/afii9840/p17/p16/afii9840/p17/p17/afii9840/p17/p18
/afii9840/p18/p16/afii9840/p18/p17/afii9840/p18/p18/p2/p11
/p571
0/p2/p4/p92/p16
(a/p24/p16/p57a/p24/p17)
/p58(a/p24/p16/p57a/p24/p17)/p17
/afii9840/p16/p16/p59/afii9840/p17/p17/p572/afii9840/p16/p17
X/p53has an asymptotic chi-square distribution with 1 degree of freedom (the
rankofLis1).
Ingeneral,totestifanytwocovariateshavethesameeffectson T,thenull
hypothesiscanbewrittenas
H/p15:a/p71/p58a/p72(ora/p71/p57a/p72/p580) (11 .2.15)
The corresponding L/p58(0,...,0,1,0,...,0, /p571, 0,...,0 )andc/p580, and the
X/p53in(11.2.14 )becomes
X/p53/p58(a/p24/p71/p57a/p24/p72)/p17
/afii9840/p71/p71/p59/afii9840/p72/p72/p572/afii9840/p71/p72(11.2.16)262
whichhasanasymptoticchi-squaredistributionwith1degreeoffreedom. H/p15isrejectedifX/p53/p57/afii9851/p17/p16/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p16/p11/p16/p92/p63/p30/p17.
To test that noneof the covariatesis relatedto the survivaltime,the null
hypothesisis
H/p15:a/p580 (11.2.17 )
The respective test statistics for this overall null hypothesis are shown in
Section9.1.Forexample,thelog-likelihoodratiostatisticstherebecomes
X/p42/p58/p572[l(0,d/p19(0),a/p24/p15(0),/afii9846/p24(0))/p57l(b/p19)] (11.2.18 )
which has an asymptotic chi-square distribution with pdegrees of freedom
underH/p15, wherepis the number of covariates; d/p19(0),a/p24/p15(0), and /afii9846/p24(0)are the
MLEof d,a/p15,and /afii9846givena/p580.
11.3 EXPONENTIAL REGRESSION MODEL
Toincorporatecovariatesintotheexponentialdistribution,weuse (11.2.4 )for
thelogsurvivaltimeandlet /afii9846/p581:
logT/p71/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71/p59/afii9830/p71/p58/afii9839/p71/p59/afii9830/p71, (11.3.1 )
where /afii9839/p71/p58a/p15/p59/afii9814/p78/p72/p14/p16a/p72x/p72/p71,/afii9830/p71’sareindependentlyidenticallydistributed (i.i.d.)
random variables with a double exponential or extreme value distributionwhichhasthefollowingdensityfunction g(/afii9830) andsurvivorshipfunction G(/afii9830):
g(/afii9830)/p58exp[/afii9830/p57exp(/afii9830)] (11.3.2 )
G(/afii9830)/p58exp[/p57exp(/afii9830)] (11.3.3 )
This model is the exponential regression model. Thas the exponential
distributionwiththefollowinghazard,density,andsurvivorshipfunctions.
h(t,/afii9838/p71)/p58/afii9838/p71/p58exp/p3/p57/p1a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71/p2/p4/p58exp(/p57/afii9839/p71)(11.3.4 )
f(t,/afii9838/p71)/p58/afii9838/p71exp(/p57/afii9838/p71t) (11.3.5 )
S(t,/afii9838/p71)/p58exp(/p57/afii9838/p71t)( 11.3.6 )
where /afii9838/p71isgivenin (11.3.4 ).Thus,theexponentialregressionmodelassumesa
linear relationship between the covariates and the logarithm of hazard. Let 263
h/p71(t,/afii9838/p71) andh/p72(t,/afii9838/p72)be thehazards ofindividuals iandj; the hazardratio of
thesetwoindividualsis
h/p71(t,/afii9838/p71)
h/p72(t,/afii9838/p72)/p58/afii9838/p71/afii9838/p72/p58exp[/p57(/afii9839/p71/p57/afii9839/p72)]/p58exp/p3/p57/p78/p26
/p73/p14/p16a/p73(x/p73/p71/p57x/p73/p72)/p4(11.3.7)
This ratio is dependent only on the differences of the covariates of the two
individualsandthe coefficients.It doesnotdependonthe time t.InChapter
12we introduce aclass of modelscalled proportionalhazardmodels inwhich
the hazard ratio of any two individualsis assumed to be a time-independentconstant. The exponential regression model is therefore a special case of theproportionalhazardmodels.
The MLE of b/p58(a/p15,a/p16,...,a/p78)is a solution of (11.2.12 ), using (11.2.10 ),
wheref(t,/afii9838) andS(t,/afii9838) aregivenin (11.3.5 )and (11.3.6 ).Computerprograms
inSASorBMDPcanbeusedtocarryoutthecomputation.
In the following we introduce a practical exponential regression model.
Suppose that there are n/p58n/p16/p59n/p17/p59/p37/p59n/p73individuals in ktreatment
groups.Lett/p71/p72bethesurvivaltimeand x/p16/p71/p72,x/p17/p71/p72,...,x/p78/p71/p72thecovariatesof the
jthindividualinthe ithgroup,where pisthenumberofcovariatesconsidered,
i/p581,...,k, andj/p581,...,n/p71. Define the survivorship function for the jth
individualinthe ithgroupas
S/p71/p72(t)/p58exp(/p57/afii9838/p71/p72t) (11 .3.8)
where
/afii9838/p71/p72/p58exp(/p57/afii9839/p71/p72)and /afii9839/p71/p72/p58/p57
/p1a/p15/p71/p59/p78/p26
/p74/p14/p16a/p74x/p74/p71/p72/p2(11.3.9)
This model was proposed by Glasser (1967 )and was later investigated by
Prentice (1973 )and Breslow (1974 ). The term exp (/p57a/p15/p71)represents the
underlyinghazardofthe ithgroupwhencovariatesareignored.Itisclearthat
/afii9839/p71/p72definedin (11.3.9 )isaspecialcaseof (11.3.4 )byaddinganewindexforthe
treatment groups. To construct the likelihood function, we use the followingindicatorvariablestodistinguishcensoredobservationsfromtheuncensored:
/afii9829/p71/p72/p58/p71i ft/p71/p72uncensored
0i ft/p71/p72censored
According to (11.2.10 )and (11.3.8 ), the likelihood function for the data can
thenbewrittenas
L(/afii9838/p71/p72)/p58/p73/p26
/p71/p14/p16/p76/p71/p147
/p72/p14/p16(/afii9838/p71/p72)/p66/p71/p72exp(/p57/afii9838/p71/p72t/p71/p72)264
Substituting (11.3.9 )in the logarithm of the function above, we obtain the
log-likelihoodfunctionof a/p15/p58(a/p15/p16,a/p15/p17,...,a/p15/p73) anda/p58(a/p16,a/p17,...,a/p78):
l(a/p15,a)/p58/p73/p26
/p71/p14/p16/p76/p71/p26
/p72/p14/p16/p3/afii9829/p71/p72/p1a/p15/p71/p59/p78/p26
/p74/p14/p16a/p74x/p74/p71/p72/p2/p57t/p71/p72exp/p1a/p15/p71/p59/p78/p26
/p74/p14/p16a/p74x/p74/p71/p72/p2/p4
/p58/p73/p26
/p71/p14/p16/p3a/p15/p71r/p71/p59/p78/p26
/p74/p14/p16a/p74s/p71/p74/p57exp(a/p15/p71)/p76/p71/p26
/p72/p14/p16t/p71/p72exp/p1/p78/p26
/p74/p14/p16a/p74x/p74/p71/p72/p2/p4(11.3.10)
where
s/p71/p74/p58/p76/p71/p26
/p72/p14/p16/afii9829/p71/p72x/p74/p71/p72
isthesumofthe lthcovariatecorrespondingtotheuncensoredsurvivaltimes
intheithgroupandr/p71isthenumberofuncensoredtimesinthatgroup.
Maximumlikelihoodestimatesof a/p15/p71’sanda/p74’scanbeobtainedbysolving
thefollowingk/p59pequationssimultaneously.Theseequationsareobtainedby
takingthederivativeof l(a/p15,a)in(11.3.10 )withrespecttothe ka/p15/p71’sandpa/p74’s:
r/p71/p57exp(a/p24/p15/p71)/p76/p71/p26
/p72/p14/p16t/p71/p72exp/p1/p78/p26
/p74/p14/p16a/p24/p74x/p74/p71/p72/p2/p580i/p581,...,k(11.3.11)
/p73/p26
/p71/p14/p16/p3s/p71/p74/p57exp(a/p24/p71)/p76/p71/p26
/p72/p14/p16t/p71/p72x/p74/p71/p72exp/p1/p78/p26
/p74/p14/p16a/p24/p74x/p74/p71/p72/p2/p4/p580l/p581,...,p(11.3.12)
ThiscanbedonebyusingtheNewton —RaphsoniterativeprocedureinSection
7.1.ThestatisticalinferencesfortheMLEandthemodelarethesameasthosestated in Section 7.1. Let a/p24/p15anda/p24be the MLE of a/p15andain(11.3.10 ), and
a/p24/p15(0)be the MLE of a/p15givena/p580. According to (11.2.18 ), the difference
betweenl(a/p24/p15,a/p24)andl(a/p24/p15(0),0)can be used to test the overall null hypothesis
(11.2.17 )that none of the covariates is related to the survival time by
considering
X/p42/p58/p572(l(a/p24/p15(0),0)/p57l(a/p24/p15,a/p24)) ( 11.3.13 )
aschi-squaredistributedwith pdegreesoffreedom.A X/p42greaterthanthe100 /afii9825
percentage point of the chi-square distribution with pdegrees of freedom
indicates significant covariates. Thus, fitting the model with subsets of thecovariatesx/p16,x/p17,...,x/p78allowsselectionofsignificantcovariatesofprognostic
variables.Forexample,if p/p582,totestthesignificanceof x/p17afteradjustingfor
x/p16,thatis,H/p15:a/p17/p580,wecompute
X/p42/p58/p572[l(a/p24/p15(0),a/p24/p16(0),0) /p57l(a/p24/p15,a/p24/p16,a/p24/p17)] 265
Table 11.1 Summary Statistics for the Five Regimens
Additive
Therapy
Geometric Median
6-MP MTX Numberof Numberin Mean /p63of Mean Remission
Regimen Cycle Cycle Patients Remission WBC Age (yr)Duration
1 A-D NM 46 20 9,000 4.61 510
2 A-D A-D 52 18 12,308 5.25 409
3 NM NM 64 18 15,014 5.70 3074 NM A-D 54 14 9,124 4.30 416
5 None None 52 17 13,421 5.02 420
1,2,4 — — 152 52 10,067 4.74 4353,5 — — 116 35 14,280 5.40 340
All — — 268 87 11.711 5.02 412
Source:Breslow (1974 ).ReproducedwithpermissionoftheBiometricSociety.
/p63Thegeometricmean of x/p16,x/p17,...,x/p76isdefinedas (/afii9811/p76/p71/p14/p16x/p71)/p16/p30/p76. Itgivesalessbiasedmeasureof
centraltendencythanthearithmeticmeanwhensomeobservationsareextremelylarge.
wherea/p24/p15(0)anda/p24/p16(0) are,respectively,theMLEof a/p15anda/p16givena/p17/p580.X/p42followsthechi-squaredistributionwith1degreeoffreedom.Asignificant X/p42value indicates the importance of x/p17. This can be done automatically by a
stepwiseprocedure.Inaddition,ifoneormoreofthecovariatesaretreatments,the equality of survival in specified treatment groups can be tested bycomparingtheresultingmaximumlog-likelihoodvalues.Havingestimatedthecoefficientsa/p15/p71anda/p74,asurvivorshipfunctionadjustedforcovariatescanthen
beestimatedfrom (11.3.9 )and (11.3.8 ).
The following example, adapted from Breslow (1974 ), illustrates howthis
modelcanidentifyimportantprognosticfactors.
Example 11.1 Twohundredandsixty-eightchildrenwithnewlydiagnosed
and previously untreated ALL were entered into a chemotherapy trial. Aftersuccessful completion of an induction course of chemotherapy designed toinduce remission, the patients were randomized onto five maintenance regi-mens designed to maintain the remission as long as possible. Maintenancechemotherapyconsistedofalternatingeight-weekcyclesof6-MPandmethot-rexate (MTX )towhichactinomycin-D (A-D)ornitrogenmustard (NM)was
added.TheregimensaregiveninTable11.1.Regimen5isthecontrol.Manyinvestigators had a prior feeling that actinomycin-D was the active additivedrug; therefore, pooled regimens 1, 2, and 4 (with actinomycin-D )were
comparedtoregimens3and5 (withoutactinomycin-D ).Covariatesconsidered
were initial WBC and age at diagnosis. Analysis of variance showed thatdifferences between the regimens with respect to these variables were notsignificant. Table 11.1 shows that the regimen with lowest (highest )WBC
geometricmeanhasthelongest (shortest )estimatedremissionduration.Figure266
Figure 11.1 Remission curves of all patients by WBC at diagnosis. (From Breslow,
1974.ReproducedwithpermissionoftheBiometricSociety. )
11.1 gives three remission curves by WBC; differences in duration were
significant.It is well known that the initial WBC is an important prognosticfactorfor patientsfollowedfromdiagnosis;however,itisinterestingtoknowif this variable will continue to be important after the patient has achievedremission.
To identify important prognostic variables, model (11.3.9 )was used to
analyzetheeffectofWBCandageatdiagnosis.Previousstudies (Pierceetal.,
1969; George et al., 1973 )showed that survival is longest for children in the
middleagerange (6—8years ),suggestingthatbothlinearandquadraticterms
in age be included. The WBC was transformed by taking the commonlogarithm.Thus,thenumberofcovariatesis p/p583.Letx/p16,x/p17,andx/p18denote
log/p16/p15(WBC ), age, and age squared, and a/p16,a/p17, anda/p18be the respective
coefficients.Insteadofusingastepwisefittingprocedure,themodelwasfittedfivetimesusingdifferentnumbersofcovariates.Table11.2givestheresults.
Theestimatedregressioncoefficientswereobtainedbysolving (11.3.11 )and
(11.3.12 ). Maximum log-likelihood values were calculated by substituting the
regression coefficients with the estimates in (11.3.10 ). TheX/p42values were
computedfollowing (11.3.13 ),whichshowtheeffectofthecovariatesincluded.
Thefirst fit did not includeany covariates.The log-likelihoodso obtainedistheunadjustedvalue l(a/p24/p15(0),0)in(11.3.13 ).Thesecondfitincludedonly x/p16or
log/p16/p15(WBC ), which yields a larger log-likelihood value than the first fit.
Following (11.3.13 ),weobtain
X/p42/p58/p572(l(a/p24/p15(0),0)/p57l(a/p24/p15,a/p24/p16))/p58/p572(/p571332.925 /p591316.399 )/p5833.05 267
Table 11.2 Regression Coefficients and Maximum Log-Likelihood Values for Five Fits
RegressionCoefficient
Covariates Maximum
Fit Included Log-Likelihood b/p16b/p17b/p18/afii9851/p17df
1 None /p571332.925
2x/p16(log/p16/p15WBC ) /p571316.399 0.72 33.05 1
3x/p16,x/p17(age) /p571316.111 0.73 0.02 33.63 2
4x/p17,x/p18(agesquared ) /p571327.920 /p570.24 0.018 10.01 2
5x/p16,x/p17,x/p18/p571314.065 0.67 /p570.14 0.011 37.72 3
Source:Breslow (1974 ).ReproducedwithpermissionoftheBiometricSociety.
with1degreeoffreedom.Thehighlysignificant (p/p580.001 )X/p42valueindicates
theimportanceofWBC.Whenageandagesquaredareincluded (fit4)inthe
model,theX/p42value,10.01,islessthanthatoffit2.ThisindicatesthatWBC
isabetterpredictorthanageastheonlycovariate.TotestthesignificanceofageeffectsafteradjustingforWBC,wesubtractthelog-likelihoodvalueoffit2fromthatoffit5andobtain
X/p42/p58/p572(/p571316.399/p591314.065)/p584.668
with3 /p571/p582degreesoffreedom.Thesignificanceofthis X/p42valueismarginal
(p/p580.10).Comparingthemaximumlog-likelihoodvalueoffit2tothatoffit
5,wefindthatlogWBCaccountsforthemajorportionofthetotalcovariateeffect. Thus, log (WBC )was identified as the most important prognostic
variable. In addition, subtracting the maximum log-likelihood value of fit 5fromthatoffit3yields
X/p42/p58/p572(/p571316.111/p591314.065)/p584.092
with 1 degree of freedom. This significant (p/p580.05)value indicates that the
agerelationshipisindeedaquadraticone,withchildren6to8yearsoldhavingthe most favorable prognosis. For a complete analysis of the data, theinterestedreaderisreferredtoBreslow (1974 ).
TouseSAStoperformtheanalysis,letTbetheremissionduration,TGan
indicatorvariable (TG/p581ifinregimengroups1,2,and4;0otherwise ),CENS
asecondindicatorvariable (CENS /p580 whentis censored;1 otherwise ),and
x1,x2,andx3belog/p16/p15(WBC ),age,andagesquared,respectively.Assumethat
thedataaresavedin‘‘C: /p33RDT.DAT’’asatextfile,whichcontainssixcolumns,
and that each row (consisting of six space-separated numbers )gives the
observedT,CENS,TG,x1,x2,andx3fromachild.Forinstance,afirstrow268
inRDT.DATmaybe
500 1 0 4.079 5.2 27.04
whichrepresentsthata5.2-year-oldchildwithinitiallog/p16/p15(WBC )/p584.079who
received regimen 3 or 5 relapsed after 500 days [i.e., t/p58500, CENS /p581,
TG/p580,x1 /p584.079,x2 /p585.2,andx3 (agesquared )/p5827.04].
Forthisdataset,thefollowingSAScodecanbeusedtoperformfits1to5
inTable11.2byusingprocedureLIFEREG.
dataw1;
infile‘c: /p33rdt.dat’missover;
inputtcenstgx1x2x3;
run;proclifereg;
model1:modelt*cens (0)/p58tg/d /p58exponential;
model2:modelt*cens (0)/p58tgx1/d /p58exponential;
model3:modelt*cens (0)/p58tgx1x2/d /p58exponential;
model4:modelt*cens (0)/p58tgx2x3/d /p58exponential;
model5:modelt*cens (0)/p58tgx1x2x3/d /p58exponential;
run;
ForBMDPprocedure2Lthefollowingcodecanbeusedforfit5.
/input file /p58‘c:/p33rdt.dat’.
variables /p586.
format /p58free.
/print level /p58brief.
/variable names /p58t,cens,tg,x1,x2,x3.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58tg,x1,x2,x3.
accel /p58exponential.
/end
11.4 WEIBULL REGRESSION MODEL
To consider the effects of covariates, we use the model (11.2.4 ); that is, the
log-survival-timeofindividual iis
logT/p71/p58a/p15/p59/p78/p26
/p73/p14/p16a/p73x/p73/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71(11.4.1 )
where /afii9839/p71/p58a/p15/p59/afii9814/p78/p73/p14/p16a/p73x/p73/p71and/afii9830has the distribution defined in (11.3.2 )and 269
(11.3.3 ). This model is the Weibull regression model. Thas the Weibull
distributionwith
/afii9838/p71/p58exp/p1/p57/afii9839/p71/afii9846/p2and /afii9828/p581
/afii9846(11.4.2)
andthefollowinghazard,density,andsurvivorshipfunctionsthatarerelated
withcovariatesvia /afii9838/p71in(11.4.2 ):
h(t,/afii9838/p71,/afii9828)/p58/afii9838/p71/afii9828t/p65/p92/p16 (11.4.3)
f(t,/afii9838/p71,/afii9828)/p58/afii9838/p71/afii9828t/p65/p92/p16exp(/p57/afii9838/p71t/p65) (11 .4.4)
S(t,/afii9838/p71,/afii9828)/p58exp(/p57/afii9838/p71t/p65) (11 .4.5)
Thehazardratioofanytwoindividuals iandj,basedon (11.4.3 )and (11.4.2 ),
is
h/p71h/p72/p58exp/p1/p57/afii9839/p71/p57/afii9839/p72/afii9846/p2/p58exp/p1/p571
/afii9846/p78/p26
/p73/p14/p16a/p73(x/p73/p71/p57x/p73/p72)/p2
which is not time-dependent. Therefore, similar to the exponentialregression
model,theWeibullregressionmodelisalsoaspecialcaseoftheproportionalhazardmodels.
The following example illustrates the use of the Weibull regression model
andofcomputersoftwarepackages.
Example 11.2 Considerthetumor-freetimeinTable3.4.Supposethatwe
wishtoknowifthreedietshavethesameeffectonthetumor-freetime.LetTbe the tumor-free time; CENS be an index (or dummy )variable with
CENS /p580ifTiscensoredand1otherwise;andLOW,SATU,andUNSAbe
index variables indicating that a rat was fed a low-fat, saturated fat, orunsaturated fat diet, respectively (e.g., LOW /p581 if fed a low-fat diet; 0
otherwise ).Thedatafromthe90ratsinTable3.4canbepresentedusingthese
fivevariables.Forexample,thethreeobservationsinthefirstrowofTable3.4canberearrangedas
TCENS LOW SATU UNSA
1 4 0 11001 2 4 1010
1 1 2 1001
Assume that the rearranged data are saved in the text file ‘‘C: /p33RAT.DAT’’,
whichcontainsthedatafromthe90ratsinfivecolumnsasaboveandthefivenumbersineachrowarespace-separated.Thisdatafileisreadyforalmostall270
of the statistical software packages for parametric survival analysis currently
available,suchasSASandBMDP.Supposethatthetumor-freetimefollowstheWeibulldistributionandthefollowingWeibullregressionmodelisused:
logT/p71/p58a/p15/p59a/p16SATU/p71/p59a/p17UNSA/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71(11.4.6 )
where /afii9830/p71has a double exponential distribution as defined in (11.3.2 )and
(11.3.3 ).Notethatfrom (11.4.3 )and (11.4.2 ),
logh(t,/afii9838/p71,/afii9828)/p58log/afii9838/p71/p59log(/afii9828t/p65/p92/p16)
/p58/p57/afii9839/p71/afii9846/p59log(/afii9828t/p65/p92/p16)
/p58/p57a/p15/p57a/p16SATU/p71/p57a/p17UNSA/p71/afii9846/p59log(/afii9828t/p65/p92/p16) (11.4.7)
Denotethehazardfunctionofaratfedanunsaturated,saturated,andlow-fat
diet ash/p83,h/p81, andh/p74, respectively. From (11.4.7 ), logh/p83/p58(/p57a/p15/p57a/p17)/
/afii9846/p59log(/afii9828t/p65/p92/p16), logh/p81/p58(/p57a/p15/p57a/p16)//afii9846/p59log(/afii9828t/p65/p92/p16), and log h/p74/p58/p57a/p15/
/afii9846/p59log(/afii9828t/p65/p92/p16). Thus, the logarithm of the hazard ratio of rats fed a low-fat
dietandthosefedasaturatedfatdietislog (h/p74/h/p81)/p58a/p16//afii9846,andthesimilarratios
of rats fed a low-fat diet and those an unsaturated fat diet, and of rats fed asaturated fat diet and those fed an unsaturated fat diet are, respectively,log(h/p74/h/p83)/p58a/p17//afii9846andlog (h/p81/h/p83)/p58(a/p17/p57a/p16)//afii9846. Theseratiosareconstantsand
are independent of time. Therefore, to test the null hypothesis that the threediets have an equal effect on tumor-free time is equivalent to testing thefollowing three hypotheses: H/p15:h/p74/h/p81/p581o ra/p16/p580,H/p15:h/p74/h/p83/p581, ora/p17/p580,
andH/p15:h/p81/h/p83/p581ora/p17/p58a/p16.ThestatisticdefinedinSection9.1.1canbeused
to test the first two null hypotheses,and the statistic defined in (11.2.16 )can
be used for the third one. Failure to reject a null hypothesis impliesthat thecorrespondinglog-hazard ratio is not statistically different from zero; that is,therearenostatisticallysignificantdifferencesbetweenthetwocorrespondingdiets. For example, failure to reject H/p15:a/p16/p580 means that there are no
significantdifferencesbetweenthe hazardsfor ratsfeda low-fatdietandratsfedasaturatedfatdiet.Whenallthreehypotheses H/p15:a/p16/p580,H/p15:a/p17/p580,and
H/p15:a/p17/p58a/p16are rejected, we conclude that the three diets have significantly
different effects on tumor-free time. Furthermore, a positive (negative )es-
timated implies that the hazard of a rat fed a low-fat diet is exp (a/p16//afii9846) times
higher (lower )thanthat ofa rat feda saturatedfat diet. Similarly,a positive
(negative )estimateda/p17and (a/p17/p57a/p16)imply, respectively, the hazard of a rat
fed a low-fat diet is exp (a/p17//afii9846)times higher (lower )than that of a rat fed an
unsaturated fat diet, and the hazard of a rat fed a saturated fat diet isexp[ (a/p17/p57a/p16)//afii9846] timeshigher (lower )thanthatofaratfedanunsaturatedfat
diet. 271
To estimate the unknown coefficients, a/p16,a/p17,a/p15, and /afii9846, we construct the
log-likelihood function by replacing /afii9839in(11.4.2 ),(11.4.4 ), and (11.4.5 )with
(11.4.6 ).Next,placetheresulting f(t/p71,/afii9838/p71,/afii9828) andS(t/p71,/afii9838/p71,/afii9828) inthelog-likelihood
function (11.2.10 ). The log-likelihood function for the observed 90 exact or
right-censoredtumor-freetimes, t/p16,t/p17,...,t/p24/p15,inthethreedietgroupsis
l(a/p15,a/p16,a/p17,/afii9828)/p58/p26log[f(t/p71,/afii9838/p71,/afii9828)]/p59/p26log[S(t/p71,/afii9838/p71,/afii9828)]
/p58/p26[log/afii9828/p59(/afii9828/p571)logt/p71/p57/afii9828/afii9839/p71/p57t/p65/p71exp(/p57/afii9828/afii9839/p71)]
/p59/p26[/p57t/p65/p71exp(/p57/afii9828/afii9839/p71)]
/p58/p26/p43log/afii9828/p59(/afii9828/p571)logt/p71/p57/afii9828(a/p15/p59a/p16SATU/p9/p59a/p17UNSA/p9)
/p57t/p65/p71exp[/p57/afii9828(a/p15/p59a/p16SATU/p9/p59a/p17UNSA/p9)]/p44
/p59/p26/p43/p57t/p65/p71exp[/p57/afii9828(a/p15/p59a/p16SATU/p9/p59a/p17UNSA/p9)]/p44
The first term in the log-likelihood function sums over the uncensored
observations,andthesecondtermsumsovertheright-censoredobservations.TheMLE(a/p24/p16,a/p24/p17,a/p24/p15,/afii9846/p24)of(a/p16,a/p17,a/p15,/afii9846) where /afii9846/p581//afii9828isasolutionof (11.2.12 )
with the above log-likelihood function by applying the Newton —Raphson
iterative procedure. The results from SAS are shown in Table 11.3, whereINTERCPT /p58a/p15and SCALE /p58/afii9846. The MLE /afii9846/p24/p580.43,a/p24/p16/p58/p570.394,
a/p24/p17/p58/p570.739,anda/p24/p17/p57a/p24/p16/p58/p570.345.H/p15:a/p16/p580(orh/p74/h/p81/p581),H/p15:a/p17/p580(or
h/p74/h/p83/p581),andH/p15:a/p17/p57a/p16/p580(orh/p81/h/p83/p581)arerejectedat significancelevel
p/p580.0065,p/p580.0001, andp/p580.0038, respectively. The conclusion that the
dataindicate significantdifferencesamong thethree diets is the same as thatobtained in Chapter 3 using the k-sample test. Furthermore, both a/p24/p16anda/p24/p17are negative and h/p19/p74/h/p19/p81/p58exp(a/p24/p16//afii9846/p24)/p58exp(/p570.916 )/p580.40,h/p19/p74/h/p19/p83/p58exp(a/p24/p17/
/afii9846/p24)/p58exp(/p571.719 )/p580.18, andh/p19/p81/h/p19/p83/p58exp((a/p24/p17/p57a/p24/p16)//afii9846/p24)/p58exp(/p570.802 )/p580.45.
Thus,basedonthedataobserved,thehazardofratsfedalow-fatdietis40%and18%ofthehazardofratsasaturatedfatdietandanunsaturatedfatdiet,respectively, and the hazard of rats fed a saturated fat diet is 45%of that ofratsfedanunsaturatedfatdiet.
Thesurvivorshipfunctionin (11.4.5 )canbeestimatedbyusing (11.4.2 )and
theMLEofa/p15,a/p16,a/p17,and /afii9846:
S/p19(t,/afii9838,/afii9828)/p58exp(/p57/afii9838/p19t
/afii9828/p24)
/p58exp/p7/p57exp/p3/p571
/afii9846/p24(a/p24/p15/p59a/p24/p16SATU /p59a/p24/p17UNSA )/p4t1//afii9846/p24/p8
/p58exp[/p57exp(/p5712.56 /p590.92/p59SATU /p591.72/p59UNSA )t/p17/p13/p18/p18]
Based onS/p19(t,/afii9838/p71,/afii9828), we can estimate the probabilityof surviving a given time
for rats fed with any of the diets. For example, for rats fed a low-fat diet,272
Table 11.3 Analysis Results for Rat Data in Table 3.4 Using a Weibull
Regression Model
Regression Standard
Variable Coefficient Error X/p42pexp(a/p24/p71//afii9846/p24)
INTERCPT(a/p24/p15)5.400 0.113 2297 .610 0.0001
TRTSA (a/p24/p16) /p570.394 0.145 7.407 0.0065 0.40
TRTUS (a/p24/p17) /p570.739 0.140 28.049 0.0001 0.18
SCALE (/afii9846/p24) 0.430 0.043
a/p24/p17/p57a/p24/p16/p570.345 0.119 8.355 0.0038 0.45
(SATU /p580andUNSA /p580),theprobabilityofbeingtumor-freefor200daysis
S/p19/p42/p45/p53(200)/p58exp[/p57exp(/p5712.56 )(200)/p17/p13/p18/p18]
/p58exp[/p570.00000353 (200)/p17/p13/p18/p18]/p580.132
and for rats fed an unsaturated fat diet, (SATU /p580 and UNSA /p581), the
probabilityis0.011.
FollowingistheSAScodeusedtoobtainTable11.3,basedontheWeibull
regressionmodelin (11.4.6 ).
dataw1;
infile‘c: /p33rat.dat’missover;
inputtcenslowsatuunsa;
run;procliferegcovout;
modelt*cens (0)/p58satuunsa/d /p58weibull;
run;
TherespectiveBMDPprocedure2Lcodebasedon (11.4.6 )is
/input file /p58‘c:/p33rat.dat’.
variables /p585.
format /p58free.
/print level /p58brief.
/variable names /p58t,cens,low,satu,unsa.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58satu,unsa.
accel /p58weibull.
/end 273
11.5 LOGNORMAL REGRESSION MODEL
Let/afii9830in(11.2.4 )be the standard normal random variable with the density
functiong(/afii9830) andsurvivorshipfunction G(/afii9830),
g(/afii9830)/p58exp(/p57/afii9830/p17/2)
/p402/afii9843(11.5.1 )
G(/afii9830)/p581/p57/afii9818(/afii9830)/p581/p571
/p402/afii9843/p16/p67
/p92/p27e/p92/p86/p130/p30/p17dx (11.5.2 )
where /afii9818is the cumulative distribution function of the standard normal
distribution. Then the model defined by (11.2.4 )for the survival time Tof
individuali,
logT/p71/p58a/p15/p59/p78/p26
/p73/p14/p16a/p73x/p73/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71
isthe lognormalregressionmodel. Thasthe lognormaldistributionwiththe
densityfunction
f(t,/afii9839/p71,/afii9846/p17)/p58exp[/p57(logt/p57/afii9839/p71)/p17/2/afii9846/p17]
/p402/afii9843/afii9846t(11.5.3)
andthesurvivorshipfunction
S(t,/afii9839/p71,/afii9846/p17)/p581/p57/afii9818/p1logt/p57/afii9839/p71/afii9846 /p2(11.5.4 )
It can be shown that the hazard function h(t,/afii9846,a/p15,a/p16,...,a/p78)ofTwith
covariatex/p16,x/p17,...,x/p78and unknown parameters and coefficients /afii9846,a/p15,
a/p16,...,a/p78canbewrittenas
logh(t,/afii9846,a/p15,a/p16,...,a/p78)/p58logh/p15[texp(/p57/afii9839)]/p57/afii9839 (11.5.5 )
whereh/p15(·)isthehazardfunctionofanindividualwithallcovariatesequalto
zero. Equation (11.5.5 )indicates that h(t,/afii9846,a/p15,a/p16,...,a/p78)is a function of h/p15evaluatedattexp(/p57/afii9839), not independentof t. Thus, the lognormal regression
modelisnotaproportionalhazardsmodel.
Example 11.3 Considerthesurvivaltimedatafrom30patientswithAML
inTable11.4.Twopossibleprognosticfactorsorcovariates,age,andcellular-274
Table 11.4 Survival Times and Data for Two Possible
Prognostic Factors of 30 AML Patients
SurvivalTime x/p16x/p17SurvivalTime x/p16x/p17
18 0 0 8 1 0
90 1 2 1 1
28/p59 00 2 6 /p59 10
31 0 1 10 1 1
39/p59 014 10
19/p59 013 10
45/p59 014 10
60 1 1 8 1 180 1 8 1 1
15 0 1 3 1 1
23 0 0 14 1 1
28/p59 003 10
70 1 1 3 1 1
12 1 0 13 1 1
91 0 3 5 /p59 10
itystatusareconsidered:
x/p16/p58/p71 if patient is /p4650 years old
0 otherwise
x/p17/p58/p71 if cellularity of marrowclot section is 100%
0 otherwise
Letususethelognormalregressionmodel
logT/p71/p58a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71/p59/afii9846/afii9830/p71(11.5.6 )
and
/afii9839/p71/p58a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71(11.5.7)
Theunknowncoefficientsandparameter a/p16,a/p17,a/p15,/afii9846needtobeestimated.
Weconstructthelog-likelihoodfunctionbyreplacing /afii9839in(11.5.3 )and (11.5.4 )
with (11.5.7 ), then replacing f(t/p71,/afii9839,/afii9846/p17) andS(t/p71,/afii9839,/afii9846/p17) in the log-likelihood
function (11.2.5 )with their expression (11.5.3 )and (11.5.4 ), respectively. The
resultinglog-likelihoodfunctionfortheexactandright-censoredsurvivaltimes 275
Table 11.5 Asymptotic Likelihood Inference for Data on 30 AML Patients Using a
Lognormal Regression Model
Regression Standard
Variable /p63 Coefficient Error X/p42p
INTERCPT(a/p15) 3.3002 0.3750 77.4675 0.0001
x/p16(a/p16) /p571.0417 0.3605 8.3475 0.0039
x/p17(a/p17) /p570.2687 0.3568 0.5672 0.4514
SCALE (/afii9846) 0.9075 0.1409
/p63x/p16/p581 if patient /p4650 yearsold, and0 otherwise; x/p17/p581 if cellularityof marrowclot section is
100%,and0otherwise.observedfromthe30patientswithAMLis
l(a/p15,a/p16,a/p17,/afii9846)/p58/p26/p3/p57(logt/p71/p57/afii9839/p71)/p17
2/afii9846/p17/p57log(/p402/afii9843/afii9846t/p71)/p4
/p59/p26/p7log/p31/p57/afii9818/p1logt/p71/p57/afii9839/p71/afii9846 /p2/p4/p8
/p58/p26/p7/p57[logt/p71/p57(a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71)]/p17
2/afii9846/p17/p57log(/afii9846t/p71/p402/afii9843)/p8
/p59/p26/p7log/p31/p57/afii9818(logt/p71/p57(a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71)
/afii9846 /p4/p8
The first term in the log-likelihood function sums over the uncensored
observations,and the second sumsover the right-censoredobservations.TheMLE (a/p24/p16,a/p24/p17,a/p24/p15,/afii9846/p24)o f(a/p16,a/p17,a/p15,/afii9846) can be obtained by applying the
Newton—Raphsoniterativeprocedure.The hypothesis-testingproceduresdis-
cussed in Section 9.1.2 can be used to test whether the coefficients a/p16anda/p17areequaltozero.Table11.5showsthat a/p16issignificantly (p/p580.0039 )different
fromzero,while a/p17isnot (p/p580.4514 ).Thesignsoftheregressioncoefficients
indicatethatageover50yearshassignificantlynegativeeffectsonthesurvivaltime,whilea100%cellularityofmarrowclotsectionalsohasanegativeeffect;however,theeffectisnotofsignificantimportancetothesurvivaltime.
LetTbethesurvivaltimeandCENSbeanindex (ordummy )variable
with CENS /p580i fTis censored and 1 otherwise. Assume that the data are
saved in a text file ‘‘C: /p33AML.DAT’’ with four numbers in each row, space-
separated,whichcontainssuccessively T,CENS,x1,andx2.
LetTbethesurvivaltimeandCENSbeanindex (ordummy )variablewith
CENS /p580ifTiscensoredand1otherwise.Assumethatthedataaresavedin
a text file ‘‘C: /p33AML.DAT’’ with four numbers in each row, space-separated,
which contains successively T, CENS, x1, and x2. The following SAS code is
usedtoobtaintheresultsinTable11.5.276
dataw1;
infile‘c: /p33aml.dat’missover;
inputtcensx1x2;
run;
proclifereg;
model1:modelt*cens (0)/p58x1x2/d /p58lnormal;
run;
IfBMDPisused,thefollowing2Lcodeissuggested.
/input file /p58‘c:/p33aml.dat’.
variables /p584.
format /p58free.
/print level /p58brief.
/variable names /p58t,cens,x1,x2.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58x1,x2.
accel /p58lnormal.
/end
11.6 EXTENDED GENERALIZED GAMMA REGRESSION MODEL
Inthissectionweintroducea regressionmodelthat isbasedonanextended
formofthegeneralizedgammadistributiondefinedinSection6.4.Assumethatthesurvivaltime Tofindividualiandcovariates x/p16,...,x/p78havetherelation-
shipgivenin (11.4.1 ),where /afii9830hasthelog-gammadistributionwiththedensity
functiong(/afii9830) andsurvivorshipfunction G(/afii9830):
g(/afii9830)/p58/p34/afii9829/p34[exp (/afii9829/afii9830)//afii9829/p17]/p16/p30/p66/p130exp[/p57exp(/afii9829/afii9830)//afii9829/p17]
/afii9772(1//afii9829/p17)(11.6.1 )
G(/afii9830)/p58/p7I/p3exp(/afii9829/afii9830)
/afii9829/p17,1
/afii9829/p17/p4if/afii9829/p580
1/p57I/p3exp(/afii9829/afii9830)
/afii9829/p17,1
/afii9829/p17/p4if/afii9829/p570/p57/p45/p58/afii9830/p58 /p59/p45(11.6.2 )
(11.6.3 )
This model is the extended generalized gamma regression model. It can be
shown thatThas the extended generalized gamma distribution with the
densityfunction
f(t,/afii9825,/afii9838,/afii9828)/p58/p34/afii9825/p34/afii9828/p65/afii9838/p63/p65/p71t/p63/p65/p92/p16exp[/p57/afii9828(/afii9838/p71t)/p63]
/afii9772(/afii9828)(11.6.4)
andsurvivorshipfunction
S(t,/afii9825,/afii9838,/afii9828)/p58/p7I(/afii9828(/afii9838/p71t)/p63,/afii9828)
1/p57I(/afii9828(/afii9838/p71t)/p63,/afii9828)if/afii9825/p580
if/afii9825/p570(11.6.5 )
(11.6.6 ) 277
where
/afii9838/p71/p58exp(/p57/afii9839/p71) /afii9825/p58/afii9829
/afii9846/afii9828/p581
/afii9829/p17(11.6.7 )
/afii9772(x) isthecompletegammafunctiondefinedin (6.2.9 ),I(a,x)istheincomplete
gamma function defined in (6.4.4 ), and /afii9829is a shape parameter. We used the
extended generalized gamma distribution in (11.6.4 )here because it is the
distribution used in SAS. The derivation is left to the reader as an exercise
(Exercise11.12 ).
The estimation procedures for the parameters, regression coefficients, and
the covariate adjusted survivorship function are similar to those discussed inSections11.3and11.4.
Example 11.4 Consider the survival times (T)in days and a set of
prognostic factors or covariates from 137 lung cancer patients, presented inAppendix I of Kalbfleisch and Prentice (1980 ). The covariates include the
Karnofskymeasureofthe overallperformancestatus (KPS )ofthe patientat
entry into the trial, time in months from diagnosis to entry into the trial(DIAGTIME ), age in years (AGE ), prior therapy (INDPRI, yes or no ),
histological type of tumor, and type of therapy. There are four histologicaltypes of tumor: adeno, small, large, and squamous cell and two types oftherapies: standard and experimental. The values of KPS have the followingmeanings: 10—30 completely hospitalized, 40 —60 partial confinement, 70 —90
able to care for self. Assume that the survival time follows the extendedgeneralizedgammaregressionmodel,wewishtoidentifythe mostsignificantprognosticvariables.
First we define several index (or dummy )variables for the categorical
variablesandthecensoringstatus.LetCENS /p580whenthesurvivaltime Tis
censoredand1otherwise;INDADE /p581,INDSMA /p581,andINDSQU /p581if
the type of cancer cell is adeno, small, and squamous, respectively, and 0otherwise; INDTHE /p581 if the standard therapy is received and 0 otherwise;
andINDPRI /p581ifthereisapriortherapyand0otherwise.Themodelis
logT/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17AGE/p71/p59a/p18DIAGTIME/p71/p59a/p19INDPRI/p71
/p59a/p20INDTHE/p71/p59a/p21INDADE/p71/p59a/p22INDSMA/p71/p59a/p23INDSQU/p71/p59/afii9846/afii9830/p71
(11.6.8 )
wherethedensityfunctionof /afii9830/p71isdefinedin (11.6.1 ).Thus,
/afii9839/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17AGE/p71/p59a/p18DIAGTIME/p71/p59a/p19INDPRI/p71/p59a/p20INDTHE/p71
/p59a/p21INDADE/p71/p59a/p22INDSMA/p71/p59a/p23INDSQU/p71(11.6.9 )
Toestimatea/p16,...,a/p23,/afii9829,a/p15,and /afii9846,weconstructthelog-likelihoodfunctionby
replacing /afii9839in(11.6.7 )and (11.6.4 )—(11.6.6 )with (11.6.9 ),thenreplacing f(t/p71,b)
andS(t/p71,b)inthelikelihoodfunction (11.2.10 )bythosein (11.6.4 )and (11.6.5 )
or(11.6.6 ). The MLE (a/p24/p16,...,a/p24/p23,/afii9829/p19,a/p24/p15,/afii9846/p24)of (a/p16,...,a/p23,/afii9829,a/p15,/afii9846) can be278
Table 11.6 Asymptotic Likelihood Inference on Lung Cancer Data Using a Generalized
Gamma Regression Model
Regression Standard
Variable Coefficient Error X/p42p
INTERCPT(a/p15)2 .176 0 .719 9 .143 0.003
INDADE(a/p21) /p570.759 0.286 7.034 0.008
INDSMA(a/p22) /p570.594 0.264 5.059 0.025
INDSQU(a/p23) 0.150 0.291 0.266 0.606
KPS(a/p16) 0.034 0.005 46.443 0.000
AGE(a/p17) 0.008 0.009 0.845 0.358
DIAGTIME(a/p18) 0.000 0.009 0.001 0.980
INDPRI(a/p19) /p570.089 0.216 0.171 0.679
INDTHE(a/p20) 0.168 0.185 0.823 0.364
SCALE (/afii9846) 1.000 0.071
SHAPE (/afii9829) 0.450 0.223
INTERCPT(a/p15) 2.748 0.396 48.247 0.000
INDADE(a/p21) /p570.766 0.280 7.492 0.006
INDSMA(a/p22) /p570.534 0.258 4.284 0.039
INDSQU(a/p23) 0.144 0.280 0.264 0.608
KPS(a/p16) 0.033 0.005 45.497 0.000
SCALE (/afii9846) 1.004 0.070
SHAPE (/afii9829) 0.473 0.206
obtained in a manner similar to that used in Examples 11.2 and 11.3. The
hypothesis-testing procedure defined in Section 11.2 can be used to testwhetherthecoefficients a/p16,a/p17,...,a/p23areequaltozero.ThefirstpartofTable
11.6 shows the results from SAS (where INTERCPT /p58a/p15, SCALE /p58/afii9846, and
SHAPE /p58/afii9829).
Table 11.6 shows that a/p16,a/p21, anda/p22are significantly (p/p580.05)different
fromzero,whereastheothercovariatesarenot (p/p570.05).Thatis,onlyKPS
and the type of cancer cell have significant effects on the survival time. Inparticular, adeno cell carcinoma and small cell carcinoma have significantnegativeeffectsonsurvivaltime.PatientswhohavebetterKarnofskyperform-ance status have a longer survival time. If we wish to include only KPS andcelltypeinthemodel,thelowerpartofTable11.6givestheresults.
Assume that the coded data are saved in ‘‘C: /p33LCANCER.DAT’’as a text
file with 10 numbers in a row, space-separated, which contains data for T,
CENS, KPS, AGE, DIAGTIME, INDPRI, INDTHE, INDADE, INDSMA,andINDSQU,inthatorder.TheSAScodeusedtoobtainTable11.6is
dataw1;
infile‘c: /p33lcancer.dat’missover;
inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu;
run; 279
proclifereg;
Model1: modelt*cens (0)/p58kpsage diagtimeindpriindtheindadeindsmaindsqu /
d/p58gamma;
Model2:modelt*cens (0)/p58kpsindadeindsmaindsqu/d /p58gamma;
run;
11.7 LOG-LOGISTIC REGRESSION MODEL
Assumethattherelationshipbetweenthesurvivaltime T/p71forindividualiand
aset of covariates, x/p16,...,x/p78canbe expressedbythe AFTmodel in (11.4.1 ),
where /afii9830/p71hasalogisticdistributionwiththedensityfunction
g(/afii9830)/p58exp(/afii9830)
[1/p59exp(/afii9830)]/p17(11.7.1 )
andsurvivorshipfunction
G(/afii9830)/p581
1/p59exp(/afii9830)(11.7.2 )
This model is the log-logistic regression model. Then Thas the log-logistic
distribution defined in Section 6.5. The parameter /afii9825in the distribution is a
functionofthecovariates:
/afii9825/p71/p58exp/p1/p57/afii9839/p71/afii9846/p2/afii9828/p581
/afii9846(11.7.3)
Substituting (11.7.3 )inthesurvivorshipfunctionin (6.5.2 ),weobtain
logS(t,b)
1/p57S(t,b)/p58/p57log(/afii9825t/p65)/p58/afii9839
/afii9846/p57/afii9828logt(11.7.4)
or
logS(t,b)
1/p57S(t,b)/p58a/p15/afii9846/p591
/afii9846/p78/p26
/p73/p14/p16a/p73x/p73/p57/afii9828logt(11.7.5)
whereb/p58(a/p15,a/p16,...,a/p78,/afii9846).SinceS(t/p71,b)istheprobabilityofsurvivinglonger
thant,S(t/p71,b)/[1/p57S(t/p71,b)]istheoddsofsurvivinglongerthan t.LetOR/p71and
OR/p72denote the odds of surviving longer than tfor individuals iandj,
respectively.Thelogarithmoftheoddsratiois
logOR/p71OR/p72/p581
/afii9846/p78/p26
/p73/p14/p16a/p73(x/p73/p71/p57x/p73/p72) (11 .7.6)
Thisratio isindependentoftime.Therefore,thelog-logisticregressionmodel
isaproportionaloddsmodel,notaproportionalhazardsmodel.280
Example 11.5 Wefitthelog-logisticregressionmodelabovetothedatain
Example11.6.1usingonlyKPSandthethreecancercelltypeindexvariables.Thatis,
logT/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71/p59/afii9846/afii9830/p71
(11.7.7 )
wherethedensityfunctionof /afii9830/p71isdefinedin (11.7.1 ).Thus,
/afii9839/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71(11.7.8 )
Toestimate b/p58(a/p15,a/p16,...,a/p19,/afii9846)/p30,weconstructthelog-likelihoodfunctionby
using the /afii9825and/afii9828in(11.7.3 )as parameters in the density and survivorship
functions of the log-logistic distribution in Section 6.5. The resulting log-likelihoodfunctionforthe137observedexactorright-censoredsurvivaltimesis
l(b)/p58/p26
/p7/p57/afii9839/p71/afii9846/p57log/afii9846/p591/p57/afii9846
/afii9846logt/p71/p572log /p31/p59exp/p1/p57/afii9839/p71/afii9846/p2t/p16/p30/p78/p71/p4/p8
/p59/p26/p7/p57log/p31/p59exp/p1/p57/afii9839/p71/afii9846/p2t/p16/p92/p78/p71/p4/p8
/p58/p26/p1/p571
/afii9846(a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71)
/p57log/afii9846/p591/p57/afii9846
/afii9846logt/p71
/p572log /p71/p59exp/p3/p571
/afii9846(a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71
/p59a/p19INDSQU/p71)/p4t/p16/p30/p78/p71/p8/p2
/p57/p26/p1log/p71/p59exp/p3/p571
/afii9846(a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71
/p59a/p19INDSQU/p71)/p4t/p16/p30/p78/p71/p8/p2
The first term in the log-likelihood function sums over the uncensored
observations,and the second sumsover the right-censoredobservations.TheMLE(a/p24/p16,...,a/p24/p19,a/p24/p15,/afii9846/p24)o f(a/p16,...,a/p19,a/p15,/afii9846) aregiveninTable11.7,withtheir- 281
Table 11.7 Asymptotic Likelihood Inference on Lung Cancer Data Using a
Log-Logistic Regression Model
Regression Standard
Variable Coefficient Error X/p42pexp(a/p71/a)
INTERCPT(a/p15)2.451 0.344 50.911 0.000 —
INDADE(a/p17) /p570.749 0.261 8.217 0.004 0.275
INDSMA(a/p18) /p570.661 0.240 7.565 0.006 0.321
INDSQU(a/p19)0.029 0.264 0.012 0.913 1.051
KPS(a/p16) 0.036 0.004 66.885 0.000 1.064
SCALE (/afii9846) 0.581 0.043 — — —
standarderrors, likelihood ratio test statistics (X/p42), andpvalues. The results
aresimilartothoseobtainedfromfittingthegeneralgammaregressionmodelinExample11.4.
In addition, using (11.7.5 )and (11.7.6 ), we can obtain odds ratios for the
covariates. For example, let the odds of surviving to time tfor four patients
withthesameKPSbutdifferentcelltype (adeno,small,squamousandlarge )
bedenotedbyOR/p31/p34,OR/p49/p43,OR/p49/p47,andOR/p42/p31,respectively;thenthelog-odds
ratiosoftheindividualswithadeno,small,andsquamouscelltypestotheonewithlargecelltypeare,respectively,
logOR/p31/p34OR/p42/p31
/p58a/p17/afii9846logOR/p49/p43OR/p42/p31/p58a/p18/afii9846logOR/p49/p47OR/p42/p31/p58a/p19/afii9846
Replacinga/p17,a/p18,a/p19,and /afii9846withtheirestimates,wehave
OR/p31/p34OR/p42/p31/p58exp/p1a/p24/p17/afii9846/p24/p2/p580.275
OR/p49/p43OR/p42/p31/p58exp/p1a/p24/p18/afii9846/p24/p2/p580.321
OR/p49/p47OR/p42/p31/p58exp/p1a/p24/p19/afii9846/p24/p2/p581.051
Theseresultsmeanthatinlungcancerpatients,personswithadenoandsmall
cell type have odds of only about one-fourth and one-third, respectively, ofthose with large cell type. The odds of persons with large cell carcinoma arenotsignificantlydifferentfromthoseofpatientswithsquamouscellcarcinoma.Further,whenignoringcelltype,exp (a/p24/p16//afii9846/p24)representsanincrease (ordecrease )
in the odds for any 1-unit increase in the KPS measure. In this case,282
exp(a/p24/p16//afii9846/p24)/p581.064; thus for a 1-unit increase in the KPS measure, the odds
increaseby6.4%.Theresultsare,ingeneral,consistentwiththoseobtainedinExample11.4.
ThefollowingSAScodecanbeusedtoobtaintheresultsinTable11.7.
dataw1;
infile‘c: /p33lcancer.dat’missover;
inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu;
run;proclifereg;
modelt*cens (0)/p58kpsindadeindsmaindsqu/d /p58llogistic;
run;
ThefollowingBMDP2Lcodeisalsoapplicable.
/input file /p58‘c:/p33lcancer.dat’.
variables /p5810.
format /p58free.
/print level /p58brief.
/variable names /p58t,cens,kps,age,diagtime,indpri,indthe,indade,indsma,indsqu.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58kps,indade,indsma,indsqu.
accel /p58llogistic.
/end
11.8 OTHER PARAMETRIC REGRESSION MODELS
Inthissectionwediscusstwomodelsinwhichthesurvivaltime Tisassumed
tofollowtheexponentialdistributionwithdensityandsurvivorshipfunctionsasdefinedin (6.1.1 )and (6.1.3 ),respectively,andthemeansurvivaltime1/ /afii9838or
hazardrate /afii9838hasthefollowinglinearrelationshipwiththecovariates:
Model1:1
/afii9838/p71/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71
Model2: /afii9838/p71/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71
Model 1 is considered by Feigl and Zelen (1965 )and extended to include
censoreddatabyZippinandArmitage (1966 ).Model2isusedbyByaretal.
(1974 ). 283
Model 1
Supposethatnpatientsareenteredinastudy; rofthesedieand s/p58n/p57rare
still alive at the end of the study. Let t/p16,...,t/p80be the exact survival times of
therdeaths andt/p62/p16,...,t/p62/p81be thescensoring times. Furthermore, let x/p72/p71,
i/p581,...,n,j/p581,...,p, be the observed value of the jth covariate of the ith
patient.Themodelassumesthat themeansurvivaltime islinearlyrelatedtothecovariates:
1
/afii9838/p71/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71/p58/p78/p26
/p72/p14/p15a/p72x/p72/p71(11.8.1)
wherex/p15/p71/p891.Theterma/p15representstheunderlyinghazardinthesensethat
1/a/p15isthehazardrate /afii9838/p71whencovariatesareignoredorall x/p72/p71’sarezero.Then
thelikelihoodfunctionofthe nsurvivaltimesunderthemodel (11.8.1 )canbe
writtenas
L(a/p15,a/p16,...,a/p78)/p58/p80/p147
/p71/p14/p16/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2/p92/p16exp/p3/p57t/p71/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2/p92/p16/p4
/p59/p81/p147
/p73/p14/p16exp/p3/p57t/p62/p73/p1/p78/p26
/p72/p14/p15a/p72x/p72/p73/p2/p92/p16/p4(11.8.2)
Thelog-likelihoodisthen
l(a/p15,a/p16,...,a/p78)/p58/p57/p80/p26
/p71/p14/p16log/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2/p57/p80/p26
/p71/p14/p16t/p71/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2/p92/p16
/p57/p80/p26
/p73/p14/p16t/p62/p73/p1/p78/p26
/p72/p14/p15a/p72x/p72/p73/p2/p92/p16(11.8.3)
The maximumlikelihood estimates of a/p72,j/p580, 1,...,p/p72, may be obtained
bysolvingsimultaneouslythe p/p591equations:
/p57/p80/p26
/p71/p14/p16/p1/p78/p26
/p72/p14/p16a/p72x/p72/p71/p2/p59/p80/p26
/p71/p14/p16t/p71/p1/p78/p26
/p72/p14/p16a/p72x/p72/p71/p2/p92/p17/p59/p81/p26
/p73/p14/p16t/p62/p73/p1/p78/p26
/p72/p14/p16a/p72x/p72/p73/p2/p92/p17/p580
/p57/p80/p26
/p71/p14/p16x/p72/p71/p1/p78/p26
/p72/p14/p16a/p72x/p72/p71/p2/p92/p16/p59/p80/p26
/p71/p14/p16t/p71x/p72/p71/p1/p78/p26
/p72/p14/p16a/p72x/p72/p71/p2/p92/p17
/p59/p81/p26
/p73/p14/p16t/p62/p73x/p72/p73/p1/p78/p26
/p72/p14/p16a/p72x/p72/p73/p2/p92/p17/p580j/p581,...,p
(11.8.4)
Again,thiscanbedonebyNewton —Raphsoniterativeproceduresdescribedin
Section7.1.284
AfterobtainingtheMLE, a/p24/p72,j/p580,1,...,p,thelog-likelihoodfunctioncan
beusedtotestthesignificanceofthecovariates.TheprocedureisexactlythesameasthoseusedinExample11.1.
The survivorship function (for theith patient )adjusted for the covariates
canbeobtainedfrom
S/p19/p71(t)/p58exp(/p57/afii9838/p19/p71t)
/p58exp/p3/p57t/p1/p78/p26
/p72/p14/p15a/p24/p72x/p72/p71/p2/p92/p16/p4(11.8.5 )
Model 2
Byaretal. (1974 )developedanotherexponentialmodelrelatingsurvivaltime
toconcomitantinformationforprostatecancerpatientsinwhichtheindividualhazardislinearlyrelatedtothepossibleprognosticvariables:
/afii9838/p71/p58a/p15/p59/p78/p26
/p72/p14/p16a/p72x/p72/p71/p58/p78/p26
/p72/p14/p15a/p72x/p72/p71(11.8.6)
wherex/p15/p71/p581. Similar to the model of Feigl and Zelen (1965 ),a/p15is the
underlyinghazardratewhen covariatesareignored,forceof mortalityortheintercept.
Supposethatrofthenpatientsaredeadand s/p58n/p57rarestillaliveatthe
endofthestudy;thenthelikelihoodfunctionis
L(a/p15,a/p16,...,a/p78)/p58/p80/p147
/p71/p14/p16/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2exp/p3/p57/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2t/p71/p4
/p59/p81/p147
/p73/p14/p16exp/p3/p57/p1/p78/p26
/p72/p14/p15a/p72x/p72/p73/p2t/p62/p73/p4(11.8.7)
Takingthelogarithmsof (11.8.7 ),weobtainthelog-likelihoodfunction
l(a/p15,a/p16,...,a/p78)/p58/p80/p26
/p71/p14/p16/p3log/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2/p57/p1/p78/p26
/p72/p14/p15a/p72x/p72/p71/p2t/p71/p4/p57/p81/p26
/p73/p14/p16/p1/p78/p26
/p72/p14/p15a/p72x/p72/p73/p2t/p62/p73
(11.8.8)
ToobtaintheMLEsofthe a/p72’s,weneedtosolvesimultaneouslythefollowing
p/p591equations:
/p80/p26
/p71/p14/p16/p1x/p72/p71/p26/p78/p72/p14/p15a/p24/p72x/p72/p71/p57x/p72/p71t/p71/p2/p57/p81/p26
/p73/p14/p16x/p72/p73t/p62/p73/p580j/p580,1,...,p(11.8.9 ) 285
TheseequationscanbesolvedsimultaneouslybyusingtheNewton —Raphson
iterativeprocedure.
After the MLEs of a/p72,j/p580, 1,...,p, are obtained, the log-likelihood
functioncanbeusedtotestthesignificanceofthecovariatesbyfollowingthesameprocedureasthatusedinExample11.1.
Thesurvivorshipfunctionforthe ithindividualadjustedforthecovariates
canbeestimatedfrom
S/p19/p71(t)/p58exp(/p57/afii9838/p19/p71t)
/p58exp/p1/p57t/p78/p26
/p72/p14/p15a/p24/p72x/p72/p71/p2(11.8.10)
11.9 MODEL SELECTION METHODS
Toidentifyimportantriskfactors usinga parametricapproach,oneneedsto
select a most appropriateparametric model and identify the most significantsubset of covariates. In this section we first discuss, for a given parametricmodel,howtochooseanoptimalsubsetofthecovariatesthathavestatisticallysignificant effects on the survival time. Second, we consider if the significantcovariates are known, how to determine which parametric model is mostappropriate.Third,wediscussamethodthatcanbeusedtocompareamongparametricmodelswithdifferentsubsetsofcovariates.
11.9.1 SelectionofMost SignificantCovariatesfor a Known ParametricModel
For a known parametricmodel, the following methods can be used to select
anoptimalsubsetofthecovariatesinthesensethatthesubsetselectedhasthemost statistically significant effects on the survival time among all subsets ofthe covariates. These methods include the forward, backward, stepwise, AIC,andBICselectionprocedurescommonlyusedinordinaryregressionanalyses.Wegiveonlyabriefoutlinehere.Interestedreadersare referredtobooksonordinaryregressionanalysis.
Forward Selection Procedure
Theforwardselectionprocedureis anadding processinwhichonecovariateisselectedandaddedtothemodelateverystep.First,wehavetoestimatethespecificparametersthatdefinetheparametricmodelandthecoefficientsoftheadjusting covariates, if any, that are forced into the model. For example, tohaveage-andgender-adjustedresults,ageandgendermustbeincludedinthemodel, whether or not they are significant. Then the adjusted chi-squarestatisticsforeachcovariatenotinthemodelarecomputedandthelargestofthesestatisticsisidentified.Ifthelargestchi-squarestatisticissignificantatthe286
/afii9825level specified (usually, /afii9825/p580.15)for entry, the corresponding covariate is
addedtothemodel.
Leta/p16bethevectoroftheparametersorcoefficientsofcovariatesalreadyin
the model and l(·) be the log-likelihood function. The forward selection
procedurewillselect x/p72,whichisnotyetinthemodel,toenterthemodelifthe
difference between the log-likelihood values with x/p72and withoutx/p72is largest
among all the x/p73’s that are not in the model. That is, the coefficient a/p72ofx/p72satisfies
X/p42/p582[(l(a/p24/p72,a/p24/p16)/p57l(a/p24/p16/p72(0))]
/p58max
/p73/p432[l(a/p24/p73,a/p24/p16)/p57l(a/p24/p16/p73(0))],foranyx/p73thatisnotinthemodel /p44
(11.9.1 )
andX/p42/p57/afii9851/p17/p16/p11/p63wherea/p73isthecoefficientof x/p73notyetinthemodel,( a/p24/p73,a/p24/p16),is
theMLEof (a/p73,a/p16),a/p24/p16/p73(0)istheMLEof a/p16givena/p73/p580,and /afii9851/p17/p16/p11/p63isthe /afii9825-level
critical point of the chi-square distribution with 1 degree of freedom. In theforwardselectionprocedure,onceacovariateisenteredintothemodel,itwillnever be removed. The process is repeated until none of the remainingcovariatesmeetthelevel /afii9825specifiedforentryoruntilapredeterminednumber
ofcovariateshavebeenentered.
Backward Selection Procedure
The backward selection procedure is an elimination process in which all thecovariatesareincludedinthemodelatthebeginningandareremovedonebyone according to a significance criterion. The specific parameters that definethe parametric model and the coefficients of all the covariates are estimatedfirst.ThentheWaldtestisusedtoexamineeachcovariate.Theleastsignificantcovariate that does not meet the specified level /afii9825(usually, /afii9825/p580.15)for
stayinginthemodelisremoved.Thatis,covariate x/p72willberemovedfromthe
modelif
X/p53/p58a/p24/p17/p72v/p17/p72/p72
/p58min
/p73/p1a/p24/p17/p73v/p17/p73/p73foranyx/p73thatisinthemodel /p2(11.9.2 )
andX/p53/p45/afii9851/p17/p16/p11/p63wherea/p72is the corresponding coefficient for x/p72andv/p17/p72/p72is the
estimatedvarianceof a/p24/p72.Inthebackwardselectionprocedure,onceacovariate
isremovedfromthemodel,itremainsexcluded.Theprocessisrepeateduntilallthecovariatesremainedinthemodelmeetthespecifiedsignificancelevel /afii9825
forstayingoruntilapredeterminednumberofcovariatesremaininthemodel.The advantages of the backward procedure have been discussed by Mantel(1970 ). 287
Stepwise Selection Procedure
The stepwise selection procedure is a combinationof forward and backwardselectionprocedures. At first, it issimilar to the forwardselection procedure;however,covariatesalreadyinthemodeldonotnecessarilyremain.Covariatesalreadyinthemodelmayberemovedlateriftheyarenolongersignificant.Thestepwiseselectionprocessterminatesifnosignificantcovariatecanbeaddedtothe model or if the covariate just entered into the model is removed and nomorecovariatescanbeadded.
Information Criterion (AIC and BIC) Procedures
The Baysian information criterion (BIC)selection procedure discussed in
Section 9.3 can be used to select the best parametric model with covariates.Thiscanbedoneeasilybyreplacingthelog-likelihoodfunction l(b/p19)in(9.3.1 )
withthelog-likelihoodfunctionwithsubsetsofcovariatesdefinedinprevioussections of this chapter. The subset of covariates that produces the largest r
value in (9.3.1 )among all possible subsets is the choice. If the number of
covariates is large, one may apply the forward, backward, and stepwiseselectionmethodfirstto reducethenumberofcandidatecovariates,thenusetheBICprocedure.TheAICcriterioncanbeappliedinasimilarmanner.
11.9.2 Selection of a Parametric Model with a Fixed Subset of Covariates
Ifthemostsignificantsubsetofcovariatesisknown,selectionofanappropriate
parametricmodelcanbecarriedoutbyusingaproceduresimilartothosebasedon the likelihood functions and discussed in Section 9.2. The procedures areexactlythesameexceptthatallthelikelihoodfunctionsarereplacedbythosewith covariates, for example, those given in (11.3.10 )and Examples 11.2 and
11.5.Withcomputersoftwarepackagesavailablecommercially,theprocedurecaneasilybeapplied.Thefollowingexampleillustratestheapplication.
Example 11.6 Considerthe lungcancerpatientswhodid not receiveany
prior therapy in Example 11.4. Assume that the three covariates KPS,INDADE,andINDSMAaremostsignificant.Forthesethreefixedcovariates,the log-likelihood values based on the exponential, Weibull, lognormal, log-logistic, and generalized gamma models are given in Table 11.8. From thistable,thelognormal,Weibullandexponentialmodels (relativetothegeneral-
izedgammamodel ),withthethreecovariates,arerejectedat /afii9825/p580.0325,0.016
and 0.024, respectively. It appears that the exponentialmodel, relative to theWeibull, is not rejected (p/p580.194 ). However, since the exponential model
belongs to the Weibull distribution family and the Weibull model has beenrejected,theexponentialmodelwiththethreecovariatesisnotappropriateforthe data, as noted earlier in Chapter 9. Thus, we conclude that none of thethree models (exponential, Weibull, and lognormal ), with covariates, provide
anappropriatefittothedata.InExample11.7wewillseethatthelog-logisticmodelisthebestfitamongallthesemodels.288
Table 11.8 Goodness-of-Fit Tests Based on Asymptotic Likelihood Inference on Lung
Cancer Data
Distribution LL /p63LLR /p63p/p63BIC AIC
Extended /p57132.793 — — /p57146.517 /p57144.793
generalizedgamma
Log-logistic /p57131.230 — — /p57142.667 /p57141.230
Lognormal /p57135.022 4.459 /p640.035 /p57146.459 /p57145.022
Weibull /p57135.669 5.752 /p650.016 /p57147.106 /p57145.669
Exponential /p57136.512 7.438 /p660.024 /p57145.661 /p57144.512
Exponential /p57136.512 1.686 /p670.194 — —
/p63LL,log-likelihood;LLR,log-likelihoodratiostatistic; p,probabilitythattherespectivechi-square
randomvariable /p57LLR.
/p64Lognormalrelativetoextendedgeneralizedgamma.
/p65Weibullrelativetoextendedgeneralizedgamma.
/p66Exponentialrelativetoextendedgeneralizedgamma.
/p67ExponentialrelativetoWeibull.
Using the data file ‘‘C: /p33LCANCER.DAT’’ described in Example 11.4, the
followingSAScodecanbeusedtoobtainTable11.8.
dataw1;
infile‘c: /p33lcancer.dat’missover;
inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu;ifindpri /p580;
run;
proclifereg;
Model1:modelt*cens (0)/p58kpsindadeindsma/d /p58exponential;
Model2:modelt*cens (0)/p58kpsindadeindsma/d /p58weibull;
Model3:modelt*cens (0)/p58kpsindadeindsma/d /p58lnormal;
Model4:modelt*cens (0)/p58kpsindadeindsma/d /p58gamma;
Model5:modelt*cens (0)/p58kpsindadeindsma/d /p58llogistic;
run;
11.9.3 Selection of a Parametric Model and an Optimal Subset of Covariates
Simultaneously: AIC and BIC Procedures
TheextendedAICandBICcriteria,whichincludecovariates,canbeapplied
notonlytoselectthemostsignificantcovariatesforagivenparametricmodel,but also, simultaneously, to select the best parametric model. The proceduremaybetediousifthenumberofcovariatestobeconsideredislarge.However,in practice, the number of covariates worthy of consideration in a model isusuallyreducedafter univariateanalyses,as describedin Section11.1.There-fore,theAICorBICproceduremaynotbetoodifficulttoapply.Withtheaidof software packages, we can apply the forward, backward, and stepwise 289
selection methods in Section 11.9.1 first to fit different parametric regression
models to the data and then include in the AIC or BIC procedure all or asubsetofthesignificantcovariatesidentifiedineachfit.
Example 11.7 Consider the lung cancer data of Example 11.6. We apply
themethodsofSection11.9.1toselectthebestsubsetofcovariatesseparatelyfor the exponential, Weibull, lognormal, log-logistic, and generalized gammamodels. The same three covariates (KPS, INDADE, and INDSMA )are
selected (at the 0.05 level )as the most significantcovariatesin every of these
parametricmodels.ThelastcolumnofTable11.8givesthe rvaluesoftheBIC
for the different parametric models with the same three covariates. Based onthesevalues,thelog-logisticmodelwiththethreecovariatesshouldbeselectedasthefinalmodelforthedatasinceitsBICorAICvalueisthelargestamongallthemodels.However,itisnotknownifthelog-logisticmodelissignificantlybetterthantheothermodels.
11.9.4 Cox--Snell Residual Procedure with Covariates
TheAFTmodelsinSections11.2to11.7assumethefollowinglinearrelation-
shipbetweenlog Tandthepcovariates:
logT/p71/p58a/p15/p59/p78/p26
/p73/p14/p16a/p73x/p73/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71(11.9.3)
where /afii9830/p71has survivalfunction G(/afii9830). Oncea specifiedparametricmodeland a
subsetofcovariatesareselected,toassessthegoodnessoffitofthismodel,oneapproachistocomputetheregressionresiduals
/afii9830/p24/p71/p58logt/p71/p57/afii9839/p24/p71/afii9846/p24
i/p581,2,...,n (11.9.4)
where
/afii9839/p24/p71/p58a/p24/p15/p59/p78/p26
/p73/p14/p16a/p24/p73x/p73/p71
anda/p24/p15,a/p24/p16,a/p24/p17,...,a/p24/p78and/afii9846/p24aretheMLEof a/p15,a/p16,a/p17,...,a/p78and/afii9846,respectively,
andt/p71’s are observed survival times. An /afii9830/p24/p71is taken as censored if the
corresponding t/p71is censored. If the model fitted is correct, the corresponding
survivalfunction G(/afii9830) isthesurvivalfunctionofthefittedmodel.Forexample,
ifTindeedfollowsthe log-logisticregressionmodel with a selectedsubset of
covariates, the corresponding /afii9830/p24/p71’s should followthe log-logistic distribution.
Moreover, if the fitted model is correct, the Cox —Snell residuals defined in
(8.4.1 )are
r/p71/p58/p57logG(/afii9830/p24/p71;d/p19)/p58/p57logG/p1logt/p71/p57/afii9839/p24/p71/afii9846/p24;d/p19/p2i/p581,2,...,n(11.9.5 )290
Figure 11.2 Cox—Snellresidualsplotfromthefittedexponentialmodelonlungcancer
data.
whered/p19istheMLEoftheparametersofthedistribution.Let S/p19(r) denotethe
estimated survival function of r/p71’s. From Section 8.4, the graph of r/p71versus
/p57logS/p19(r/p71),i/p581, 2,...,n, should be closed to a straight line with unit slope
and zero intercept if the fitted model for the survival time Tis correct. This
graphicalmethodcan be used to assessthe goodness of fit of the parametric
regressionmodel.
Example 11.8 Figures 11.2 to 11.6 showthe Cox —Snell residuals plots
from fitting the exponential, Weibull, lognormal, log-logistic, and extendedgeneralized gamma models, respectively with the three covariates KPS, IN-DADE, and INDSMA, to the lung cancer data in Example 11.6. The fivegraphslooksimilar,andallareclosetoastraightlinewithunitslopeandzerointercept. No significant differences are observed in these graphs. The resultsobtained are similar to those from Examples 11.6 and 11.7. The differencesamongthe fivedistributionsaresmall withthelog-logisticdistributionbeingslightlybetterthantheothers.
Using the same data file ‘‘C: /p33LCANCER.DAT’’ as in Example 11.6.1, the
following SAS code can be used to obtain the Cox —Snell residuals based on
the exponential, Weibull, lognormal, log-logistic, and generalized gammamodelwiththethreecovariates,KPS,INDADE,andINDSMA. 291
Figure 11.3 Cox—Snell residuals plot from the fitted Weibull model on lung cancer
data.
Figure 11.4 Cox—Snellresidualsplotfromthefittedlognormalmodelonlungcancer
data.292
Figure 11.5 Cox—Snellresidualsplotfromthefittedlog-logisticmodelonlungcancer
data.
dataw1;
infile‘c: /p33lcancer.dat’missover;
inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu;ifindpri /p580;
run;
procliferegnoprint;
a:modelt*cens (0)/p58kpsindadeindsma/d /p58exponential;
outputout /p58wacdf /p58f;
b:modelt*cens (0)/p58kpsindadeindsma/d /p58weibull;
outputout /p58wbcdf /p58f;
c:modelt*cens (0)/p58kpsindadeindsma/d /p58lnormal;
outputout /p58wccdf /p58f;
d:modelt*cens (0)/p58kpsindadeindsma/d /p58gamma;
outputout /p58wdcdf /p58f;
e:modelt*cens (0)/p58kpsindadeindsma/d /p58llogistic;
outputout /p58wecdf /p58f;
run;
datawa;
setwa;model /p58‘Exponential’;
datawb; 293
Figure 11.6 Cox—Snell residuals plot from the fitted extended generalized gamma
modelonlungcancerdata.
setwb;
model /p58‘Weibull’;
datawc;
setwc;model /p58‘LNnormal’;
datawd;
setwd;model /p58‘Gamma’;
datawe;
setwe;model /p58‘LLogistic’;
dataw2;
setwawbwcwdwe;rcs/p58/p57log(1/p57f);
run;procsort;
bymodel;
run;
proclifetestnotableouts /p58wsnoprint;
timercs*cens (0);
bymodel;294
run;
dataws;
setws;
mls/p58/p57log(survival );
run;title‘Cox-SnellResiduals (rcs)and-log (estimatedsurvivalfunctionofrcs )(mls)’;
procprintdata /p58ws;
varmodelrcsmls;
run;
Bibliographical Remarks
Anexcellentexpositorypaperonstatisticalmethodsfortheidentificationand
use of prognostic factors has been written by Armitage and Gehan (1974 ).
Manystudiesofprognosticfactorshavebeenpublished,includingSirottetal.(1993 ), Brancato et al. (1997 ), Linka et al. (1998 ), and Lassarre (2001 ). The
accelerated failure time (AFT )model was introduced by Cox (1972 ). The
detailedstatisticalinferenceoftheAFTmodelsandthetheoreticalaspectsofmodel-selectingmethodsareincludedintheworkscitedinthebibliographicalremarks at the end of Chapter 9 and in the papers and books cited in thischapter.
EXERCISES
11.1Consider the data given in Exercise Table 3.1. In addition to the five
skintests,ageandgendermayalsohaveprognosticvalue.Examinetherelationshipbetweensurvivalandeachofthesevenpossibleprognosticvariables as in Table 3.12. For each variable, group the patientsaccording to different cutoff points. Estimate and drawthe survivalfunctionforeachsubgroupbytheproduct-limitmethodandthenusethemethodsdiscussedinChapter5 tocomparesurvivaldistributionsofthesubgroups.PrepareatablesimilartoTable3.12.Interpretyourresults. Is there a subgroup of any variable that shows significantlylongersurvivaltimes? (Fortheskintestresults,usethelargerdiameter
ofthetwo. )
11.2Consider the seven variables in Exercise 11.1. Use the Weibull re-
gressionmodeltoidentifythemostsignificantvariables.CompareyourresultswiththatobtainedinExercise11.1.
11.3ConsiderthedatagiveninExerciseTable3.3.Examinetherelationship
between remission duration and survival time and each of the ninepossibleprognosticvariables:age,gender,familyhistoryofmelanoma,andthesixskintests.Groupthepatientsaccordingtodifferentcutoff 295
points. Estimate and drawremission and survival curves for each
subgroup. Compare the remission and survival distributions of sub-groups using the methods discussed in Chapter 5. Prepare tablessimilartoTable3.8.
11.4Perform the following analyses: (1)Use the exponential, Weibull,
lognormal,generalizedgamma,orlog-logisticregressionmodelssepa-ratelytoidentifythesignificantvariablesinExercise11.3fortheirrelative
importancetoremissiondurationandsurvivaltime. (2)Selectamodel
amongthesefinalmodelsusingtheBICorAICmethod. (3)Calculate
separatelytherespectivelikelihoodfortheexponential,Weibull,lognor-mal,orgeneralizedgammaregressionmodelwiththefixedvariablesinthemodelselectedinstep (2),thenusethemethodinSection11.9.2tochoose
amodelandseewhetherthismodelisthemodelselectedinstep (2).
11.5Performthesameanalysesas inExercise11.4for survivaltimein the
157diabeticpatientsgiveninExerciseTable3.4.
11.6Using the notations in Example 11.2, show that if we use the follow-
ing model to replace the model defined in (11.4.6 ), logT/p71/p58a/p15/p59a/p16LOW/p9/p59a/p17UNSA/p9/p59/afii9846/afii9830/p9/p58/afii9839/p9/p59/afii9846/afii9830/p9, the hypothesis H/p15:h/p81/p58h/p83is
equivalenttoH/p15:a/p17/p580.
11.7Following Examples 11.2 and 11.3, obtain the log-likelihood function
basedon (11.6.8 )forthe137observedexactandright-censoredsurvival
timesfromthelungcancerpatients.
11.8Using the same notation as in Example 11.5, showthat if w e
use the model log T/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDLAR/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71/p59/afii9846/afii9830/p71,whereINDLAR /p581ifthetypeofcancerislarge,
and0otherwise,toreplacethemodeldefinedin (11.7.8 ),thehypotheses
H/p15:a/p18/p580 andH/p15:a/p19/p580 are equivalent to H/p15:OR/p49/p43/p58OR/p31/p34and
H/p15:OR/p49/p47/p58OR/p31/p34, respectively. Moreover, if we use the model
logT/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDLAR/p71/p59a/p18INDADE/p71/p59a/p19INDSQU/p71/p59/afii9846/afii9830/p71toreplacethemodeldefinedin (11.7.8 ),thehypothesis H/p15:a/p19/p580
isequivalentto H/p15:OR/p49/p47/p58OR/p49/p43.
11.9Let/afii9830beasurvivaltimewiththedensityfunction g(/afii9830)/p58exp[/afii9830/p57exp(/afii9830)].
Showthat the survival time Tdefined by log T/p58/afii9839/p59/afii9846/afii9830has the
Weibulldistributionwith /afii9838/p58exp(/p57/afii9839//afii9846)and/afii9828/p581//afii9846byapplyingthe
densitytransformationrulein (11.2.11 ).
11.10Let/afii9830beasurvivaltimewiththestandardnormaldistribution N(0,1).
Showthat the survival time Tdefined by log T/p58/afii9839/p59/afii9846/afii9830has the
lognormaldistributionbyapplyingthe densitytransformationrule in(11.2.11 ).296
11.11Letubeasurvivaltimewiththedensityfunction f(u),
f(u)/p58exp[u//afii9829/p17/p57exp(u)]
/afii9772(1//afii9829/p17)
where /afii9772(·)isthegammafunctiondefinedin (6.2.8 ).
(a)Showthat the survival time /afii9830defined by /afii9830/p58/afii9839//afii9829/p59(log/afii9829/p17)//afii9829has
thefollowingdensityfunction,
g(/afii9830)/p58/p34/afii9829/p34[exp (/afii9829/afii9830)//afii9829/p17]/p16/p30/p66/p130exp[/p57exp(/afii9829/afii9830)//afii9829/p17]
/afii9772(1//afii9829/p17)/p57/p45/p58/afii9830/p58 /p59/p45
andsurvivalfunction,
G(/afii9830)/p58/p7I/p1exp(/afii9829/afii9830)
/afii9829/p17,1
/afii9829/p17/p2if/afii9829/p580
1/p57I/p1exp(/afii9829/afii9830)
/afii9829/p17,1
/afii9829/p17/p2if/afii9829/p570
whereI(·,·)istheincompletegammafunctiondefinedasin (6.4.4 ).
(b)Showthat the survival time Tdefined by log T/p58/afii9839/p59/afii9829/afii9830has the
extendedgammadensityfunctiondefinedin (11.6.4 ).
11.12If/afii9830hasalogisticdistributionwiththedensityfunction
g(/afii9830)/p58exp(/afii9830)
[1/p59exp(/afii9830)]/p17
showthatthesurvivaltime TdefinedbylogT/p58/afii9839/p59/afii9846/afii9830hasthelog-logistic
distribution with /afii9825/p58exp(/p57/afii9839//afii9846)and/afii9828/p581//afii9846by applying the density
transformationrulein (11.2.11 ). 297
CHAPTER 12
Identification of Prognostic Factors
Related to Survival Time:Cox Proportional Hazards Model
In Chapter 11 we discussed parametric survival methods for model fitting andforidentifyingsignificant prognosticfactors. Thesemethodsare powerfulif theunderlying survival distribution is known. The estimation and hypothesistesting of parameters in the models can be conducted by applying standardasymptotic likelihood techniques. However, in practice, the exact form of theunderlying survival distribution is usually unknown and we may not be ableto find an appropriate model. Therefore, the use of parametric methods inidentifying significant prognostic factors is somewhat limited. In this chapterwediscussamostcommonlyused model,theCox (1972 )proportionalhazards
model, and its related statistical inference. This model does not requireknowledge of the underlying distribution. The hazard function in this modelcan take on any form, including that of a stepfunction, but the hazardfunctions of different individuals are assumed to be proportional and indepen-dentof time.Theusual likelihoodfunctionis replacedby thepartial likelihoodfunction.Theimportantfactisthatthestatisticalinferencebasedonthepartiallikelihood function is similar to that based on the likelihood function.
12.1 PARTIALLIKELIHOODFUNCTIONFORSURVIVALTIMES
The Cox proportional hazards model possesses the property that different
individuals have hazard functions that are proportional, i.e., [ h(t/p34x/p16)/h(t/p34x/p17)],
the ratio of the hazardfunctions for two individuals with prognostic factors orcovariatesx/p16/p58(x/p16/p16,x/p17/p16,...,x/p78/p16)/p30, andx/p17/p58(x/p16/p17,x/p17/p17,...,x/p78/p17)/p30is a constant
(does not vary with time t). This means that the ratio of the risk of dying of
two individuals is the same no matter how long they survive. In Sections 11.3
298
and 11.4, we showed that the exponential and Weibull regression models
possess this property. This property implies that the hazard function given aset of covariates x/p58(x/p16,x/p17,...,x/p78)/p30can be written as a function of an
underlying hazard function and a function, say g(x/p16,...,x/p78), of only the
covariates, that is,
h(t/p34x/p16,...,x/p78)/p58h/p15(t)g(x/p16,...,x/p78)o rh(t/p34x)/p58h/p15(t)g(x)(12.1.1 )
The underlying hazard function, h/p15(t), represents how the risk changes with
time,andg(x)representsthe effect of covariates. h/p15(t) can be interpreted as the
hazard function when all covariates are ignored or when g(x)/p581, and is also
called thebaseline hazard function . The hazard ratio of two individuals with
different covariates x/p16andx/p17is
h(t/p34x/p16)
h(t/p34x/p17)/p58h/p15(t)g(x/p16)
h/p15(t)g(x/p17)/p58g(x/p16)
g(x/p17)(12.1.2)
which is a constant, independent of time.
The Cox (1972 )proportional hazard model assumes that g(x)in(12.1.1 )is
an exponential function of the covariates, that is,
g(x)/p58exp/p1/p78/p26
/p72/p14/p16b/p72x/p72/p2/p58exp(b/p30x)
and the hazard function is
h(t/p34x)/p58h/p15(t) exp /p1/p78/p26
/p72/p14/p16b/p72x/p72/p2/p58h/p15(t) exp (b/p30x)( 12.1.3 )
whereb/p58(b/p16,...,b/p78)denotes the coefficients of covariates. These coefficients
can be estimated from the data observed and indicate the magnitude of theeffects of their corresponding covariates. For example, if there is only onecovariate treatment, let x/p16/p580 if a person receives placebo and x/p16/p581i fa
personreceivestheexperimentaldrug.Thehazardratioofthepatientreceivingthe experimental drug and the one receiving placebo based on (12.1.2 )and
(12.1.3 )is
h(t/p34x/p16/p581)
h(t/p34x/p16/p580)
/p58exp(b/p16)
Thus, the two treatments are equally effective if b/p16/p580 and the experimental
drugintroduceslower (higher )riskforsurvivalthanplaceboif b/p16/p580(b/p16/p570).
It can be shown that (12.1.3 )is equivalent to
S(t/p34x)/p58[S/p15(t)]exp(/afii9814/p78/p72/p14/p16b/p72x/p72)/p58[S/p15(t)]exp(b/p30x)(12.1.4 ) 299
Thus the covariates can be incorporated into the survivorship function. The
use of (12.1.3 )can be exemplified as follows.
1.Two-sample problems. Suppose that p/p581; that is, there is only one
covariate,x/p16, which is an indicator variable:
x/p16/p71/p58/p70 if theith individual is from group0
1 if theith individual is from group1
Then according to (12.1.3 ), the hazard functions of groups 0 and 1 are,
respectively, h/p15(t) andh/p16(t)/p58h/p15(t) exp(b/p16). Thehazardfunctionofgroup
1 is equal to the hazard function of group0 multip lied by a constantexp(b/p16), or the two hazard functions are proportional. In terms of the
survivorshipfunction,
S(t)/p58[S/p15(t)]/p65
wheretheconstant c/p58exp(b/p16)(Nadas,1970 ).Thetwo-sampletestdevelop-
ed from (12.1.3 )is the Cox—Manteltest discussedin Chapter 5. It is now
apparent that the test is based on the assumption of a proportionalhazard between the two groups.
2.Two-sampleproblemswithcovariates. The covariates in (12.1.3 )can either
be indicator variables such as x/p16in the two-groupp roblem above or
prognosticfactors. Havingone or more covariates representing prognos-tic factors in (12.1.1 )enables us to examine the relation between two
groups, adjusting for the presence of prognostic factors.
3.Regression problems. Dividing both sides of (12.1.3 )byh/p15(t) and taking
its logarithm, we obtain
logh/p71(t)
h/p15(t)
/p58b/p16x/p16/p71/p59b/p17x/p17/p71/p59/p37/p59b/p78x/p78/p71/p58/p78/p26
/p72/p14/p16b/p72x/p72/p71/p58b/p30x/p71(12.1.5 )
wherethex’sarecovariatesforthe ithindividual.Theleftsideof (12.1.5 )is
a function of hazard ratio (or relative risk )and the right side is a linear
function of the covariates and their respective coefficients.
As mentioned earlier, h/p15(t) is the hazard function when all covariates are
ignored.If thecovariates are standardizedaboutthe meanand the modelusedis
logh/p71(t)
h/p15(t)/p58b/p16(x/p16/p71/p57x/p21/p16)/p59b/p17(x/p17/p71/p57x/p21/p17)/p59/p37/p59b/p78(x/p78/p71/p57x/p21/p78)/p58b/p30(x/p71/p57x/p21)
(12.1.6 )300
wherex/p21/p30/p58(x/p21/p16,x/p21/p17,...,x/p21/p78)andx/p21/p72is the average of the jth covariate for all
patients, the left side of (12.1.6 )is the logarithm of the ratio of risk of failure
for a patient with a given set of values x/p30/p71/p58(x/p16/p71,x/p17/p71,...,x/p78/p71)to that for an
average patient who has an average value for every covariate.
In this chapter we focus on the use of (12.1.5 ), and the main interest here is
to identify important prognostic factors. In other words, we wish to identifyfrom thepcovariates a subset of variables that affect the hazard more
significantly, and consequently, the length of survival of the patient. We are
concerned with the regression coefficients. If b/p71is zero, the corresponding
covariateis notrelatedto survival.If b/p71is notzero,it representsthe magnitude
of the effect of x/p71on hazard when the other covariates are considered
simultaneously.
To estimate the coefficients, b/p16,...,b/p78, Cox (1972 )proposes a partial
likelihoodfunctionbasedon aconditionalprobabilityoffailure,assumingthattherearenotiedvaluesinthesurvivaltimes.However,inpractice,tiedsurvivaltimes are commonly observed and Cox’s partial likelihood function wasmodified to handle ties (Kalbfleisch and Prentice, 1980; Breslow, 1974; Efron,
1977 ). In the following we describe the estimation procedure without and with
ties.
12.1.1 EstimationProcedureswithoutTiedSurvivalTimes
Suppose that kof the survival times from nindividuals are uncensored and
distinct, and n/p57kare right-censored. Let t/p7/p16/p8/p58t/p7/p17/p8/p58/p37/p58t/p7/p73/p8be the ordered
kdistinct failure times with corresponding covariates x/p7/p16/p8,x/p7/p17/p8,...,x/p7/p73/p8. Let
R(t/p7/p71/p8) be the risk set at time t/p7/p71/p8.R(t/p7/p71/p8)consists of all persons whose survival
times are at least t/p7/p71/p8. For the particular failure at time t/p7/p71/p8, conditionallyon the
risksetR(t/p7/p71/p8),theprobabilitythatthefailureisontheindividualasobservedis
exp/p37
/p26/p78/p72/p14/p16b/p72x/p72/p7/p71/p8/p38
/p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74) /p1/p58exp(b/p30x/p7/p71/p8)
/p26l/p43R(t/p7/p71/p8)exp(b/p30x/p74)/p2
Each failure contributes a factor and hence the partial likelihood function is
L(b)/p58/p73/p147
/p71/p14/p16exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p7/p71/p8/p38
/p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74) /p1/p58/p73/p147
/p71/p14/p16exp(b/p30x/p7/p71/p8)
/p26l/p43R(t/p7/p71/p8)exp(b/p30x/p74)/p2(12.1.7 )
and the log-partial likelihood is
l(b)/p58logL(b)/p58/p73/p26
/p71/p14/p16/p78/p26
/p72/p14/p16b/p72x/p72/p71/p57/p73/p26
/p71/p14/p16log/p3/p26
l/p43R(t/p7/p71/p8)exp/p1/p78/p26
/p72/p14/p16b/p72x/p72/p74/p2/p4
/p58/p73/p26
/p71/p14/p16/p7b/p30x/p7/p71/p8/p57log/p3/p26
l/p43R(t/p7/p71/p8)exp(b/p30x/p74)/p4/p8(12.1.8 ) 301
The maximum partial likelihood estimator (MPLE )b/p19ofbcan be obtained by
the steps shown in (7.1.2 )—(7.1.4 ). That is,b/p19/p16,...,b/p19/p78are obtained by solving
the following simultaneous equations:
/p42(l(b))
/p42b/p580
or
/p42l(b)
/p42b/p83/p58/p73/p26
/p71/p14/p16[x/p83/p7/p71/p8/p57A/p83/p71(b)]/p580u/p581, 2,...,p(12.1.9)
where
A/p83/p71(b)/p58/p26l/p43R(t/p7/p71/p8)x/p83/p74exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74)
/p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74/p38/p58/p26l/p43R(t/p7/p71/p8)x/p83/p74exp(b/p30x/p74)
/p26l/p43R(t/p7/p71/p8)exp(b/p30x/p74)(12.1.10 )
by applying the Newton —Raphson iterated procedure. The second partial
derivatives of l(b)with respective to b/p83andb/p84,u,v/p581, 2,...,p, in the
Newton—Raphson iterative procedure are
I/p83/p84(b)/p58/p42/p17l(b)
/p42b/p83/p42b/p84/p58/p57/p73/p26
/p71/p14/p16C/p7/p83/p84/p71/p8(b/p16,...,b/p78)/p58/p57/p73/p26
/p71/p14/p16C/p7/p83/p84/p71/p8(b)u,v/p581, 2,...,p
(12.1.11)
where
C/p7/p83/p84/p71/p8(b)/p58/p26l/p43R(t/p7/p71/p8)x/p83/p74x/p84/p74exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74)
/p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74)/p57A/p83/p71(b)A/p84/p71(b) (12.1.12)
The covariance matrix of the MPLE b/p19, defined similarly as V/p19(b) defined in
(7.1.5 ),i s
V/p19(b/p19)/p58Cov/p19(b/p19)/p58/p3/p57/p42/p17l(b/p19)
/p42b/p42b/p30/p4/p92/p16(12.1.13 )
where /p57/p42/p17l(b/p19)//p42b/p42b/p30is called the observedinformationmatrix with /p57I/p83/p84(b/p19)as
its(u,v)element and where I/p83/p84(b)is defined in (12.1.11 ). Let the (i,j) element
ofV/p19(b/p19)in(12.1.13 )bev/p71/p72; then the 100 (1/p57/afii9825)% confidence interval for b/p71is,
according to (7.1.6 ),
/p37b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71,b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71/p38 (12.1.14)
12.1.2 EstimationProcedurewithTiedSurvivalTimes
Suppose that among the nobserved survival times there are kdistinct
uncensored times t/p7/p16/p8/p58t/p7/p17/p8/p58/p37/p58t/p7/p73/p8. Letm/p7/p71/p8denote the number of people302
who fail att/p7/p71/p8or the multiplicity of t/p7/p71/p8;m/p7/p71/p8/p571 if there are more than one
observation with value t/p7/p71/p8;m/p7/p71/p8/p581 if there is only one observation with value
t/p7/p71/p8. LetR(t/p7/p71/p8)denote the set of people at risk at time t/p7/p71/p8[i.e.,R(t/p7/p71/p8)consists of
those whose survival times are at least t/p7/p71/p8] andr/p71be the number of such
persons.Forexample,in thefollowingset ofsurvivaltimes fromeightsubjects,/p4315, 16 /p59, 20, 20, 20, 21, 24, 24 /p44,n/p588,k/p584,t/p7/p16/p8/p5815,t/p7/p17/p8/p5820,t/p7/p18/p8/p5821,
t/p7/p19/p8/p5824,m/p7/p16/p8/p581,m/p7/p17/p8/p583,m/p7/p18/p8/p581, andm/p7/p19/p8/p582. ThenR(t/p7/p16/p8) includes all
eight subjects. R(t/p7/p17/p8)/p58/p43those subjects with survival times 20, 21, and 24 /p44,
R(t/p7/p18/p8)/p58/p43those subjects with survival times 21 and 24 /p44, andR(t/p7/p19/p8)/p58/p43those
subjects with survival time 24 /p44; thus,r/p16/p588,r/p17/p586,r/p18/p583, andr/p19/p582.
To discuss the methods for ties, we introduce a few additional notations.
From everyR(t/p7/p71/p8), we can randomly select m/p7/p71/p8subjects. Donate each of these
m/p7/p71/p8selections byu/p7/p72/p8. There are/p80/p71C/p75
/p7/p71/p8/p58r/p71!/[m/p7/p71/p8!(r/p71/p57m/p7/p71/p8)!] possibleu/p7/p72/p8’s. Let
U/p71denote the set that contains all the u/p7/p72/p8’s. For example, from R(t/p7/p17/p8), we can
randomlyselectany m/p7/p17/p8/p583 out of the 6 (r/p17/p586) subjects. There are a total of
/p21C/p18/p5820 such selections (or subsets ), and one ofu/p7/p72/p8’s is, for example, /p43three
subjects with survival times 20, 20, and 24 /p44.U/p17/p58/p43u/p7/p16/p8,u/p7/p17/p8,...,u/p7/p17/p15/p8/p44contains
all 20 subsets. Now let us focus on the tied observations. Letx/p73/p58(x/p16/p73,x/p17/p73,...,x/p78/p73)/p30denote the covariates of the kth individual,
z
u(j)/p58/p26k/p43u/p7/p72/p8x/p73/p58(z1u(j),z2u(j),...,zpu(j))/p30,wherezlu(j)isthesumof the lth covariate
of them/p7/p71/p8persons who are in u/p7/p72/p8. Letu*/p7/p71/p8denotes the set of m/p7/p71/p8people who
failed at time t/p7/p71/p8, andzu*(i)/p58/p26k/p43u*/p7/p71/p8x/p73/p58(z*1u*(i),z*2u*(i),...,z*pu*(i))/p30, wherez*lu*(i)be
the sum of the lth covariate of the m/p7/p71/p8persons who are in u*/p7/p71/p8(failed at time
t/p7/p71/p8). For example, for the set of survival times above, z*1u*(2)equals the sum of
the first covariate values of three persons who failed at time 20. With thesenotations we are ready to introduce the following method for ties.
Continuous Time Scale
In the case of a continuous time scale, for the m/p7/p71/p8persons failing at t/p7/p71/p8,i ti s
reasonable to say that the survival times of the m/p7/p71/p8people are not identical
sincethe ties are most likelyto be the resultsof imprecisemeasurements.If theprecisemeasurementscouldbemade,these m/p7/p71/p8survivaltimescouldbeordered
and we could use the likelihood function in (12.1.7 ). In the absence of
knowledge of the true order (the real case ), we have to consider all possible
orders of these observed m/p7/p71/p8tied survival times. For each t/p7/p71/p8, the observed m/p7/p71/p8tied survival time can be ordered in m/p7/p71/p8!(m/p7/p71/p8factorial )different possible ways.
For each of these possible orders we will have a product as in (12.1.7 )for the
corresponding m/p7/p71/p8survival times. Therefore, when the survival time is meas-
ured at a continuous time scale, construction and computation of the exactpartial likelihood function is a very tedious task if m/p7/p71/p8is larger. Readers
interested in the details of the exact partial likelihood function are referred toKalbfleischandPrentice (1980 )andDelongetal. (1994 ).Theformulaprovided
by Delong et al. makes computation of the partial likelihood function for tiedcontinuous survival times more feasible. We will not discuss the exact partiallikelihood function further due to its complexity. Among the statistical sof- 303
twarepackages,SASincludesaprocedurebasedontheexactpartiallikelihood
function. Use of this procedure is illustrated in Example 12.3.
To approximate the exact partial likelihood function, the following two
likelihood functions can be used when each m/p7/p71/p8is small compared to r/p71.
Breslow (1974 )provided the following approximation:
L/p32(b)/p58/p73/p147
/p71/p14/p16exp(z/p30u*(i)b)
[/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b)]m/p7/p71/p8(12.1.15 )
An alternative approximation was provided by Efron (1977 ):
L/p35(b)/p58/p73/p147
/p71/p14/p16exp(z/p30u*(i)b)
/p147/p75/p7/p71/p8/p72/p14/p16[/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b)/p57[(j/p571)/m/p7/p71/p8]/p26l/p43u*/p7/p71/p8exp(x/p30/p74b)](12.1.16 )
Discrete Time Scale
If survival times are observed at discrete times, the tied observations are trueties: that is, these events really happen at the same time. Cox (1972 )proposed
the following logistic model:
h/p71(t)dt
1/p57h/p71(t)dt/p58h/p15(t)dt
1/p57h/p15(t)dtexp/p1/p78/p26
/p72/p14/p16b/p72x/p72/p71/p2/p58h/p15(t)dt
1/p57h/p15(t)dtexp(b/p30x/p71)
This model reduces to (12.1.3 )in the continuous time scale. Using the model
and replacing the ith term in (12.1.7 )with the following term with tied
observations at t/p7/p71/p8:
exp(z/p30u*(i)b)
/p26u/p7/p72/p8/p43U/p71exp(z/p30u(j)b)
the partial likelihoodfunction with tied observations at a discrete time scale is
L/p66(d)/p58/p73/p147
/p71/p14/p16exp(z/p30u*(i)b)
/p26u/p7/p72/p8/p43U/p71exp(z/p30/p83/p7/p72/p8b)(12.1.17 )
Theith term in this expression represents the conditional probability of
observing the m/p7/p71/p8failures given that there are m/p7/p71/p8failures at time t/p7/p71/p8and the
risk setR(t/p7/p71/p8)att/p7/p71/p8. The number of terms in the denominator of the ith term
is/p80/p71C/p75/p7/p71/p8/p58r/p71!/[m/p7/p71/p8!(r/p71/p57m/p71)!], as noted earlier, and will be very large if the m/p7/p71/p8is large. Fortunately, a recursive algorithm proposed by Gail et al. (1981 )
makes the calculation manageable. Equation (12.1.17 )can also be considered
as an approximation of the partial likelihood function for continuous survivaltimes with ties by assuming that the ties are true as if they were observed at adiscrete time scale.304
Asshowninmanypapersinliterature,inmostpracticalsituations,thethree
partial likelihood functions above are reasonably good approximations of theexact partial likelihood function for continuous survival time with ties. Whenthere are no ties on the event times (i.e.,m/p7/p71/p8/p891), (12.1.15)—(12.1.17 )reduce to
(12.1.7 ). The maximum partial likelihood estimate of bin(12.1.15 )—(12.1.17 )
can be estimated using procedures similar to those in (12.1.8 )—(12.1.14 ).
Once the coefficients are estimated, relative risk (or relative hazard )in
(12.1.2 )or(12.1.5 )can be obtained.For example, if x/p16represents hypertension
and is defined as
x/p16/p58
/p71 if patient is hypertensive
0 otherwise
thehazardrateforhypertensivepatientsis exp (b/p19/p16)timesthat fornormotensive
patients. That is, the risk associated with hypertension is exp (b/p19/p16)adjusting for
the other covariates in the model. A 100 (1/p57/afii9825)% confidence interval for the
relative risk can be obtained by using the confidence interval for b/p16. Let
(b/p16/p42,b/p16/p51)be the 100 (1/p57/afii9825)% confidence interval for b/p16; a 100 (1/p57/afii9825)%
confidence interval for the relative risk is (exp(b/p16/p42),exp(b/p16/p51))according to
(7.1.8 ). This application of the proportional hazards model has been used
extensively, particularly by epidemiologists.
The following three examples illustrate the use of Cox’s regression model.
Example 12.1 Consider the survival data from 30 patients with AML in
Table 11.4. Recall that the two possible prognostic factors are
x/p16/p58/p71 if patient is /p4650 years
0 otherwise
x/p17/p58/p71 if cellularity of marrow clot section is 100%
0 otherwise
We fit the Cox proportional hazard model to the data. The results are
presented in Table 12.1 In this case, Breslow’s approximation in (12.1.15 )is
usedtohandleties.Thepositivesignsoftheregressioncoefficientsindicatethatthe older patients (/p4650 years )and patients with 100% cellularity of the
marrow clot section have a higher risk of dying. Furthermore, age is signifi-cantly related to survival after adjustment for cellularity. The results areconsistent with those from fitting the lognormal regression model in Example11.3. The coefficients of the binary covariates can be interpreted in terms ofrelative risk. The estimated risk of dying for patients at least 50 years of age is2.75 times higher than that for patients younger than 50. Patients with 100%cellularity have a 42%higher risk of dying than patients with less than 100%cellularity. 305
Table12.1 ResultsofaProportionalHazardsRegressionAnalysisofDatain
Table11.4
Regression Standard
Covariate Coefficient Error pValue exp (coefficient )
x/p16(age) 1.01 0.46 0013 2.75
x/p17(cellularity ) 0.35 0.44 0.212 1.42
The95%confidenceintervalsfor b/p16(age)andb/p17(cellularity )are1.01 /p601.96
(0.46)or(0.11, 1.91 )and 0.35 /p601.96 (0.44)or(/p570.51, 1.21 ), respectively.
Consequently, the 95% confidence intervals for the relative risks are(e/p15/p13/p16/p16,e/p16/p13/p24/p16)or(1.12, 6.75 )and (e/p92/p15/p13/p20/p16,e/p16/p13/p17/p16)or(0.60, 3.35 ), respectively. The
small number of patients (30)may have contributed to the large standard
errors ofb/p19/p16andb/p19/p17and consequently, the wide confidence intervals. The lower
bound of the confidence interval for age is only slightly above 1. This suggeststhat the importance of age should be interpreted carefully. In general, if thenumber of subjects is small and the standard errors of the estimates are large,the estimates may be unreliable.
When the two covariates are considered simultaneously, the risk for a
patientwith x/p16/p581 andx/p17/p581 relative to patientswith x/p16/p580 andx/p17/p580 can
be estimated. The relative risk is estimated as exp (1.01/p590.35)/p583.90 for a
patient who is over 50 years of age and whose cellularity is 100%, comparedto patients who are younger than 50 and whose cellularity is less than 100%.
Using the same data set ‘‘C: /p33AML.DAT’’ defined in Example 11.3, the
following SAS code can be used to obtain the results in Table 12.1.
data w1;
infile ‘c: /p33aml.dat’ missover;
input t cens x1 x2;
run;proc phreg;
model t*cens (0)/p58x1 x2 / rl;
run;
If BMDP 2L is used, the following code is applicable.
/input file /p58‘c:/p33aml.dat’ .
variables /p584.
format /p58free.
/print cova./variable names /p58t, cens, x1, x2.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58x1, x2.306
If SPSS is used, the following code suffices.
data list file /p58‘c:/p33aml.dat’ free
/ t cens x1 x2.
coxreg t with x1 x2
/status /p58cens event (1)
/print /p58all.
Example 12.2 In a study (Buzdar et al., 1978 )to evaluate a combination
of 5-flourouracil, adramycin, cyclophosphamide, and BCG (FAC-BCG )as
adjuvant treatment in stage II and III breast cancer patients with positiveaxillary nodes, 131 patients receiving FAC-BCG after surgery and radiationtherapy were compared with 151 patients receiving surgery and radiationtherapy only (control group ).
Cox’s regression model was used to identify prognostic factors and to
evaluate the comparability of the two treatment groups. The model was fittedto the data from 151 patients to determine the variables related to length ofremission. The possible prognostic variables considered were age (years ),
menopausal status (1, premenopausal; 0, other ), size of primary tumor (2,/p583
cm; 4, 3—5 cm; 7, /p575c m ), state of disease (2, stage II; 3, stage III ), location of
surgery (1, M. D. Anderson Hospital; 0, other ), number of nodes involved (2,
/p584; 7, 4—10; 12, /p5710), and race (1, Caucasian; 2, other ). The covariates were
selected by the forward selection method outlined in Section 11.9. Threevariables—number of nodes involved, state of disease, and menopausalstatus—were selected for use in the model,all related significantly (p/p580.1)to
disease-free time. The regression equation including these variables only is
logh/p71(t)
h/p15(t)
/p580.111(number of nodes /p576.16)/p590.8122 (stage /p572.39)
/p590.872 (menopausal /p570.26)
Table 12.2 gives the details of the fit. Relative risk was taken as h/p71(t)/h/p15(t), the
ratio of the risk of death per unit of time for a patient with a given set ofprognostic variables to the risk for a patient whose prognostic variables wereaverage in value. The relative risk for each variable was calculated byconsidering favorable or unfavorable values of that variable, assuming thatother variables were at their average value. Note that the risk of relapse perunit time for a patient with 12 positive nodes is 3.04 (ratio or risk )times that
fora patientwith only twopositive nodes.The riskof relapseper unittime fora stage III patient was 2.25 times that of a stage II patient.
The Cox’s regression model was also fitted to the combined groupof
FAC-BCG and control patients, including type of treatment (0, control; 1,
FAC-BCG ),menopausalstatus,sizeofprimarytumor,andnumberofinvolved
nodes as potential prognostic variables. The regression equation with three 307
Table12.2 PatientCharacteristicsRelatedtoDisease-FreeTimeinCox’sRegression
ModelFittoControlPatients
Maximum Relative Risk /p63
Prognostic Regression Significance Log Ratio of
Variable Coefficient Level ( p) Likelihood Favorable Unfavorable Risks
Number of
nodes 0.1110 /p580.01 /p57257.407 0.63 1.91 3.04
Stage 0.8122 0.016 /p57254.533 0.73 1.64 2.25
Menopausal
status 0.8720 /p580.1 /p57250.576 0.80 1.91 2.39
Source:Buzdar et al. (1978 ). Reprinted by permission of the editor.
/p63Favorable variables: number of nodes /p582, stage II, postmenopausal. Unfavorable variables:
number of nodes /p5812, stage III, premenopausal.
Table12.3 PatientCharacteristicsRelatedtoSurvival,TreatmentIncluded
Maximum Relative Risks /p63
Prognostic Regression Significance Log Ratio of
Variable Coefficient Level ( p) Likelihood Favorable Unfavorable Risks
Treatment /p571.8792 /p580.01 /p57201.200 0.37 2.42 6.55
Menopausal 0.9644 0.01 /p57197.719 0.73 1.91 2.62
status
Size of 0.1611 0.05 /p57195.865 0.72 1.61 2.24
primary tumor
Source:Buzdar et al. (1978 ). Reprinted by permission of the editor.
/p63Favorable variables: treatment—FAC-BCG, postmenopausal, size of primary tumor 2cm.
Unfavorable variables: no adjuvant treatment, premenopausal, size of primary tumor 7cm.significant (p/p580.05)variables obtained was as follows:
logh/p71(t)
h/p15(t)/p58/p571.8792 (treatment /p570.47)/p590.9644 (menopausal status /p570.33)
/p590.1611 (size of primary tumor /p574.04)
Table12.3givesthedetailsofthefit.Themostimportantvariableinpredicting
survival time was the type of treatment (FAC-BCG favorable ); other signifi-
cantly important variables were menopausalstatus and size of primary tumor.Therisk ofdeathperunitof timefora patientreceivingno adjuvanttreatment(control group )was 6.55 times that for a patient receiving the treatment,
showing that FAC-BCG can prolong life considerably.308
Example 12.3 Suppose that demographic, personal, clinical, and labora-
tory data are collected from an interview and physical examination of 200participants in a study of cardiovascular disease (CVD ). These participants,
aged 50—79 years and free of CVD at the time of the baseline examination, are
then followed for 10 years. During the follow-upp eriod, 96 of the 200participants develop or die of CVD. We use this set of simulated data toillustrate further the use of the proportional hazards model in identifyingimportant risk factors. Table 12.4 gives a subset of the simulated data of 68
participants.
The event time Tof interest is CVD-free time, which is defined as the time
in years from baseline examination to the first time that a participant wasdiagnosed as having CVD or confirmed as a CVD death. CVD includescoronary heart disease (CHD )and stroke. The covariates of interest are age
(AGE ), gender (SEX /p581 if male and /p580 if female ); smoking status
(SMOKE /p581 if current smoker, and 0 otherwise ); body mass index
(BMI /p58weight in kilograms divided by height in meter squared ); systolic
blood pressure (SBP ); logarithm of ratio of urinary albumin and creatinine
(LACR ); logarithm of triglycerides (LTG ); hypertension status (HTN /p581i f
SBP/p46140 mmHg or DBP /p4690 mmHg or under treatments of hypertension,
and /p580 otherwise ); and diabetes status (DM /p581 if fasting glucose /p46126
mg/dL or under the treatments of diabetes, and /p580 otherwise ). For the CVD
outcome of interest, we let DG denote the type of CVD. DG /p580 if the
participant is free of CVD at the end of the study or confirmed as a non-CVDdeath (thus the CVD-freetime is censored ),/p581 if the participanthad a stroke,
/p582 if the participant had a CHD, and /p583 if the participant had other CVDs.
ItisofinteresttocomparetheriskofCVDamongthethreeagegroups:50 —59,
60—69, and70—79. We create twodummy variables:AGEA /p581 if aged 50—69,
/p580 otherwise; and AGEB /p581 if aged 60—69, and /p580 otherwise. Thus for a 70
to79-year-old,AGEA /p580andAGEB /p580.Wealsocreateavariabletodenote
the censoring status: CENS /p580 if t is censored, and /p581 if uncensored.
Toillustratethedifferentmethodstohandleties,wefittheCoxproportional
hazards model with the following six covariates: AGEA, AGEB, SEX,SMOKE, BMI, and LACR. The approximated partial likelihood functiondefined in (12.1.15 )—(12.1.17 )as well as the exact partial likelihood function
(Delong et al., 1994 )are applied. As noted in Sections 11.3 and 11.4, the
exponential and Weibull regression models are also proportional hazardmodels. Therefore, for comparisons we also fit an exponential and a Weibullregression model with the same covariates to the data. The estimated re-gression coefficients obtained from the proportional hazards model withapproximated discrete, Breslow, Efron, and exact partial likelihood functionsas well as those from the exponential and Weibull regression models are givenin Table 12.5. All of the estimates based on the Cox model and an approxi-mated partial likelihood function are very closed to those based on the exactpartial likelihood. Those based on Efron’s approximation are almost identicalto those (different only at the fourth decimal place )based on the exact partial 309
Table12.4 ASubsetoftheSimulatedDataforaCardiovascularDiseaseStudyinExample12.3 /p63
ID T CENS DG AGEA AGEB SEX SMOKE BMI SBP LACR LTG AGE HTN DM
1 7.4 0 0 0 0 0 0 31.78 141 4.23 3.94 77.8 1 0
2 7.9 0 0 0 0 0 0 25.02 124 4.31 4.66 76.9 0 1
3 6.4 0 0 0 0 0 1 26.05 111 4.38 4.27 76.3 0 04 7.1 0 0 0 0 0 1 26.92 140 1.11 4.51 72.2 1 0
5 6.0 0 0 0 0 0 1 34.30 146 1.19 4.82 76.0 1 0
6 6.5 0 0 0 0 0 1 31.76 142 1.20 4.88 74.5 1 07 8.3 0 0 0 0 1 0 25.01 154 3.53 4.10 70.7 1 1
8 7.9 0 0 0 0 1 0 28.21 136 3.73 4.12 75.2 1 0
9 7.6 0 0 0 1 0 0 28.13 127 2.92 4.24 64.9 0 0
10 8.4 0 0 0 1 0 0 25.68 118 2.47 4.41 60.2 0 0
11 7.4 0 0 0 1 0 0 34.34 118 2.37 4.46 64.4 0 1
12 7.7 0 0 0 1 0 0 28.92 127 3.58 4.55 68.8 1 113 6.9 0 0 0 1 0 1 24.68 100 2.11 4.33 64.4 0 0
14 7.2 0 0 0 1 0 1 21.93 121 3.39 4.64 60.8 0 1
15 6.3 0 0 0 1 0 1 29.47 98 1.96 4.69 64.4 0 0
16 7.4 0 0 0 1 0 1 28.65 150 2.59 4.95 61.6 1 0
17 4.5 0 0 0 1 1 0 32.28 128 2.99 4.73 65.3 0 0
18 7.0 0 0 0 1 1 0 29.21 117 2.17 4.91 65.7 0 119 2.8 0 0 0 1 1 0 28.82 136 4.04 4.92 65.4 0 1
20 7.2 0 0 0 1 1 0 30.58 121 2.84 4.94 64.5 0 1
21 7.4 0 0 1 0 0 0 27.83 95 1.85 4.44 52.0 0 022 5.2 0 0 1 0 0 0 26.61 128 2.87 4.51 50.7 0 0
23 7.7 0 0 1 0 0 0 30.32 96 2.41 4.60 52.5 0 1
24 7.8 0 0 1 0 0 0 30.41 130 1.45 4.73 55.9 0 025 7.6 0 0 1 0 0 1 29.98 140 1.88 4.51 53.4 1 0
26 7.9 0 0 1 0 0 1 26.00 118 2.34 4.53 51.0 0 0
27 7.3 0 0 1 0 0 1 29.05 110 1.44 4.67 50.6 0 0
310
288.2 0 0 1 0 0 1 27.21 131 2.50 4.68 57.7 0 0
29 3.8 0 0 1 0 1 0 36.97 141 4.60 4.25 58.7 1 0
30 6.9 0 0 1 0 1 0 29.44 115 2.89 4.26 53.6 0 131 6.1 0 0 1 0 1 0 33.85 154 3.48 4.48 51.2 1 0
32 7.2 0 0 1 0 1 0 32.13 122 2.92 4.48 55.2 0 0
33 8.4 0 0 1 0 1 1 27.52 135 2.39 4.42 53.7 0 034 5.0 0 0 1 0 1 1 30.64 114 1.39 4.45 54.9 1 0
35 6.5 0 0 1 0 1 1 29.94 120 2.96 4.49 50.7 0 0
36 6.4 0 0 1 0 1 1 29.89 115 1.68 4.52 51.3 0 037 2.6 1 1 0 0 0 0 30.88 189 5.38 4.72 73.9 1 1
38 2.7 1 1 0 0 0 1 25.05 200 3.37 4.86 77.2 1 1
39 2.7 1 1 0 0 1 0 26.80 130 2.31 5.10 73.5 0 040 3.3 1 1 0 0 1 1 21.67 111 3.53 4.18 71.1 0 0
41 2.9 1 1 0 1 0 0 36.83 114 2.64 4.52 68.2 0 0
42 0.2 1 1 0 1 0 1 21.49 125 4.61 4.69 67.3 0 043 2.1 1 1 0 1 1 0 31.05 131 1.38 4.48 69.1 0 0
44 6.8 1 1 0 1 1 1 26.78 134 4.36 4.90 61.0 1 0
45 5.7 1 1 1 0 0 0 35.78 132 9.93 5.11 52.5 0 146 1.1 1 1 1 0 0 1 28.44 134 3.54 4.32 55.7 0 0
47 6.6 1 1 1 0 1 0 24.38 124 4.16 4.00 51.8 0 1
48 1.3 1 1 1 0 1 1 34.13 126 5.87 3.95 53.1 0 149 4.6 1 2 0 0 0 0 43.23 128 5.08 5.25 72.2 0 1
50 6.3 1 2 0 0 0 1 38.67 126 5.16 4.50 76.8 1 1
51 2.0 1 2 0 0 1 0 34.49 130 2.69 3.95 76.7 1 152 4.2 1 2 0 0 1 1 20.78 127 4.40 4.54 73.1 0 0
53 3.6 1 2 0 1 0 0 28.40 118 5.43 4.66 69.3 1 1
54 3.2 1 2 0 1 0 1 28.73 154 1.94 5.24 68.9 1 155 4.5 1 2 0 1 1 0 44.25 97 2.01 4.40 68.6 0 1
56 4.5 1 2 0 1 1 1 32.46 141 0.74 4.39 63.5 1 0
57 6.1 1 2 1 0 0 0 39.72 118 2.39 3.93 52.6 0 1
(Continuedoverleaf )
311
Table12.4 Continued
ID T CENS DG AGEA AGEB SEX SMOKE BMI SBP LACR LTG AGE HTN DM
58 3.0 1 2 1 0 0 1 27.90 117 7.45 5.61 56.0 0 1
59 2.1 1 2 1 0 1 0 27.77 119 7.03 4.71 54.3 0 1
60 1.3 1 2 1 0 1 1 31.03 151 3.94 4.43 59.2 1 161 4.9 1 3 0 0 0 0 25.22 129 6.69 3.90 75.4 1 0
62 2.5 1 3 0 0 0 1 45.29 130 2.46 4.40 75.7 0 1
63 3.8 1 3 0 0 1 0 25.03 188 6.25 5.63 71.7 1 164 5.0 1 3 0 1 1 0 46.76 96 3.93 4.12 65.6 1 0
65 1.5 1 3 0 1 1 1 28.53 126 3.09 4.65 68.6 0 1
66 4.1 1 3 1 0 0 0 23.63 144 8.24 4.82 59.4 1 167 0.5 1 3 1 0 1 0 31.39 134 6.96 4.11 54.2 1 0
68 2.7 1 3 1 0 1 1 30.29 115 4.70 4.98 59.1 1 1
/p63ID, participant id number; T, CVD event time (CVD-free time );C E N S /p580 if censored, and /p581 if uncensored; DG /p580i fn o n - C V Da t
the end of the study or non-CVD death, /p581i fs t r o k e , /p582 if coronary heart disease (CHD ),a n d /p583 if the other CVDs; AGEA /p581i f
aged 50—59 and /p580o t h e r w i s e ;A G E B /p581i fa g e d6 0—69 and /p580o t h e r w i s e ;S E X /p581i fm a l ea n d /p580 if female; SMOKE /p581i fc u r r e n t
smoker and 0 otherwise; BMI, body mass index; SBP, systolic blood pressure; LACR, logarithm of the ratio of urinary albumin and
creatinine; LTG, logarithm of triglycerides; HTN /p581i fS B P /p46140mmHg or DBP (diastolic blood pressure )/p4690mmHg and /p580
otherwise; DM /p581 if fasting glucose /p46126mg/dL and /p580o t h e r w i s e .
312
Table12.5 ResultsfromFittingaCoxProportionalHazardsModelBasedonDifferent
MethodsforTiesontheCVDData
Regression Coefficient
Variable Breslow Discrete Efron Exact Exponential Weibull
AGEA /p571.3478 /p571.3662 /p571.3558 /p571.3560 /p571.2550 /p571.0436
AGEB /p570.7709 /p570.7828 /p570.7753 /p570.7755 /p570.7107 /p570.5966
SEX 0.7134 0.7233 0.7187 0.7189 0.6862 0.5659
SMOKE 0.3762 0.3810 0.3776 0.3776 0.3440 0.2855BMI 0.0253 0.0256 0.0255 0.0255 0.0233 0.0194
LACR 0.1735 0.1759 0.1739 0.1740 0.1658 0.1357
likelihood function. The estimated regression coefficients based on the two
parametric models, particularly the exponential regression model, are alsoclose to those based on the Cox hazards model. From the signs of thecoefficients,we see that men, current smokers,and persons with high BMIandalbumin—creatinine ratios have a higher hazard (risk)of CVD and shorter
CVD-free time. The coefficients of the two age variables are both negative,indicating that persons in the younger age groups have a lower hazard (risk)
of CVD.
Supposethat‘‘C: /p33EX12d2d1.DAT’’containseightsuccessivecolumns,forT,
CENS,AGEA,AGEB,SEX,SMOKE,BMI,andLACR,andthatthenumbersin each row are space-separated.The following code for the SAS PHREG andLIFEREG procedures can be used to obtain the results in Table 12.5.
data w1;
infile ‘c: /p33ex12d2d1.dat’ missover;
input t cens agea ageb sex smoke bmi lacr;
run;proc phreg;
model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58breslow;
run;proc phreg;
model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58discrete;
run;proc phreg;
model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58efron;
run;proc phreg;
model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58exact;
run;proc lifereg;
Model a: model t*cens (0)/p58agea ageb sex smoke bmi lacr / d /p58exponential;
Model b: model t*cens (0)/p58agea ageb sex smoke bmi lacr / d /p58weibull;
run; 313
12.2 IDENTIFICATIONOFSIGNIFICANTCOVARIATES
As noted earlier, one principal interest is to identify significant prognostic
factors or covariates. This involves hypothesis testing and covariate selectionprocedures, similar to those discussed in Chapter 11 for parametric methods.The differences are that the Cox proportional hazard model has a partiallikelihood function in which the only parameters are the coefficientsassociated with the covariates. However, statistical inference based on the
partial likelihood function has asymptotic properties similar to those basedon the usual likelihood. Therefore, the estimation procedure (discussed in
Section 12.1 )is similar to those in Section 7.1, and the hypothesis-testing
procedures are similar to those in Sections 9.1 and 11.2. For example, theWald statistic in (9.1.4 )can be used to test if any one of the covariates has no
effect on the hazard, that is, to test H/p15:b/p71/p580. By replacing the log-likelihood
function with the log partial likelihood function, the log-likelihood ratiostatistic, the Wald statistic, and the score statistic in (9.1.10 ),(9.1.11 ), and
(9.1.12 )can be used to test the null hypothesis that all the coefficients are
equal to zero, that is, to test
H/p15:b/p16/p580,b/p17/p580,...,b/p78/p580
orH/p15:b/p580in(9.1.9 ). Similarly the forward, backward, and stepwise selection
procedures discussed in Section 11.9.1 are applicable to the Cox proportionalhazard model.
The following example, using the SAS PHREG procedure, illustrates these
procedures.
Example 12.4 We use the entire CVD data set in Example 12.3 to
demonstrate how to identify the most important risk factors among all thecovariates. Suppose that the effects of age, gender, and current smoking statusonCVDriskareoffundamentalinterestandwewishtoincludethesevariablesin the model. In epidemiology this is often referred to as adjusting for thesevariables. Thus, AGEA, AGEB, SEX, and SMOKE are forced into the modeland we are to select the most important variables from the remainingcovariates (BMI,SBP,LACR,LTG,HTN,andDM ),adjustingforage,gender,
and current smoking status.
The SAS procedure PHREG is used with Breslow’s approximation for ties
(default procedure )and three variable selection methods (forward, backward,
and stepwise ). Two covariates, BMI and LACR, are selected at the 0.05
significance level by all three selection methods. The final model, in the formof(12.1.5 ),includingonlythefourcovariatesthatwepurposefullyincludedand
the two most significant ones identified by the selection method, is314
Table12.6 AsymptoticPartialLikelihoodInferenceontheCVDDatafromtheFinal
CoxProportionalHazardsModel /p63
95%Confidence
Interval
Regression Standard Wald Relative
Variable Coefficient Error Statistic pHazards Lower Upper
FinalModelfortheCohortCVDData
AGEA /p571.3558 0.2712 24.9910 0.0001 0.258 0.151 0.439
AGEB /p570.7753 0.2618 8.7709 0.0031 0.461 0.276 0.769
SEX 0.7187 0.2193 10.7457 0.0010 2.052 1.335 3.153
SMOKE 0.3776 0.2208 2.9235 0.0873 1.459 0.946 2.249
BMI 0.0255 0.0124 4.2113 0.0402 1.026 1.001 1.051LACR 0.1739 0.0446 15.2112 0.0001 1.190 1.090 1.299
b/p16/p57b/p17/p570.580 4 .9443 0.0262 0.560
b/p18/p59b/p191.096 11 .5409 0.0007 2.993
b/p18/p57b/p190.341 1 .3001 0.2542 1.407
HypothesisTestingResults (H/p15:allb/p71/p580)
Log-partial-likelihood ratio statistic 42.1130 0.0001
Score statistic 43.1750 0.0001Wald statistic 41.3830 0.0001
/p63The covariates, except AGEA, AGEB, SEX, and SMOKE, in the final model are selected among
BMI, SBP, LACR, LTG, HTN, and DM.logh(t/p71)
h/p15(t/p71)/p58b/p16AGEA/p71/p59b/p17AGEB/p71/p59b/p18SEX/p71/p59b/p19SMOKE/p71
/p59b/p20BMI/p71/p59b/p21LACR/p71
/p58/p571.3558AGEA/p71/p5707753AGEB/p71/p590.7187SEX/p71
/p590.3776SMOKE/p71/p590.0255BMI/p71/p590.1739LACR/p71(12.2.1 )
The regression coefficients, their standard errors, the Wald test statistics, p
values, and relative hazards (relative risks as they are termed by many
epidemiologists )are given in Table 12.6. The estimated regression coefficients
b/p19/p71,i/p581, 2,...,6, are solutions of (12.1.9 )using the Newton —Raphson iterated
procedure (Section 7.1 ). The estimated variances of b/p19/p71,i/p581, 2,...,6, are the
respective diagonal elements of the estimated covariance matrix defined in(12.1.13 ).The square rootsof these estimatedvariancesare thestandarderrors
in the table. The Wald statistics are for testing the null hypothesis that thecovariate is not related to the risk of CVD or H/p15:b/p71/p580,i/p581,...,6, respect-
ively. For example, the Wald statistic equals 10.7457 for gender with a pvalue 315
of 0.0010 and b/p580.7187. It indicates that after adjusting for all the variables
in the model (12.2.1 ), gender is a significant predictor for the development of
CVD,withmenhavingahigherriskthanwomen.Therelativehazard (orrisk )
isexp(b/p19/p16),andforthecovariategender,it isexp (0.7187 )/p582.052,whichimplies
that men aged 50 —79 years have about twice the risk of developingCVD in 10
years. The 95%confidence interval for the relative risk is (1.335, 3.153 ), which
is calculated according to (7.1.8 ). For a continuous variable, exp( b/p19/p71)represents
the increase in risk corresponding to a 1-unit increase in the variable. For
example,for BMI,exp (0.0255 )/p581.026; that is, for every unitincreasein BMI,
the risk for CVD increases 2.6%.
To compare hazards among different age groups, between genders, or
between smokers and nonsmokers, let h/p31/p37/p35/p31(t),h/p31/p37/p35/p32(t),h/p31/p37/p35/p33(t),h/p43/p31/p42(t),
h/p36/p35/p43(t),h/p49/p43(t), andh/p44/p49/p43(t) denote hazard functions for participants that are
50—59, 60—69, 70—79 years old, male, female, current smoker, and not current
smoker, respectively. The log hazard ratio of a person in the 50 to 59-yearage group to a person in the 70 to 79-year group assuming the two people areof same gender and the same current smoking status, BMI and LACR, islog[h/p31/p37/p35/p31(t)/h/p31/p37/p35/p33(t)]/p58b/p16; similarly, log[ h/p31/p37/p35/p32(t)/h/p31/p37/p35/p33(t)]/p58b/p17and
log[h/p31/p37/p35/p31(t)/h/p31/p37/p35/p32(t)]/p58b/p16/p57b/p17. Assuming that the two people are in the
same age groupand have the same BMI and LACR, the log hazard ratio ofmale to females is
logh/p43/p31/p42(t)
h/p36/p35/p43(t)
/p58b/p18
Similarly,assuming thatthe two people arein the same agegroup,of the same
gender, and have the same BMI and LACR, the hazard ratio of a smoker to anonsmoker is
logh/p49/p43(t)
h/p44/p49/p43(t)/p58b/p19
Thus, testing whether risk of CVD are the same among different age groups is
equivalent to testing H/p15:b/p16/p580,H/p15:b/p17/p580, andH/p15:b/p16/p57b/p17/p580. Similarly, to
test if the risk of CVD is the same between males and females or betweensmokersandnonsmokersisequivalenttotastingthenullhypothesis H/p15:b/p18/p580
orH/p15:b/p19/p580, respectively.
To consider more than one covariate, we also can formulate the null
hypothesis by using (12.2.1 ). For example, if we wish to compare male
nonsmokers to female smokers, from (12.2.1 ),
logh/p43/p31/p42/p92/p44/p49/p43h/p36/p35/p43/p92/p49/p43 /p58b/p18/p57b/p19316
assuming that they are in the same age groupand have the same BMI and
LACR. Thus to test if these two groups of people have the same risk of CVD,we test the null hypothesis H/p15:b/p18/p57b/p19/p580. Similarly, to compare male
smokerstofemalenonsmokers,wecantestthenullhypothesis H/p15:b/p18/p59b/p19/p580.
Thesenullhypothesesareintheformoflinearcombinationsofthecoefficients.Using the notations in Section 11.2, the hypotheses H/p15:b/p16/p57b/p17/p580 and
H/p15:b/p18/p59b/p19/p580 are the hypotheses in (11.2.13 )withc/p580,L/p58(1/p5710000 ),
andL/p58(001100 ), respectively. The Wald statistics in Table 12.6 are
calculated according to (11.2.14 ). By assuming that the patients have the same
BMI and LACR, we can construct hypotheses to compare subgroups definedby age groups, gender, and current smoking status.
The last part of Table 12.6 shows the results of testing the null hypothesis
that none of these covariates have any effect on the development of CVD. Thelog partial likelihood ratio, Wald, and score statistics, X/p42,X/p53, andX/p49are
calculated according to (9.1.10 ),(9.1.11 ), and (9.1.12 ), respectively. Table 12.6
indicates that the hypotheses, H/p15:b/p16/p580,H/p15:b/p17/p580,H/p15:b/p16/p57b/p17/p580,
H/p15:b/p18/p580,H/p15:b/p20/p580,H/p15:b/p21/p580,andH/p15:b/p18/p59b/p19/p580 are rejected ata signifi-
cance level of p/p580.05. However, the hypotheses H/p15:b/p19/p580 and
H/p15:b/p18/p57b/p19/p580 are not rejected at a 0.05 level. The null hypothesis
H/p15:allb/p71/p580,i/p581,...,6, is rejected with p/p580.0001 by using any of these
tests.
Assuming that the other covariates are the same, based on the relative
hazards shown in the table, we conclude that (1)participants aged 50 —59
and 60—69 have, respectively, about 25% and 50% lower CVD risk than
those aged 70 —79(H/p15:b/p16/p580 andH/p15:b/p17/p580 are rejected );(2)participants
aged50—59have50%lowerCVDriskthanthoseaged60 —69(H/p15:b/p16/p57b/p17/p580
is rejected );(3)men’s CVD risk is twice as high as that of women (H/p15:b/p18/p580
is rejected );(4)BMI and LACR have a significant effect on CVD risk
(H/p15:b/p20/p580 andH/p15:b/p21/p580 are rejected )and the risk increases about 3% and
19%,respectively,forevery1-unitincreaseinBMIandLACR,respectively; (5)
male smokers have a CVD risk three times higher than that of femalenonsmokers (H/p15:b/p18/p59b/p19/p580 is rejected );(6)male nonsmokers have CVD risk
similartothatoffemalesmokers (H/p15:b/p18/p57b/p19/p580isnotrejected );(7)consider-
ing current smoking status alone, smokers had similar CVD risk as non-smokers (H/p15:b/p19/p580 is not rejected ). This example is solely for the purpose of
illustrating the use of the proportional hazards model and the interpretationof its results. Other hypotheses of interest can be constructed in a similarmanner. The construction of null hypotheses for comparisons among sub-groups defined by AGEGROUP*SEX*SMOKE are left to the reader asexercises.
Suppose that ‘‘C: /p33EX12d4d1.DAT’’ is a text data file that contains
12 successive columns for T, CENS, AGEA, AGEB, SEX, SMOKE, BMI,LACR,SBP,LTG,HTN,andDM.ThefollowingSAScodeisusedtoobtainedthe results in Table 12.6. 317
data w1;
infile ‘c: /p33ex12d4d1.dat’ missover;
input t cens agea ageb sex smoke bmi lacr sbp ltg htn dm;
run;
proc phreg data /p58w1;
model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm /
include /p584 selection /p58f;
run;
proc phreg data /p58w1;
model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm /
include /p584 selection /p58b;
run;proc phreg data /p58w1 outest /p58wcov covout;
model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm /
include /p584 selection /p58s;
run;
proc phreg data /p58w1;
model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm /
include /p584 selection /p58score best /p583;
run;data wcov;
set wcov;
if-type-/p58‘cov’;
keepagea ageb sex smoke bmi lacr sbpltg htn dm;
run;title ‘The estimated covariance of the estimated coefficients’;proc print data /p58wcov;
run;
The following SPSS code can be used to select an optimal subset of
covariates among all covariates by the forward and backward selectionmethods defined in Section 11.9.1 and to obtain the estimated coefficients andthe other results in Table 12.6.
data list file /p58‘c:/p33ex12d4d1.dat’ free
/ t cens agea ageb sex smoke bmi lacr sbpltg htn dm.
coxreg t with agea ageb sex smoke bmi lacr sbpltg htn dm
/status /p58cens event (1)
/method /p58fstepbmi lacr sbpltg htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all.
coxreg t with agea ageb sex smoke bmi lacr sbpltg htn dm
/status /p58cens event (1)
/method /p58bstepbmi lacr sbpltg htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all.318
If BMDP 2L is used, the following code is applicable when selecting an
optimal subset of covariates among all covariates by the stepwise selectionmethod defined in Section 11.9.1 and to obtain the results in Table 12.6.
/input file /p58‘c:/p33ex12d4d1.dat’ .
variables /p5812.
format /p58free.
/print cova.
/variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg,
htn, dm.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, htn,
dm.
Step /p58phh.
Example 12.5 If we do not force age, gender, and current smoking status
on the model and are not interested in the three age groups, we can fit theproportional hazard model with age as a continuous variable and the othercovariates: SEX, SMOKE, BMI, SBP, LACR, LTG, HTN, and DM. UsingBreslow’s method for ties, the stepwise selection method, and the SAS pro-cedure PHREG, the final model with significant (p/p580.05)covariates is
logh(t)
h/p15(t)
/p580.697AGE /p590.7528SEX /p590.1111LACR /p590.3987LTG
(12.2.2 )
The details are given in Table 12.7; all four covariates in the model have
positive coefficients, indicating that the risk of developing CVD increases withage, gender, albumin/creatinine ratio, and triglyceride values. The relativehazards represent the increase in risk of CVD per unit increase in thecovariates. For example, for every 1-unit increase in log (albumin/creatinine ),
the risk of developing CVD increases 12%after adjusting for age, gender, andlog triglyceride. Men have more than twice the risk of CVD as women. Theglobal null hypothesis that all four coefficients equal zero ( H/p15:allb/p71/p580)is
rejected by all three tests, as given in the lower part of Table 12.7.
12.3 ESTIMATIONOFTHESURVIVORSHIPFUNCTIONWITH
COVARIATES
Whenparametricregressionmodels (Chapter11 )areused,wecanestimatethe
survivorshipfunctionsimplybyreplacingtheparametersandcoefficientsinthe
survival function with their estimates. This is not the case when the Cox 319
Table12.7 AsymptoticPartialLikelihoodInferenceontheCVDDatafromtheFinal
CoxProportionalHazardsModelSelectedbytheStepwiseModelSelectionMethod /p63
95%Confidence
Interval for
Relative Hazards
Regression Standard Chi-Square Relative
Variable Coefficient Error Statistic pHazards Lower Upper
AGE 0.0697 0.0136 26.1393 0.0001 1.07 1.04 1.10
SEX 0.7528 0.2192 11.7893 0.0006 2.12 1.38 3.26LACR 0.1111 0.0459 5.8602 0.0155 1.12 1.02 1.22
LTG 0.3987 0.1976 4.0722 0.0436 1.49 1.01 2.20
H/p15:All coefficients equal zero
Log-partial-likelihood ratio statistic 44.002 0.0001
Score statistic 44.278 0.0001
Wald statistic 42.527 0.0001
/p63The covariates in the final model are selected among AGE, SEX, SMOKE, BMI, LACR, LTG,
HTN, and DM using the stepwise selection method.
proportional hazards model is used since we do not know the exact form of
the baseline hazard function or the survival function. In this section weintroduce briefly two estimators of the survival function, one proposed byBreslow (1974 )and the other by Kalbfleisch and Prentice (1980 ). These
estimates are available in commercialsoftware packages. Readers interested indetails are referred to the corresponding publications.
As indicated earlier, under the Cox model, the survivorshipfunction with
covariatesx/p72’s is
S(t,x)/p58[S/p15(t)]
exp(/afii9814/p78/p72/p14/p16b/p72x/p72)(12.3.1)
Once the regression coefficients, the b/p72’s, are estimated, we need only estimate
the underlying survivorshipfunction, S/p15(t). From the estimated survivorship
function,wecaneasilyestimatetheprobabilityofsurvivinglongerthanagiventime for a patient with a given set of covariates x/p16,...,x/p78.
Byassumingthatthebaselinehazardfunctionisconstantbetweeneachpair
of successive observed failure times, Breslow has proposed the followingestimator of the baseline cumulative hazard function:
H/p19/p15(t)/p58/p26
t/p7/p71/p8/p45tm/p7/p71/p8/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)(12.3.2 )320
Following (2.15), the baseline survival function can be estimated as
S/p19/p15(t)/p58exp[ /p57H/p19/p15(t)]/p58/p147
t/p7/p71/p8/p45t/p7exp/p3m/p7/p71/p8/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)/p4/p8(12.3.3 )
and the survivorshipfunction for a p erson with a set of covariates
x/p58(x/p16,...,x/p78)is
S/p19(t,x)/p58[S/p19/p15(t)]exp(/afii9814/p78/p72/p14/p16b/p19/p72x/p72)/p58[S/p19/p15(t)]exp(b/p19/p30x)(12.3.4 )
Under mild assumptions, S/p19(t,x)has an asymptotic normal distribution with
meanS(t,x). SinceS(t,x)/p58exp[ /p57H(t,x)], the variance estimator Var /p19(S/p19(t,x))
ofS/p19(t,x)is
Var/p19(S/p19(t,x))/p60[S/p19(t,x)]/p17Var/p19(H/p19(t,x))
We will not give H/p19(t,x)here because of its complexity. The asymptotic
confidence bands for the survivorshipfunction is
/p37S/p19(t,x)/p57Z/p63/p30/p17/p40Var/p19(S/p19(t,x)),S/p19(t,x)/p59Z/p63/p30/p17/p40Var/p19(S/p19(t,x))/p38 (12.3.5 )
whereZ/p63/p30/p17is the upper 100 (1/p57/afii9825/2)percentile point of the standard normal
distribution.
An alternative estimator has been suggested by Kalbfleisch and Prentice in
whichthebaselinesurvivorshipfunction S/p15(t) isestimatedtobeastepfunction
and
S/p19/p15(t)/p58/p71/p92/p16/p147
/p72/p14/p15/afii9825/p24/p72t/p7/p71/p92/p16/p8/p58t/p45t/p7/p71/p8,i/p581,...,k/p591 (12.3.6 )
where /afii9825/p24/p15/p891and /afii9825/p24/p16,/afii9825/p24/p17,...,/afii9825/p24/p73arethesolutionofthefollowing ksimultaneous
equations:
/p26
j/p43u*/p7/p71/p8exp(x/p30/p72b/p19)
1/p57/afii9825/p24/p71exp(x/p30/p72b/p19)/p58/p26
l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)i/p581,...,k(12.3.7)
When there are no ties,
/afii9825/p24/p71/p58/p31/p57exp(x/p30/p7/p71/p8b/p19)
/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)/p4exp(/p57x/p30/p7/p71/p8b/p19)
i/p581,...,k(12.3.8)
and
S/p19/p15(t)/p58/p71/p92/p16/p147
/p72/p14/p15/p31/p57exp(x/p30/p7/p72/p8b/p19)
/p26l/p43R(t/p7/p72/p8)exp(x/p30/p74b/p19)/p4exp(/p57x/p30/p7/p72/p8b/p19)
t/p7/p71/p92/p16/p8/p45t/p58t/p7/p71/p8i/p581,...,k/p591 321
Thus,
S/p19(t,x)/p58[S/p19/p15(t)]exp(b/p19/p30x)(12.3.9 )
Undermildassumptions,theKalbfleischandPrenticeestimatorin (12.3.9 )also
follows an asymptotic normal distribution with mean S(t,x)and a variance
thatcanbeestimated.Thusconfidencebandsforthesurvivorshipfunctioncanalso be constructed.
Using (12.3.4 )withS/p15(t)i n(12.3.3 )or(12.3.6 ),the survivorshipfunctioncan
be estimated with any given values of x/p16,...,x/p78. If the observed average of
every covariate, x/p21/p16,...,x/p21/p78is used, the estimated survivorshipfunction can be
interpreted as the survivorship function of an ‘‘average’’ person.
Both the Breslow and Kalbfleisch —Prentice estimators are available in the
SAS procedure PHREG. The Breslow estimator is also available in BMDP(program 2L )and SPSS (program COXREG ). The following example illus-
trates the procedures.
Example 12.6 Again, we use the CVD data in the Example 12.3, the data
set‘‘C: /p33EX12d2d1.DAT’’,andtheSASprocedurePHREG.Weusetheaverage
of each of the covariates in (12.2.1 ), and therefore the estimated survivorship
function is for an average person. The Kalbfleisch —Prentice and Breslow
estimates of the survival function, defined in (12.3.9 )and (12.3.4 )(Efron
adjustment for ties is used ), and the lower and upper 95% confidence bands,
calculated based on (12.3.5 ), are shown in Figures 12.1 and 12.2. These
estimatedsurvivalfunctions,using all the covariatesin the model with averagevalues, are often referred to as the global covariate —adjusted survivorship
functions. The two figures are almost identical, which indicates that the twomethods produce very similar results for this set of data. From Figure 12.1 itappears that the global covariates —adjusted survivorshipfunction decreases
somewhatmore rapidly after 3.5 years. This means that the process to developCVD accelerates after 3.5 years.
Using the data set ‘‘C: /p33EX12d2d1.DAT’’ defined in Example 12.3, the SAS
code used for this example is the following.
data w1;
infile ‘c: /p33ex12d2d1.dat’ missover;
input t cens agea ageb sex smoke bmi lacr;
run;
proc phreg data /p58w1 noprint;
model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58efron;
baseline out /p58base1 survival /p58survival l /p58lowb u /p58uppb / method /p58pl;
run;title ’K-P estimate of the survival function and its lower and upper bands’;
proc print data /p58base1;
var t survival lowb uppb;
run;322
Figure 12.1 Kalbfleisch—Prentice estimate of survivorshipfunction and its 95%
confidence bands at the averages of the covariates from the fitted Cox proportional
hazards model on the CVD data.
proc phreg data /p58w1 noprint;
model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58efron;
baseline out /p58base1 survival /p58survival l /p58lowb u /p58uppb / method /p58ch;
run;
title ’Breslow estimate of the survival function and its lower and upper bands’;proc print data /p58base1;
var t survival lowb uppb;
run;
The following SPSS code can be used to obtain the Breslow estimate of the
survival function and its standard error at each uncensored observation. Theconfidence bands can then be calculated according to (12.3.5 ).
data list file /p58‘c:/p33ex12d2d1.dat’ free
/ t cens agea ageb sex smoke bmi lacr.
coxreg t with agea ageb sex smoke bmi lacr
/status /p58cens event (1)
/print /p58all. 323
Figure 12.2 Breslow estimate of the survivorshipfunction and its 95% confidence
bands at the averages of the covariates from the fitted Cox proportionalhazards model
on the CVD data.
The corresponding BMDP 2L code is
/input file /p58‘c:/p33ex12d2d1.dat’ .
variables /p588.
format /p58free.
/print cova.
Survival.
/variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58agea, ageb, sex, smoke, bmi, lacr.
In addition to the global covariates —adjusted survivorshipfunction defined
asS/p19(t,x/p21),wherex/p21/p58(x/p21/p16,x/p21/p17,...,x/p21/p78),thesurvivorshipfunctioncanbeestimated
with any specific values of one or more of the covariates and interactions. Wecan also estimate the probability of surviving longer than a given time forindividuals with a given set of values for covariates. The following is anexample.324
Figure12.3 Breslow estimate of survivorshipfunctions at the averages of BMI and
LACRfromSEX*SMOKERsubgroupsin aged 70 —79 participantsfrom thefitted Cox
proportional hazards model on the CVD data.
Example 12.7 ForthesamemodelasinExample12.6,we canestimatethe
covariate-specific survivorship function for female nonsmokers, femalesmokers,malesmokers,andmalenonsmokers.Let ususethe70 —79agegroup
and assume that BMI and LACR are at the average of the respectiveSEX—SMOKE subgroup. Thus, the specific covariate vector (AGEA, AGEB,
SEX, SMOKE, BMI, LACR )for female nonsmokers is (0, 0, 0, 0, 30.69, 4.62 ),
where 30.69 and 4.62 are the average values of BMI and LACR for femalenonsmokers. Similarly, the specific covariate vectors for female smokers, malenonsmokers, and male smokers are, respectively, (0, 0, 0, 1, 31.19, 2.67 ),(0, 0,
1, 0, 28.19, 3.43 ), and (0, 0, 1, 1, 25.76, 3.47 ). The estimated survival curves are
shown in Figure 12.3. Similarly, Figures 12.4 and 12.5 give the estimatedsurvival curves of the four groups in persons aged 60 —69 years and 50 —59
years, respectively. The groups show that in all these age groups, females havea lower risk of developing CVD (longer CVD-free time )than males. Female
nonsmokershave a slightly lower risk than female smokers and the differencesincrease as age decreases. However, among males, the differences in the risk ofCVD between smokers and nonsmokers are almost negligible in the youngestgroupandmuchlargerinthetwooldergroups.Malesmokershavethehighestrisk of developing CVD (shortest CVD-free time )among the four groups. 325
Figure12.4 Breslow estimate of survivorshipfunctions at the averages of BMI and
LACRfromSEX*SMOKERsubgroupsin aged 60 —69 participantsfrom thefitted Cox
proportional hazards model on the CVD data.
12.4 ADEQUACYASSESSMENTOFTHEPROPORTIONAL
HAZARDSMODEL
Thevalidityofstatisticalinferencesthatleadstotheidentificationofimportant
risk or prognostic factors depends largely on the adequacy of the modelselected. The proportional hazards model is used widely in medical andepidemiologicalstudies. The adequacyof this model, includingthe assumptionof proportional hazards and the goodness of fit, needs to be assessed. In thissection we introduce several methods for this purpose. A major reason forselectingthese methodsto present here is the availabilityof computer softwarethat can perform the calculations.
12.4.1 CheckingtheProportionalHazardsAssumption
The proportional hazards models defined in (12.1.1 )and (12.1.3 )assume that
the hazard ratio of two people is independent of time. This requires thatcovariates not be time-dependent. If any of the covariates varies with time, theproportional hazards assumption is violated. This fact can be used to test theassumption by including a time —covariate interaction term in the model and326
Figure12.5 Breslow estimate of survivorshipfunctions at the averages of BMI and
LACRfromSEX*SMOKERsubgroupsin aged 50 —59 participantsfrom thefitted Cox
proportional hazards model on the CVD data.
testing if the coefficient for interaction is significantly different from zero. For
example, we can add an interaction term x/p71torx/p71logtin the model, that is,
logh(t)
h/p15(t)/p58b/p16x/p16/p59/p37/p59b/p71x/p71/p59b/p71/p16x/p71t/p59b/p71/p62/p16x/p71/p62/p16/p59/p37/p59b/p78x/p78
or
logh(t)
h/p15(t)/p58b/p16x/p16/p59/p37/p59b/p71x/p71/p59b/p71/p16x/p71logt/p59b/p71/p62/p16x/p71/p62/p16/p59/p37/p59b/p78x/p78
With the added interaction term, the partial likelihood function becomes more
complicated. Fortunately, computer software is available to carry out thecalculations. Testing procedures similar to those discussed earlier (e.g., the
Waldtest ), can be used to testthe null hypothesis H/p15:b/p71/p16/p580. IfH/p15is rejected,
we conclude that Cox’s proportional hazard model is not appropriate for thedata. The interaction term with log tcan be included in the model for each of
the covariates separately. If none of the corresponding pnull hypotheses
H/p15:b/p71/p16/p580isrejected,wemayconcludethattheproportionalhazardsassump-
tion is appropriate. 327
Table12.8 AsymptoticPartialLikelihoodInferenceontheCVDDatafromtheCox
ProportionalHazardsModelwithTime-DependentCovariate
95%Confidence
Interval for
Relative Hazards
Regressor Regressor Standard Wald Relative
Variable Coefficient Error Statistic pHazards Lower Upper
(a)
AGE 0.068 0.014 25.249 0.0001 1.07 1.04 1.1
SEX 0.759 0.218 12.056 0.0005 2.14 1.39 3.28LACR 0.111 0.046 5.781 0.0162 1.12 1.02 1.22
LTG 0.915 0.435 4.420 0.0355 2.50 1.06 5.86
LTG* /p570.390 0.298 1.710 0.1910 0.68 0.38 1.22
log(t/p591)
(b)
AGE 0.071 0.014 26.635 0.0001 1.07 1.05 1.1
SEX 0.741 0.220 11.327 0.0008 2.10 1.36 3.23LACR /p570.087 0.120 0.519 0.4714 0.92 0.72 1.16
LTG 0.395 0.199 3.917 0.0478 1.48 1 2.19
LACR* 0.143 0.079 3.269 0.0706 1.15 0.99 1.35
log(t/p591)
(c)
AGE 0.038 0.033 1.330 0.2488 1.04 0.97 1.11
SEX 0.764 0.220 12.020 0.0005 2.15 1.39 3.31LACR 0.111 0.046 5.888 0.0152 1.12 1.02 1.22
LTG 0.417 0.197 4.469 0.0345 1.52 1.03 2.24
AGE*log(t
/p591) 0.023 0.023 1.046 0.3064 1.02 0.98 1.07Example 12.8 Considerthefittedproportionalhazardsmodelin (12.2.2 )for
the CVD data. To check the proportional hazards assumption, we add a termLTG/p59log(t/p591) to the model. We use t/p591 instead of tto avoid negative
values. Table 12.8 (a)gives the results. The pvalue for the interaction term is
0.1910. Similarly, the results in Table 12.9 (b)and (c)suggest that
LACR/p59log(t/p591) andAGE /p59log(t/p591) arenotsignificanteither.Sincegender
is time-independent, we may conclude that the data satisfy the proportionalhazards assumption since every covariate in the model is time-independent.
Anothermethodtochecktheproportionalhazardsassumptionis tostratify
the data based on some values of a covariate, fit a stratified Cox proportionalhazards model (this is discussed in Chapter 13 ), and then construct the
survivorship function separately for the each stratum and plot
log(/p57log(S/p19/p72(t;x/p21/p72)))j/p581, 2,...,m328
Figure 12.6 Log[ /p57log(S(t))] plots for the age-stratified Cox proportional hazards
model on the CVD data.againsttimet,wheremisthenumberofstratadefinedbythecovariate, x/p21/p72isthe
vectoroftheaveragevaluesoftheothercovariatesforthe jthstratum,and S/p19/p72(t;x/p21/p72)
istheestimatedsurvivorshipfunctionofthe jthstratumevaluatedat tandx/p21/p72.If
thehazardsareproportional,the mcurvesshouldbeparallel.Nonparallelcurves
indicatedeparturefromtheproportionalhazardsassumption.Thisisbecauseifhazard functions from any two people are proportional, it can be shown from(12.1.1 )that,forany j/p34kand1 /p45j,k/p45m,thereexistsaconstant d/p72/p73suchthat
S/p19/p72(t;x/p21/p72)/p58(S/p19/p73(t;x/p21/p73))/p66
/p72/p73 (12.4.1 )
Taking the logarithm twice, we have
log[/p57log(S/p19/p72(t;x/p21/p72))]/p58logd/p72/p73/p59log[/p57log(S/p19/p73(t;x/p21/p73))] (12.4.2 )
Thusthecurvesoflog[ /p57log(S/p19/p72(t;x/p21/p72))]andlog[ /p57log(S/p19/p73(t;x/p21/p73))]versustshould
be parallel.
Example 12.9 Consider again the fitted model in (12.2.2 ); using the
stratified analysis (more details are given in Chapter 13 ), we plot
log[/p57logS/p19/p72(t;x/p21/p72)] againsttfor two age strata (50—64 and 65—79 years )and
two gender strata separately, where x/p21/p72denotes the average values of the other
covariatesfor the jth stratum. These graphs are givenin Figures 12.6 and 12.7, 329
Figure 12.7 Log[ /p57log(S(t))] plots for gender-stratified Cox proportional hazards
model on the CVD data.
respectively.ThetwocurvesinFigure12.6areroughlyparallel.Thetwocurves
in Figure 12.7 are also parallel over time. The results suggest that theproportional hazards assumption holds.
InChapter11wediscussedseveralparametricmodels.Amongthesemodels,
the exponential and the Weibull are proportional hazards models, but theothers are not. Thus, if one of the other models providesa good fit to data, wewould know that the data do not meet the proportional hazards assumption.This procedure can also be served as an alternative for checking the propor-tional hazards assumption.
12.4.2 AssessingGoodnessofFitbyResiduals
There are several other graphicalmethods availablefor assessing the goodness
of fit of a proportional hazards model. These graphical methods are based onresiduals and are often used as diagnostic tools. In multiple regressionmethods, residuals are referred to as the difference between the observed andthepredictedvalues (basedontheregressionmodel )ofthedependentvariable.
However,whencensoredobservationsarepresentandonlyapartiallikelihoodfunction is used in the proportional hazards model, the usual concept ofresiduals is not applicable. In the following we introduce three different types330
ofresiduals:theextendedCox —Snell,deviance,andSchoenfeldresiduals.These
canbe plottedversusthesurvivaltime ora covariate.Thepatternofthegraphprovides some information about the appropriateness of the proportionalhazards model. It also provides information about outliers and other patterns.Similarto other graphical methods, interpretation of the residual plots may besubjective.
The Cox—Snell method discussed in Section 8.4 can easily be extended to
the proportional hazards model. The extended Cox —Snell residual, R/p71, for the
ith individual with observed survival time tand covariates at values x/p71is
defined asR/p71/p58/p57logS/p19(t/p71;x/p71), which is the estimated accumulated hazard
based on the proportional hazards model. If the t/p71observed is censored, the
corresponding R/p71is also censored. If the proportional hazards model is
appropriate, the plot of R/p71and its Kaplan —Meier estimate of survival function
(S/p19(R)) would appear as a 45° straight line. The Cox —Snell residual method is
useful in assessing the goodness of fit of a parametric model (Section 11.9.4 ).
However, it is not so desirable for a proportional hazards model where apartial likelihood function is used and the survivorship function is estimatedby nonparametric methods.
The deviance residuals (Therneau et al., 1990 )are defined as
R/p34/p71/p58sign(R/p43/p71)/p402[/p57R/p43/p71/p57/afii9829/p71log(/afii9829/p71/p57R/p43/p71)]
i/p581, 2,...,n(12.4.3)
where sign (·)is the sign function, which takes value 1 if its argument is
positive, 0 if zero, and /p571 if negative, R/p43/p71is the martingale residual (Fleming
and Harrington, 1991 )for theith individual,
R/p43/p71/p58/afii9829/p71/p57R/p71i/p581,...,n
and/afii9829/p71/p581 if the observed survival time t/p71is uncensored and 0 otherwise.
The martingale residuals have a skewed distribution with mean zero
(AndersonandGill,1982 ).Thedevianceresidualsalsohaveameanofzerobut
are symmetrically distributed about zero when the fitted model is adequate.Devianceresidualsare positiveforpersonswho survivefora shorter timethanexpectedandnegative forthose whosurvive longer.Thedevianceresidualsareoften used in assessing the goodness of fit of a proportional hazards model.
Another residual method was proposed by Schoenfeld (1982 )and modified
by Grambsch and Therneau (1994 ). The original Schoenfeld residuals are
definedforeachpersonandeachcovariateandarebasedonthefirstderivativeof the log-likelihood function in (12.1.9 ). A Schoenfeld residual for the jth
covariate of the ith person with the observed survival time t/p71is
R/p72/p71/p58/afii9829/p71
/p3x/p72/p71/p57/p26l/p43R(t/p7/p71/p8)x/p72/p74exp(b/p19/p30x/p74)
/p26l/p43R(t/p7/p71/p8)exp(b/p19/p30x/p74)/p4j/p581, 2,...,p;i/p581, 2,...,n
(12.4.4) 331
whereb/p19is the maximum partial likelihood estimator of b. The Schoenfeld
residuals are defined only at uncensored survival times; for censored observa-tions they are set as missing. Since b/p19is the solution of (12.1.9 ), the sum of the
Schoenfeld residuals for a covariate is zero. Thus asymptotically, the Schoen-feld residuals have a mean of zero. It can also be shown that these residualsare not correlated with one another.
Grambsch and Therneau (1994 )suggested that the Schoenfeld residuals be
weighted by the inverse of the estimated covariance matrix of R/p71/p58
(R/p16/p71,...,R/p78/p71)/p30denoted byV/p19(R/p71), that is,
R*/p71/p58[V/p19(R/p71)]/p92/p16R/p71(12.4.5 )
The weighted Schoenfeld residuals have better diagnostic power and are used
moreoftenthantheunweightedresidualsinassessingtheproportionalhazardsassumption. To simplify the computations, Grambsch and Therneau (1994 )
suggested an approximation of [ V/p19(R/p71)]/p92/p16in(12.4.5 ):
[V/p19(R/p71)]/p92/p16/p60rV/p19(b/p19)
whereris thenumberofeventsorthenumberofobserveduncensoredsurvival
times andV/p19(b/p19)is the estimated covariance matrix of b/p19in(12.1.13 ). With this
approximation, the weighted Schoenfeld residuals in (12.4.5 )can be approxi-
mated by
R*/p71/p58rV/p19(b/p19)R/p71(12.4.6 )
The graphs of deviance and Schoenfeld residuals against survival time or a
covariatecanbeusedtochecktheadequacyoftheproportionalhazardsmodel.The presence of certain patterns in these graphs may indicate departures fromthe proportional hazards assumption, while extreme departures from the maincluster indicate possible outliers or potential stability problems of the model.
Example 12.10 Consider the proportional hazards model (12.2.2 )for the
CVD data. Using the estimated survivorshipfunction with covariates, weobtain the extended Cox —Snell residual R/p71values and plot the Kaplan —Meier
estimateof the survivorshipfunctionof the R/p71’s. Figure12.8 givesthe extended
Cox—Snellresidualplot.Theconfigurationisveryclosetoa45°line,indicating
that the proportional hazards model (12.2.2 )provides a reasonable fit to the
data.
Figure 12.9 plots the deviance residuals against t. Roughly speaking, the
residualsaredistributedsymmetricallyaroundzerobetween /p573and3 withno
peculiar patterns. Larger positive (negative )residuals are associated with
smaller (larger )tvalues. The deviance residuals suggest that the proportional
hazards model provides a reasonable fit to the data.332
Figure12.8 Cox—Snell residuals plot from the fitted Cox proportional hazards model
on the CVD data.
The weighted Schoenfeld residuals versus AGE, LACR, and LTG are given
in Figures 12.10 to 12.12. In all these graphs, the residuals are distributedsymmetrically around zero except that in Figure 12.12, there are two outliersin the upperright corner. These extremelylarge residualsare from people with
exceptionallyhigh values of triglyceride. A large number of the residuals equalzero or are very close to zero, particularly those for AGE and LACR,suggestingthat the model is accuratein predicting the risk of developing CVDfor these people.
We also fit several parametric models to the data. Table 12.9 gives the
goodnessoffitassessmentsforfiveparametricmodels.Thelikelihoodratiotestresults suggest that the Weilbull regression model provides an adequate fit(p/p580.2534 ). The Weilbull fit also gives the largest BIC and AIC values,
suggesting that the Weilbull fit is best among these five models. As mentionedearlier, the Weilbull model is a proportional hazards model. Thus, theparametric model fitting provides additional evidence that the proportionalhazards model is adequate.
Usingthedataset‘‘C: /p33EX12d4d1.DAT’’inExample12.4,thefollowingSAS
code is used to obtain the Cox —Snell, deviance, and weighted Schoenfeld
residuals for AGE, LTG, and LACR in Example 12.10. 333
Figure12.9 DevianceresidualsfromthefittedCoxproportionalhazardsmodelonthe
CVD data.
data w1;
infile ‘c: /p33ex12d4d1.dat’ missover;
input t cens agea ageb sex smoke bmi lacr sbp ltg age htn dm;
run;proc phreg data /p58w1 noprint;
model t*cens (0)/p58age sex lacr ltg / ties /p58efron;
output out /p58out1 logsurv /p58ls resdev /p58rdev wtressch /p58rage r2 rlacr rltg;
run;
data out1;
set out1;rcs/p58-ls;
run;proc lifetest data /p58out1 notable outs /p58ws noprint;
time rcs*cens (0);
run;data ws;
set ws;mls/p58-log(survival );
run;
title ‘Cox-Snell Residuals (rcs)and -log (estimated survival function of rcs )(mls)’;
proc print data /p58ws;
var rcs mls;334
Figure12.10 Weighted Schoenfeld residuals from the fitted Cox proportional hazards
model on the CVD data.
Figure12.11 Weighted Schoenfeld residuals from the fitted Cox proportional hazards
model on the CVD data. 335
Figure12.12 Weighted Schoenfeld residuals from the fitted Cox proportional hazards
model on the CVD data.
run;
title ‘Deviance residuals (rdev)and weighted Schoenfeld residuals for AGE, LACR and
LTG’;proc print data /p58out1;
var t age lacr ltg rage rlacr rltg rdev;
run;
The following SPSS code can be used to obtain Cox —Snell and Schoenfeld
residuals for AGE, and LACR and LTG.
data list file /p58‘c:/p33ex12d4d1.dat’ free
/ t cens agea ageb sex smoke bmi lacr sbpltg age htn dm
coxreg t with age sex lacr ltg
/status /p58cens event (1)
/print /p58all
/save /p58hazard resid presid.
BibliographicalRemarks
An excellent expository paper on statistical methods for the identification and336
Table12.9 Goodness-of-FitTestsBasedonAsymptoticLikelihoodInferenceinFitting
theCVDData /p63
Model LL LLR p BIC AIC
Generalized
gamma /p57198.842 — — /p57217.113 /p57212.842
Log-logistic /p57203.322 — — /p57218.983 /p57215.322
Lognormal /p57206.017 14.3505 0.0002 /p57221.678 /p57218.017
Weibull /p57199.494 1.3046 0.2534 /p57215.155 /p57211.494
Exponential /p57203.061 8.4385 /p640.0147 /p57216.112 /p57213.061
Exponential /p57203.061 7.1339 /p650.0076 /p57216.112 /p57213.061
/p63LL,loglikelihood;LLR,log-likelihoodratio statistic; p,probabilitythatthe respectivechi-square
random variable /p57LLR.
/p64Compared to the generalized gamma fit.
/p65Compared to the Weibull fit.
use of prognostic factors is that of Armitage and Gehan (1974 ). Many studies
of prognostic factors have been published. A few recent ones are cited here:Well et al. (1998 ), Shipley et al. (1999 ), Marrison and Siu (2000 ), Seaman and
Bird (2001 ), Bolard et al. (2001 ), Vasan et al. (2001 ), Young et al. (2001 ),
Meisinger et al. (2002 ), Feskanich et al. (2002 ), Williams et al. (2002 ), and
Bliwise et al. (2002 ).
Cox’s regression model has stimulated the interest of many statisticians. A
large number of papers on this model and related areas have been publishedsince 1972. In addition to the articles cited earlier, the following are a fewexamples: Sasieni (1996 ), Alioum and Commenges (1996 ), Farrington (2000 ),
Vaida and Xu (2000 ), and Zhang and Klein (2001 ). Survival data analysis
methodsarecloselyrelatedtocountingprocesses,particularlytheproportionalhazards model and residual analysis. The counting process approach requiresa strong background in probability theory and stochastic processes and isbeyond the scope of this book. Interested readers are referred to Fleming andHarrington (1991 )and Andersen et al. (1993 ).
EXERCISES
12.1 (a) Consider the data in Exercise Table 3.1. In addition to the five skin
tests, age and gender may also have prognostic values. Examine therelationship between survival and each of seven possible prognosticvariables, as in Table 3.8. For each variable, groupthe p atientsaccording to different cutoff points. Estimate and draw the survivalfunctionforeachsubgroupusingtheproduct-limitmethod andthenuse the methods discussed in Chapter 5 to compare the survival 337
distribution of the subgroups. Prepare a table similar to Table 3.8.
Interpretyourresults.Isthereasubgroupofanyvariablethatshowssignificantly longer survival times? (For the skin test results, use the
larger diameter of the two. )
(b)Considerthe seven variables in part (a). Use Cox’smodel to identify
the most significant variables. Compare your results with thoseobtained in part (a).
12.2 (a) Consider the data givenin Exercise Table3.3. Examine the relation-
shipbetween remission duration and survival time for each of thenine possible prognostic variables: age, gender, family history ofmelanoma, and the six skin tests. Groupthe p atients according todifferent cutoff points. Estimate and draw remission and survivalcurves for each subgroup. Compare remission and survival distribu-tionofsubgroupsusingthemethodsdiscussedin Chapter5.Preparetables similar to Table 3.8.
(b)Use Cox’s regression model to identify the significant variables in
part (a)for their relative importance to remission duration and
survivaltime.Checktheappropriatenessoftheproportionalhazardsmodel using the significant variables identified and the stratifiedanalysis and weighted Shoenfeld residuals. Interpret the results.
12.3Use the proportional hazards model to identify the most important
factors related to survival time in the 157 diabetic patients in ExerciseTable3.4.Checktheappropriatenessof themodelusing allthe methodsdiscussed in Section 12.4 and interpret the results.
12.4 (a) Construct a table similar to Table 3.8 using the data given in Table
3.6.
(b)Use the proportional hazards model to identify the most important
factors related to survival time.
(c)Is the proportional hazards model appropriate for this data set?
12.5Using the data given in Table 12.4, perform similar analyses as in
Examples 12.3 to 12.10 and discuss the results obtained.338
CHAPTER 13
Identification of Prognostic Factors
Related to Survival Time:Nonproportional Hazards Models
In Chapter 12 we discussed the proportional hazards model for the identifica-tion of important prognostic factors, in which the covariates are assumed tobe independent of time. We also assume that there is only one cause of failure;that is, the event or failure is allowed to occur only once for each person, andthere is no correlation among failure times of different persons. However, inpractice,thecovariatesmay beobservedmore thanonceduring thestudy,andtheir values change with time, failure may be due to more than one event orcause, the same event or failure may recur during a follow-upstudy, and theevent or failure time observed may be from related persons in a family or fromthe same person at different times. In this chapter we discuss several modelsfor these situations. The first two models are extensions of the proportionalhazards model to handle time-dependent covariates and to perform stratifiedanalysis. Other models introduced in this chapter are for multiple causes offailure, recurrent events, and related observations.
13.1 MODELSWITHTIME-DEPENDENTCOVARIATES
In the Cox proportional hazards model, the ratio of hazard functions for any
two persons is assumed to be independent of time t, or the covariates are not
time-dependent. However, it is common in practice that a study include bothtime-dependent and time-independent covariates. For example, in a longitudi-nal study of heart disease, certain demographic variables, such as gender andrace, do not change with time and are usually collected only once at thebaseline examination. Other variables, such as lipids, may vary with time andare often collected in subsequent examinations.The partial likelihoodfunctionallowingtime-dependentcovariateshasthesameformasthatin (12.1.7 )except
339
that the covariates are now a function of time. That is, the partial likelihood
function with time-dependent covariates is
L(b)/p58/p73/p147
/p71/p14/p16exp[/p26/p78/p72/p14/p16b/p72x/p72/p7/p71/p8(t/p7/p71/p8)]
/p26l/p43R(t/p7/p71/p8)exp[/p26/p78/p72/p14/p16b/p72x/p72/p74(t/p7/p71/p8)]/p3/p58/p73/p147
/p71/p14/p16exp[b/p30x/p7/p71/p8(t/p7/p71/p8)]
/p26l/p43R(t/p7/p71/p8)exp[b/p30x/p74(t/p7/p71/p8)]/p4(13.1.1 )
wherekisthenumberofdistinctfailuretimes, R(t/p7/p71/p8)istherisksetthatcontains
all persons at risk at time t/p7/p71/p8,x/p74(t/p7/p71/p8)/p58(x/p16/p74(t/p7/p71/p8),x/p17/p74(t/p7/p71/p8),...,x/p78/p74(t/p7/p71/p8))/p30denotes
thecovariatesobservedfromperson lattheordereduncensoredeventtime t/p7/p71/p8,
andb/p30/p58(b/p16,b/p17,...,b/p78)/p30denotes the unknown coefficients. For covariates that
are not time varying, their values are constant over time. For example, let x/p16/p73denotegenderofperson k,thenx/p16/p73(t)/p58x/p16/p73(0)/p58x/p16/p73forallt.Thus,inpractice,
we usually have a mixture of non-time-dependent and time-dependent covari-atesinthelikelihoodfunction.Theestimationprocedureforthecoefficients, b/p72,
is similar to that discussed in Chapter 12. We can also apply the modelselection methods mentioned in Chapter 11 to select the optimal subset ofcovariates as the most important prognostic or risk factors.
There are two kinds of time-dependent covariates: (1)covariates that are
observed repeatedly at different follow-up time points prior to the occurrenceof the event or the end of a study or the censored time; and (2)covariates that
change with time according to a known mathematical function and covariatesthat have different values due to therapy, age, or the changes in medicalconditions.
The following example illustrates how the Cox proportional hazards model
is extended to fit observed survival or event time data with the first kind oftime-dependentcovariates, that is, covariates observed several times before theevent.
Example 13.1 A study was conducted to examine whether biomarker
profiles could be used for risk assessment and bladder cancer detection in acohort of workers occupationally exposed to benzidine and at risk of bladdercancer (Hemstreet et al., 2001 ). These workers were free of bladder cancer at
the time of initial (or baseline )examinationand were reexaminedat least once
based on their risk assessments in a seven-year period. The event timeconsidered in this study is the cancer-free time from baseline examination tolast follow-up. To simplify the analysis, we consider only four covariates: age,level of benzidine exposure, and two biomarkers, M1 and M2. The level ofbenzidine exposure (LEX )is scored based on the worker’s job position in the
factory and is considered fixed (time independent ). In addition, age (AGEB )
and the two biomarkers M1B (/p580 is negative, /p581 if positive )and M2B (/p580
ifnegative, /p581ifpositive )weremeasuredatbaselineexamination (theyarenot
changed with time ). At subsequent examinations, age (AGET )and the two
biomarkers,M1TandM2T,weremeasuredagain (theyarechangedwithtime )
with the status of bladder cancer and the cancer-free time from baselineexamination to subsequent examination (TR). We selected a subset of 61340
persons from this study for this example. The data reproduced in Table 13.1
are solely for the purpose of illustrating the proportional hazards model withtime-dependent covariates. Thus, the results should not be interpreted as thetrue findings of this large study.
Table 13.1 gives the baseline and follow-up data from the 61 participants
selected from the study. We use ID numbers to distinguish the data observedfrom different participants. For example, the person in the table with ID /p584
had LEX /p5836, diagnosed as M1 positive and M2 negative (M1B /p581 and
M2B /p580), and was 47.82 years old (AGEB /p5847.82 )at the baseline examin-
ation (time 0 ). He was diagnosed with negative M1 and M2 (M1T /p580 and
M2T /p580)and without cancer at 42.94 months (TR/p5842.94 )from the baseline
examination and at 51.39 years of age (AGET /p5851.39 )(thus 42.94 was
considereda censored event time, CS /p580). His third examinationwasconduc-
ted at 67.06 months (TR/p5867.06 )and he was still cancer free with both M1
and M2 negative (M1T /p580 and M2T /p580)at 53.40 years old (AGET /p5853.40 )
(thus 67.06 was considered a censored event time, CS /p580). In other words, for
this person, AGET /p5851.39, M1T /p580, and M2T /p580 during the time interval
(0,42.94]andAGET /p5853.40,M1T /p580,andM2T /p580duringthetimeinterval
(42.94, 67.06]. The event time TR was censored at the end of the first time
interval (TR/p5842.94 months, CS /p580)and also at the end of the second time
interval (TR/p5867.06 months, CS /p580). The left endpoint of a time interval is
denoted as TL in the table. Thus, in this example, covariates LEX, AGEB,M1B, and M2Bare fixedfor all time intervals, but AGET,M1T, and M2T aretime-dependent covariates, which may change from one interval to another.
To facilitate better understanding of (13.1.1. ), we use only the data from the
first six people to illustrate how to construct the likelihood function (13.1.1. ).
If we have only the data fromthe first six people,thereare two ( k/p582) distinct
uncensored cancer-free times, t/p7/p16/p8/p5814.65(observed from the person with
ID/p582)andt/p7/p17/p8/p5824.61(observed from the persons with ID /p581).A tt/p7/p16/p8, all
six people are at risk and R(t/p7/p16/p8) contains all six. At time t/p7/p17/p8, only four people
(ID/p581,3,4,and6 )areatriskand R(t/p7/p17/p8) containsthesefour.Thepersonwith
ID/p585 is censored at 14.78 months, prior to t/p7/p17/p8. Table 13.2 gives those in the
risk sets fort/p7/p16/p8andt/p7/p17/p8with values of the seven covariates.
Let
x/p74(t/p7/p71/p8)/p58(LEX/p74, AGEB/p74, M1B/p74, M2B/p74, AGET/p74(t/p7/p71/p8), M1T/p74(t/p7/p71/p8), M2T/p74(t/p7/p71/p8))/p30
denote the covariates from person l(ID/p58l)evaluated at the ordered uncen-
soredeventtime t/p7/p71/p8,i/p581,2;l/p581,2,...,6, andb/p30/p58(b/p16,b/p17,...,b/p22)/p30denotethe
unknown coefficients. Then the first term in (13.1.1 )fort/p7/p16/p8is
exp[b/p30x/p17(t/p7/p16/p8)]
/p26/p21/p74/p14/p16exp[b/p30x/p74(t/p7/p16/p8)] - 341
Table13.1 Cancer-FreeTimesforWorkersExposedtoSomeChemicalElements /p63
ID LEX AGEB M1B M2B AGET M1T M2T CS TR TL
1 180 58.64 0 0 60.70 1 0 1 24.61 0.00
2 69 40.99 0 0 42.21 1 0 1 14.65 0.00
3 36 57.14 0 0 60.72 0 0 0 42.97 0.00
4 36 47.82 1 0 51.39 0 0 0 42.94 0.004 36 47.82 1 0 53.40; 0 0 0 67.06 42.94
5 36 34.85 1 0 36.08 0 0 0 14.78 0.00
6 15 64.24 0 0 67.66 0 1 0 41.03 0.007 15 60.72 0 0 64.14 0 0 0 41.00 0.00
8 15 58.97 0 0 61.54 0 0 0 30.82 0.00
8 15 58.97 0 0 62.01 1 0 0 36.40 30.828 15 58.97 0 0 62.41 1 0 0 41.26 36.40
8 15 58.97 0 0 63.00 0 0 0 48.33 41.26
8 15 58.97 0 0 63.54 0 0 0 54.83 48.338 15 58.97 0 0 64.03 1 0 0 60.71 54.83
8 15 58.97 0 0 64.49 0 0 0 66.17 60.71
8 15 58.97 0 0 65.06 1 0 0 73.07 66.179 15 49.95 0 0 49.95 0 0 0 41.00 0.00
10 15 69.19 0 0 72.61 0 0 0 41.03 0.00
11 15 48.98 0 0 52.41 0 0 0 41.20 0.0012 15 65.52 0 0 68.95 0 0 0 41.17 0.00
13 15 47.86 0 0 47.86 0 0 0 41.43 0.00
14 15 47.82 0; 0 51.28 0 1 0 41.43 0.0014 15 47.82 0 0 52.41 0 0 0 54.97 41.43
15 15 43.49 1 0 46.53 0 0 0 36.50 0.00
15 15 43.49 1 0 46.94 0 0 0 41.43 36.5015 15 43.49 1 0 47.53 0 0 0 48.46 41.43
15 15 43.49 1 0 48.56 0 0 0 60.85 48.46
16 15 41.28 0 0 44.74 0 1 0 41.56 0.0016 15 41.28 0 0 45.86 0 0 0 54.93 41.56
17 15 49.09 0 0 52.54 0 0 0 41.43 0.00
18 15 46.03 0 0 49.45 0 0 0 41.03 0.0019 15 64.41 0 0 67.85 0 0 0 41.23 0.00
20 164 52.52 0 0 53.54 1 1 1 12.32 0.00
21 15 61.51 0 0 64.94 0 0 0 41.10 0.0022 144 64.59 0 0 68.01 1 0 0 41.13 0.00
22 144 64.59 0 0 68.60 1 1 1 48.16 41.13
23 192 62.26 0 1 64.88 1 0 0 31.47 0.0023 192 62.26 0 1 65.27 0 0 1 36.17 31.47
24 54 57.56 0 0 57.95 1 0 1 4.67 0.00
25 264 60.03 0 0 73.03 1 0 0 36.07 0.0025 264 60.03 0 0 73.71 1 0 0 44.19 36.07
25 264 60.03 0 0 64.17 1 0 0 49.68 44.19
25 264 60.03 0 0 65.15 1 0 1 61.44 49.6826 40 44.30 0 0 45.48 1 0 0 14.13 0.00
26 40 44.30 0 0 46.49 1 0 1 26.25 14.13
27 265 52.84 0 1 53.98 0 0 0 13.73 0.0027 265 52.84 0 1 55.43 0 0 0 31.18 13.73
27 265 52.84 0 1 55.83 0 0 0 35.91 31.18
27 265 52.84 0 1 56.42 0 0 1 42.97 35.91342
Table13.1Continued
ID LEX AGEB M1B M2B AGET M1T M2T CS TR TL
28 132 68.19 0 1 69.31 0 0 1 13.50 0.00
29 24 62.22 1 1 64.39 0 0 0 26.02 0.00
29 24 62.22 1 1 64.85 0 0 0 31.54 26.02
29 24 62.22 1 1 65.22 0 0 0 36.01 31.5429 24 62.22 1 1 65.82 0 1 0 43.27 36.01
29 24 62.22 0 1 66.89 1 0 1 56.02 43.27
30 132 68.27 0 0 70.12 0 0 1 22.14 0.0031 178 64.07 0 0 64.07 1 0 1 21.95 0.00
32 50 65.88 0 0 65.88 0 0 0 25.43 0.00
33 50 70.82 0 1 74.40 0 0 0 42.97 0.0034 50 60.53 0 1 63.54 0 0 0 36.14 0.00
34 50 60.53 0 1 64.67 0 0 0 49.68 36.14
34 50 60.53 0 1 66.18 0 0 0 67.88 49.6835 50 62.99 0 0 66.00 0 0 0 36.11 0.00
36 50 63.01 1 1 65.15 0 0 0 25.76 0.00
36 50 63.01 1 1 66.01 0 0 0 36.04 25.7636 50 63.01 1 1 66.60 0 0 0 43.07 36.04
36 50 63.01 1 1 67.68 1 0 0 56.05 43.07
37 50 63.86 0 0 66.89 0 0 0 36.40 0.0038 50 61.15 0 0 62.33 0 0 0 14.16 0.00
38 50 61.15 0 0 63.32 0 0 0 26.02 14.16
38 50 61.15 0 0 63.78 0 0 0 31.57 26.0238 50 61.15 0 0 64.75 0 0 0 43.20 31.57
38 50 61.15 0 0 65.30 1 0 0 49.87 43.20
39 50 61.02 0 0 64.03 1 0 0 36.14 0.0040 50 61.08 0 0 61.08 0 0 0 36.17 0.00
41 50 49.50 0 1 52.51 0 0 1 36.14 0.00
42 50 49.81 0 0 52.81 0 0 0 35.94 0.0043 50 49.09 0 0 52.10 0 0 0 36.17 0.00
44 50 47.07 0 0 50.08 0 0 0 36.14 0.00
45 50 63.69 0 1 64.84 0 0 0 13.90 0.0045 50 63.69 0 1 66.30 0 0 0 31.41 13.90
45 50 63.69 0 1 66.69 0 0 0 36.01 31.41
45 50 63.69 0 1 67.28 0 0 0 43.10 36.0146 50 55.77 0 0 58.77 0 0 0 36.01 0.00
47 50 60.84 0 1 61.99 1 0 0 13.83 0.00
47 50 60.84 0 1 64.98 0 0 0 49.71 13.8348 50 50.09 1 1 51.24 1 0 0 13.90 0.00
48 50 50.09 1 1 52.70 0 0 0 31.41 13.90
48 50 50.09 1 1 53.09 0 0 0 36.01 31.4148 50 50.09 1 1 54.23 0 0 0 49.77 36.01
48 50 50.09 1 1 54.76 0 0 0 56.05 49.77
48 50 50.09 1 1 55.75 0 0 0 67.98 56.0549 50 62.41 1 0 63.53 0 0 0 13.50 0.00
49 50 62.41 1 0 65.38 0 0 0 35.61 13.50
50 50 73.88 0 1 78,03 0 0 0 49.81 0.0051 50 44.68 0 0 47.68 1 0 0 35.98 0.00
51 50 44.68 0 0 49.36 0 0 0 56.12 35.98
52 50 62.67 0 0 65.66 0 0 0 35.91 0.00
(Continued overleaf ) - 343
Table13.1 Continued
ID LEX AGEB M1B M2B AGET M1T M2T CS TR TL
53 275 74.28 0 1 75.34 1 1 0 12.75 0.00
53 275 74.28 0 1 77.28 0 0 0 36.04 12.75
53 275 74.28 0 1 77.70 1 0 0 41.07 36.04
53 275 74.28 0 1 78.26 1 0 1 47.80 41.0754 57 39.52 0 0 43.50 0 0 0 47.80 0.00
5 57 76.22 1 0 79.23 1 0 0 36.07 0.00
5 57 76.22 1 0 79.64 0 0 0 41.10 36.075 57 76.22 1 0 80.21 0 0 0 47.84 41.10
5 57 76.22 1 0 81.24 0 0 0 60.29 47.84
56 57 62.41 0 0 65.83 0 0 0 41.10 0.0057 57 67.64 0 0 71.06 0 0 0 41.10 0.00
58 57 80.61 0 0 84.03 0 1 0 41.10 0.00
58 57 80.61 0 0 85.14 0 0 0 54.37 41.1059 57 67.78 1 0 67.68 1 0 0 72.12 0.00
60 0 47.35 0 1 47.35 0 1 0 13.83 0.00
60 0 47.35 0 1 49.84 0 0 0 43.70 13.8361 0 40.98 1 0 42.13 0 0 0 13.83 0.00
61 0 40.98 1 0 43.59 0 0 0 31.38 13.83
61 0 40.98 1 0 44.62 0 0 0 43.70 31.3861 0 40.98 1 0 46.55 0 0 0 66.92 43.70
/p63ID,participantIDnumber;LEX,levelofexposure;AGEB,ageatthebaselineexamimation;M1B
and M2B, index functions of measure 1 and 2 at the baseline; M1B /p581 if measure 1 is positive
and 0 if not; M2B /p581 if measure 2 is positive and 0 if not; AGET, age at the end of each time
interval; M1T and M2T, index functions of measure 1 and 2 at the end of each time interval;M1T /p581ifmeasure1ispositiveand0ifnot;M2T /p581ifmeasure2ispositiveand0ifnot;CS /p580
if censored and 1 if not; TR, cancer-free time in months (or the right endpoint of time interval );
TL, left endpoint of time interval.
where
x/p17(t/p7/p16/p8)/p58(69, 40.99, 0, 0, 42.21, 1, 0) /p30
is the column vector of covariates from person 2, whose cancer-free time is
t/p7/p16/p8/p5814.65.Thex/p74(t/p7/p16/p8)’sinthedenominatorarethecovariatevectorsobserved
for the six people in the risk set R(t/p7/p16/p8) and are listed in Table 13.2. For
example,x/p18(t/p7/p16/p8)/p58(36, 57.14, 0, 0, 60.72, 0, 0).The second and also the last
term in (13.1.1 )fort/p7/p17/p8/p5824.61 is
exp[b/p30x/p16(t/p7/p17/p8)]
exp[b/p30x/p16(t/p7/p17/p8)]/p59exp[b/p30x/p18(t/p7/p17/p8)]/p59exp[b/p30x/p19(t/p7/p17/p8)]/p59exp[b/p30x/p21(t/p7/p17/p8)]
wherex/p16(t/p7/p17/p8)/p58(180, 58.64, 0, 0, 60.70, 1, 0) and x/p74(t/p7/p17/p8)’s in the denominator
are the observed covariate vectors from the four persons (ID/p581, 3, 4, and 6 )344
Table 13.2 Construction of the Partial Likelihood for the Cancer-Free Times from the First Six People with Time-Dependent
Covariates
Ordered t/p7/p16/p8/p5814.65 t/p7/p17/p8/p5824.61
Event time (observed from individual with ID /p582)( observed from individual with ID /p581)
ID LEX AGEB M1B M2B AGET M1T M2T LEX AGEB M1B M2B AGET M1T M2T
1 180 58.64 0 0 60.70 1 0 180 58.64 0 0 60.70 1 0
2 69 40.99 0 0 42.21 1 0
3 36 57.14 0 0 60.72 0 0 36 57.14 0 0 60.72 0 04 36 47.82 1 0 51.39 0 0 36 47.82 1 0 51.39 0 05 36 34.85 1 0 36.08 0 06 15 64.24 0 0 67.66 0 1 15 64.24 0 0 67.66 0 1
345
Table13.3 AsymptoticPartialLikelihoodInferenceonCancer-FreeTime Datafrom
FittedModelwithTime-DependentCovariates
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
LEX 0.007 0.003 5.593 0.018 1.01 1.00 1.01
M1T 1.361 0.645 4.449 0.035 3.90 1.10 13.81
inR(t/p7/p17/p8)(see Table 13.2 for details ). The partial likelihood function for this
reduced data set is the product of these two terms.
The partial likelihood function for the entire data set in Table 13.1 can be
constructed in a similar way and estimates of the coefficients can be obtainedusing the Newton —Raphson method. The data format style in Table 13.1 is
referred to as a counting process data format. The results of fitting this modelwith time-dependent covariates and a stepwise selection method are given inTable 13.3. The coefficients indicate that high levels of exposure and positiveM1 at follow-upexamination are p ositively related to the risk of a shortcancer-free time. Assuming that other measures are the same, a person with apositiveM1atfollow-upexaminationwillhave3.9timeshigherrisktodevelopbladder cancer than will someone with a negative M1. For every 1-unitincrease in LEX, the risk will increase by 1%.
Suppose that the text data file ‘‘C: /p33EX1311.DAT’’ contains the data in
Table 13.1 and the successive 11 columns give ID, LEX, AGEB, M1B, M2B,AGET, M1T, M2T, CS, TR, and TL. The following SAS code can be used toobtain the results in Table 13.3.
data w1;
infile ‘c: /p33ex13d1d1.dat’ missover;
input id lex ageb m1b m2b aget m1t m2t cs tr tl;
run;
title ‘‘Selected Cox proportional hazards model with time dependent covariates’’;proc phreg data /p58w1;
model (tl,tr)*cs(0)/p58lex ageb m1b m2b aget m1t m2t / rl ties /p58efron selection /p58s;
where tl /p58tr;
run;
For the second type of time-dependent covariate (i.e., the covariate known
to change with time according to a mathematical function ), we simply use the
known mathematical function to replace the covariate. Following is a hypo-thetical example to illustrate the use of SAS, SPSS, and BMDP.346
Example 13.2 Suppose that we wish to fit the proportional hazards model
to a set of survival data that has been saved in a text file ‘‘C: /p33EX1312.DAT’’.
This set of data consists of survival time t, an indicator variable CENS
(/p581 for an uncensored observation and 0 for a censored observation )and
three covariates, X1, X2, and X3. Furthermore, assume that X3 is known tochange with time according to the function X3 /p59log(t/p591). In this case, the
following SAS, SPSS, and BMDP code can be used to incorporate thistime-dependent covariate with known mathematical relationship with time
into the model.
data w1;
infile ‘c: /p33ex1312.dat’ missover;
input t cens x1 x2 x3;
run;proc phreg data /p58w1;
model t*cens (0)/p58x1 x2 z/ rl ties /p58efron;
z/p58x3*log (t/p591)
run;
If the SPSS COXREG procedure is used, the code is
data list file /p58‘c:/p33ex1312.dat’ free
/ t cens x1 x2 x3.
time program
Compute z /p58x3*log (t/p591).
coxreg t with x1 x2 z
/status /p58cens event (1)
/print /p58all
/save /p58hazard resid presid.
For the BMDP 2L procedure, the code is
/input file /p58‘c:/p33ex1312.dat’ .
variables /p586.
format /p58free.
/print cova./variable names /p58t,cens, x1, x2, x3.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58x1, x2, z.
add/p58z.
/function z /p58x3*ln (time/p591). - 347
13.2 STRATIFIEDPROPORTIONALHAZARDSMODELS
The proportional hazards model in (12.1.3 )assumes that the ratio of the
hazard functions of any two people with prognostic variables x/p16andx/p17is a
constant, independent of time. This assumption may not always be met inpractical situations. To accommodate the nonproportional cases, Cox’s modelcanbegeneralizedusingtheconceptofstratification (KalbfleischandPrentice,
1980 ).Thedatacanbestratifiedbyacovariate:forexample,age.Ifweconsider
two strata, say age /p4650 and /p5850 years, the model in (12.1.3 )becomes two
models:
h/p71(t/p34x)/p58h/p15/p71(t) exp
/p1/p26b/p72x/p72/p2/p58h/p15/p71(t) exp (b/p30x/p71)( 13.2.1 )
wherei/p581, 2 for the two age strata. Notice that the underlying hazard
functionh/p15/p71(t) is assumed to be different for the two strata; however, the
regression coefficients are the same for all strata. That is, we assume that thehazards for patients may be proportional within each stratum but not amongdifferent strata (or levels ). The partial (marginal )likelihood function for all
observations from the mstrata is defined as
L(b)/p58/p75/p147
/p72/p14/p16L/p72(b)( 13.2.2 )
whereL/p72(b)is the partial (marginal )likelihood function for the jth stratum.
The regression coefficients bcan be estimated by the Newton —Raphson
method. For stratified models, the baseline survivorshipfunction for eachstratum is estimated separately based on the estimated regression coefficientsb/p19andthedatain thatstratumalonebyusingthemethodsdiscussedin Section
12.3.
Example 13.3 Consider the data givenin Example12.1.1. Suppose that we
are not sure if the risk of dying for patients at least 50 years of age isproportionaltothatforpatientslessthan50yearsanddecidetodoastratifiedanalysis. Two regression equations are therefore assumed:
h/p16(t/p34x/p17)/p58h/p15/p16(t) exp (b/p17x/p17)
h/p17(t/p34x/p17)/p58h/p15/p17(t) exp(b/p17x/p17)
whereh/p16(t/p34x/p17), the hazard function for patients under 50 years of age, and
h/p17(t/p34x/p17), the hazard function for patients at least 50 years, are functions of
cellularity,and h/p15/p16(t) andh/p15/p17(t) aretheunderlyinghazardfunctionsforthetwo
groups.Theresultsofthestratifiedanalysis, b/p19/p17/p580.22,SE (b/p19/p17)/p580.44,p/p580.31,
andexp (b/p19/p17)/p581.24,areclosetothoseobtainedearlierintheunstratifiedmodel.348
Table13.4 AsymptoticPartialLikelihoodInferenceonCVD-freeTimeDatafrom
FittedModels
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
UnstratifiedModelforAllCVDs
AGE 0.070 0.014 26.139 0.0001 1.07 1.04 1.10
SEX 0.753 0.219 11.789 0.0006 2.12 1.38 3.26
LACR 0.111 0.046 5.860 0.0155 1.12 1.02 1.22
LTG 0.399 0.198 4.072 0.0436 1.49 1.01 2.20
Gender-StratifiedModelforAllCVDs
AGE 0.063 0.013 23.430 0.0001 1.07 1.04 1.09
LACR 0.149 0.043 11.828 0.0006 1.16 1.07 1.26
Gender-SpecificProportionalHazardsModels
Female
SBP 0.022 0.006 13.962 0.0002 1.02 1.01 1.03
DM 0.986 0.373 6.973 0.0083 2.68 1.29 5.57
Male
AGE 0.069 0.018 14.453 0.0001 1.07 1.03 1.11
LACR 0.125 0.058 4.555 0.0328 1.13 1.01 1.27
However, this may not always be the case. Because the model is stratified by
age groupand no sp ecific relationshipis assumed between the hazard ratio ofpatients at least 50 years old and those under 50, tests of significance of theregression coefficients for the other variables are adjusted for age.
Example 13.4 In Example 12.5 we used the stepwise selection method to
fit the proportional hazards model to the CVD data in Example 12.3. Wereanalyze the data using a gender-stratified proportional hazards model.Results from the unstratified model and the stratified model (with a stepwise
selection procedure )are given in Table 13.4.
The unstratified model identifies AGE, SEX, LACR (logarithm of the ratio
of urinary albumin and creatinine ), and LTG (logarithm of triglycerides )as
significant covariates for the time to CVD. The gender-stratified model withthe stepwise selection procedure identifies AGE and LACR as the mostsignificant covariates. The coefficients are close to those obtained in theunstratified model. The log[-log (S(t))] at the averages of covariates AGE and
LACR for the two strata are plotted in Figure 12.7. The two curves look 349
paralleltoeachother.Itsuggeststhatstratificationforthissetofdatadoesnot
provide more information for the study. Moreover, the sex-specific propor-tional hazards model (at the bottom of Table 13.4 )show that systolic blood
pressure (SBP )and diabetes are significant covariates related to the risk of
CVDinwomenandAGEandLACRin men.Thus,thegender-specificmodelsprovide more information and suggest that there are differences in CVD riskfactors among men and women.
The method of stratification is useful in cases when the observations from
different strata are considered independent, conditional on the stratifiedvariable, or one is not interested in the effect of the stratified variable itself onthe outcome but in the interactions of the stratified variable with the othercovariatesin the model and does not knowthe exact forms of the interactions.It is clear that modeling observations from different strata separately canprovide more information than either stratification or unstratification if thesample size in each stratum is large enough.
Using the data file ‘‘C: /p33EX12d4d1.DAT’’ defined in Example 12.4, the
following SAS code can be used to obtain the results in Table 13.4 andFigure 12.7.
data w1;
infile ‘c: /p33ex12d4d1.dat’ missover;
input t cens agea ageb sex smoke bmi lacr sbp ltg age htn dm;
run;title ‘‘Unstratified model’’;proc phreg data /p58w1;
model t*cens (0)/p58age sex bmi lacr sbpltg smoke htn dm / selection /p58b
ties/p58efron;
run;title ‘‘gender stratified model’’;proc phreg data /p58w1;
model t*cens (0)/p58age bmi lacr sbpltg smoke htn dm / selection /p58b
ties/p58efron;
strata sex;
run;proc phreg data /p58w1 noprint;
model t*cens (0)/p58age lacr / ties /p58efron;
strata sex;baseline out /p58bas1 loglogs /p58lmls;
run;
title ‘‘Log-logS from fitting a gender stratified model’’;proc print data /p58bas1;
var sex age lacr t lmls;
run;proc sort data /p58w1;
by sex;
run;title ‘‘gender-specific models’’;350
proc phreg data /p58w1;
model t*cens (0)/p58age bmi lacr sbpltg smoke htn dm / selection /p58b
ties/p58efron;
by sex;
run;
The followingSPSS code can also be used. In this case, the data for women
and men are assumed to be in the files ‘‘C: /p33EX12d4d1a.DAT’’ and
‘‘C:/p33EX12d4d1b.DAT’’ separately.
data list file /p58‘‘c:/p33ex12d4d1.dat’’ free
/ t cens agea ageb sex smoke bmi lacr sbpltg age htn dm.
coxreg t with age sex bmi lacr sbpltg smoke htn dm
/status /p58cens event (1)
/method /p58bstepage sex bmi lacr sbpltg smoke htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all
coxreg t with age bmi lacr sbpltg smoke htn dm
/status /p58cens event (1)
/strata /p58sex
/method /p58bstepage bmi lacr sbpltg smoke htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all
coxreg t with age lacr
/status /p58cens event (1)
/strata /p58sex
/print /p58all
/save /p58lml.
data list file /p58‘‘c:/p33ex12d4d1a.dat’’ free
/ t cens agea ageb sex smoke bmi lacr sbpltg age htn dm.
coxreg t with age bmi lacr sbpltg smoke htn dm
/status /p58cens event (1)
/method /p58bstepage bmi lacr sbpltg smoke htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all
data list file /p58‘‘c:/p33ex12d4d1b.dat’’ free
/ t cens agea ageb sex smoke bmi lacr sbpltg age htn dm.
coxreg t with age bmi lacr sbpltg smoke htn dm
/status /p58cens event (1)
/method /p58bstepage bmi lacr sbpltg smoke htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all
For the BMDP 2L procedure, the following code can be used.
/input file /p58‘c:/p33ex12d4d1.dat’ .
variables /p5813.
format /p58free. 351
/print cova.
Survival.
/variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg,
age, htn, dm.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58age, smoke, bmi, lacr, sbp, ltg, htn, dm.
strata /p58sex.
step/p58phh.
/input file /p58‘c:/p33ex12d4d1a.dat’ .
variables /p5813.
format /p58free.
/print cova.
Survival.
/variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg,
age, htn, dm.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58age, smoke, bmi, lacr, sbp, ltg, htn, dm.
step/p58phh.
/input file /p58‘c:/p33ex12d4d1b.dat’ .
variables /p5813.
format /p58free.
/print cova.
Survival.
/variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg,
age, htn, dm.
/form time /p58t.
status /p58cens.
response /p581.
/regress covariates /p58age, smoke, bmi, lacr, sbp, ltg, htn, dm.
step/p58phh.
13.3 COMPETINGRISKSMODEL
All the methods for prognostic factor analysis discussed so far deal with a
single type of failure time for each study subject. This may be a perfectlyacceptable way to proceed in many cases. However, in some situations, failureon an person may be due to several distinct causes. It may be desirable todistinguish different kinds of events that may lead to failure and treat themdifferently in the analysis. For example, to evaluate the efficacy of hearttransplants, one would certainly want to treat deaths due to heart failuredifferently from deaths due to other causes, such as accident and cancer. In a352
mortality study, it may be more interesting to study separately deaths due to
heartdisease,diabetes,cancer,andothersthantocombineallthecauses.Thesedifferent causes of failure are considered as competing events, which introducecompeting risks. Thus, problems arising in the analysis of data with multiplecauses are commonly referred to as competingrisk problems. We will see later
that competing risk analysis, in general, requires no inference methods otherthan those introduced in Chapters 11 and 12. We focus on using theproportional hazards model to identify significant prognostic or risk factors
when competing risks are present. Readers interested in additional details arereferred to Kalbfleisch and Prentice (1980 ).
LetTbe the survival time, xthe covariate vector, and Jthe type or cause
of failure. We define a type- or cause-specific hazard function h/p72(t;x)
h/p72(t;x)/p58lim
/p9/p82/p29/p15P(t/p45T/p58t/p59/afii9773t,J/p58j/p34T/p46t,x)
/afii9773t
,j/p581,...,m (13.3.1 )
In words,h/p72(t;x)is the instantaneous failure rate of cause jat timetgivenx
and in the presence of other (m/p571)causes of failure. The only difference
between (13.3.1 )and the hazard function defined in Chapter 2 is the appear-
ance ofJ/p58j. Equation (13.3.1 )is a type- or cause-specific hazard function,
which is very much the same as the ordinary hazard function except that theevent is of a specific type. The overall hazard of failure is the sum of all thetype-specific hazards, that is,
h(t;x)/p58/p26
/p72h/p72(t;x)( 13.3.2 )
provided that the failure types are mutually excluded. Based on (2.15), we can
define the function
.
S/p72(t;x)/p58exp
/p3/p57/p16/p82
/p15h/p72(u;x)du/p4,j/p581,...,m(13.3.3)
However, these functions cannot, in general, be interpreted as survivorship
functions when m/p571. Let t/p72/p16/p58t/p72/p17/p58/p37/p58t/p72/p73/p72denote the failure times for
failures of type j,j/p581,...,m. Assuming proportional hazards, the hazard
function in (13.3.1 )can be written as
h/p72(t;x)/p58h/p15/p72(t) exp(b/p30/p72x),j/p581,...,m (13.3.4 )
which can be generalized for time-dependent covariates by replacing xwith
x(t), that is,
h/p72(t;x)/p58h/p15/p72(t) exp[b/p30/p72x(t)],j/p581,...,m (13.3.5 ) 353
The partial likelihood function for the model in (13.3.5 )is
L/p58/p75/p147
/p72/p14/p16/p73/p72/p147
/p71/p14/p16exp[b/p30/p72x/p72/p71(t/p72/p71)]
/p26l/p43R(t/p72/p71)exp[b/p30/p72x/p74(t/p72/p71)](13.3.6 )
whereR(t/p72/p71)is the risk set at t/p72/p71. The estimation of the coefficients and
identification of significant covariates can be carried out exactly the same way
asdescribedinChapters11and12bytreatingfailuretimesoftypesotherthanjas censored observations. This is perhaps the most important concept in
competing risks analysis. It is because the basic assumption for a competingrisksmodelisthattheoccurrenceofonetypeofeventremovesthepersonfromrisk of all other types of events and the person will no longer contributeto thesuccessiveriskset. Furthermore,thereis nothingtopreventone fromchoosingdifferent types of models for different h/p72(t;x)’s. For example, in a mortality
study we might choose a proportional hazards model for heart disease and aparametric model for diabetes.
The coefficient vector b/p72in(13.3.6 )indicates the effects of the covariates for
eventtypej. Ifany covariatesarenotrelatedtoaparticulartypeorcause, they
may be set to 0. If b/p72are the same for all j, the model in (13.3.5 )reduces to the
proportional hazards model in Chapter 12. The following example illustratesthe proportional hazards model with competing risks.
Example 13.5 Let us again use the CVD data in Example 12.3. The event
typesarenon-CVD (DG/p580),stroke (DG/p581),CHD (DG/p582),andtheother
CVDs (DG/p583). If one is interestedin all CVD nomatter whether it isstroke,
CHD,ortheotherCVDs,thecompetingrisksmodelreducestoageneralCVDevent model, the times (T)to CVD for DG /p581, 2, 3 are uncensored event
times, and the other times are censored (DG/p580). An indicator variable, CS,
can be used to indicate the censoring status; that is, CS /p581i fD G /p581, 2, 3,
and CS /p580 otherwise. The result from fitting the proportional hazards model
with the backward selection method is given in section (a)of Table 13.5.
If one considers strokes only, the indicator variable CS has to be defined
differently; that is, CS /p581i fD G /p581 and CS /p580i fD G /p580, 2 and 3. This
means that in addition to non-CVD, the event time of CHD and the otherCVDsaretreatedas censoredobservations.Notethatwewillremoveapersonfrom the risk set after his or her first CVD event time in constructing thelikelihoodfunctionforthemodifieddataeveniftheeventwasnotafatalevent.For stroke, age is the only significant variable [section (b)in Table 13.5]. We
call this model a marginal model for strokes. Similarly, if only CHD, otherCVDs,or either stroke or CHD are of interest,the respective modificationwillbe CS /p581i fD G /p582 and CS /p580 otherwise (CHD only );C S/p581i fD G /p583
and CS /p580 otherwise (other CVD only );o rC S /p581i fD G /p581, 2 and CS /p580
otherwise (either stroke or CHD ). The results of these three fits with the
backwardselectionmethodareshowninTable13.5 (c)—(e).Theresultssuggest354
Table13.5 AsymptoticPartialLikelihoodInferenceonCVDEventTimeDatafrom
theFittedCompetingRisksModels
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
(a) Model for All CVDs
AGE 0.070 0.014 26.139 0.0001 1.07 1.04 1.10
SEX 0.753 0.219 11.789 0.0006 2.12 1.38 3.26
LACR 0.111 0.046 5.860 0.0155 1.12 1.02 1.22
LTG 0.399 0.198 4.072 0.0436 1.49 1.01 2.20
(b) Marginal Model for Strokes
AGE 0.072 0.021 12.092 0.0005 1.08 1.03 1.12
(c) Marginal Model for CHDs
AGE 0.069 0.020 11.622 0.0007 1.07 1.03 1.12
SEX 0.970 0.329 8.716 0.0032 2.64 1.39 5.02
BMI 0.040 0.017 5.162 0.0231 1.04 1.01 1.08
LTG 1.106 0.266 17.234 0.0001 3.02 1.79 5.09
(d) Margial Model for Other CVDs
AGE 0.087 0.033 6.874 0.0087 1.09 1.02 1.17
SEX 1.100 0.555 3.937 0.0472 3.01 1.01 8.91
LACR 0.315 1.101 9.745 0.0018 1.37 1.12 1.67
(e) Marginal Model for Strokes or CHDs
AGE 0.072 0.015 23.555 0.0001 1.07 1.04 1.11
SEX 0.692 0.239 8.362 0.0038 2.00 1.25 3.20LTG 0.665 0.200 11.095 0.0009 1.94 1.32 2.88
that significant risk factors differ for different types of CVD events. Age is the
only factor common to all the CVD events.
Thus, competing risks models provide an opportunity to separate any one
ormore specifictypes ofevent orcauseof deathfrom allothertypes orcauses.In practice, it is not necessary to fit a model to every type or cause.
Suppose that ‘‘C: /p33EX13d3d1.DAT’’ is a text data file that contains 14
columns similar to Table 12.4 and the successive columns give T, CENS, DG,AGEA, AGEB, SEX, SMOKE, BMI, LACR, SBP, LTG, AGE, HTN, andDM. The following SAS code can be used to obtain the model for stroke inTable 13.5. These codes can easily be modified to obtain the results for CHD, 355
other CVD, and stroke/CHD.
data w1;
infile ‘c: /p33ex13d3d1.dat’ missover;
input t cens dg agea ageb sex smoke bmi lacr sbp ltg age htn dm;
run;
title ‘‘Model for stroke event times’’;
proc phreg data /p58w1;
model t*dg (0, 2, 3 )/p58age sex smoke bmi lacr sbpltg htn dm
/ rl selection /p58b ties /p58efron;
run;
The following SPSS code can be used.
data list file /p58‘‘c:/p33ex13d3d1.dat’’ free
/ t cens dg agea ageb sex smoke bmi lacr sbpltg age htn dm.
coxreg t with age sex bmi lacr sbpltg smoke htn dm
/status /p58dg event (1)
/method /p58bstepage sex bmi lacr sbpltg smoke htn dm
/criteria pin (0.05)pout (0.05)
/print /p58all
The following code is for the BMDP 2L procedure.
/input file /p58‘c:/p33ex13d3d1.dat’ .
variables /p5814.
format /p58free.
/print cova.
Survival.
/variable names /p58t,cens, dg, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg,
age, htn, dm.
/form time /p58t.
status /p58dg.
response /p581.
/regress covariates /p58age, sex, smoke, bmi, lacr, sbp, ltg, htn, dm.
step/p58phh.
13.4 RECURRENTEVENTSMODELS
So far we have considered events or failures that are allowed to occur only
once. Even in competing risks models, the occurrence of one type of eventremoves a person from the risk set thereafter. However, in practice the failureson an individual may be recurrences of essentially the same event, such astumor recurrences after surgeries, or may be successive events of entirelydifferent types, such as strokes and heart attacks. When data include recurrentevents, regression models such as the proportional hazards model becomemuch more mathematically complicated and often involve counting process356
theory,whichisbeyondthe scopeof thisbook.Anumber ofregressionmodels
have been proposed in the literature. In this section we introduce three modelsthat can be considered as extensions of the Cox proportional hazards model.We keepthe mathematics to a minimum and use examp les to show how thesemodels can be used to identify important prognostic or risk factors with theaid of available computer software. The three models are based on Prentice etal.(1981 ),AndersenandGill (1982 ),andWeietal. (1989 ).Allthreemodelsare
proportional hazards models, and the likelihood functions of these models are
constructeddifferently,primarilyintherisksetattheuncensoredobservations.Readers interested in details are referred to the papers cited above.
Prentice et al. Model
In their 1981 paper, Prentice, Williams, and Peterson (PWP )proposed two
modelsforrecurrentevents.BothPWPmodelscanbeconsideredasextensionsof the stratified proportionalhazardsmodelwith strata defined by the numberand time of the recurrent events. The hazard function is extended beyond theperson’s first event to cover subsequent events. In the first PWP model,follow-uptimestarts atthebeginningofthestudy (truetime0 )andthehazard
function of the ith person can be written as
h(t/p34b/p81,x/p71(t))/p58h/p15/p81(t) exp[b/p30/p81x/p71(t)] (13.4.1 )
wherethe subscript srepresents thestratumthat thepersonis inat time t. The
first stratum includes people who have at least one recurrence or are censoredwithout recurrence, the second stratum includes people who have at least tworecurrences or are censored after the first recurrence, and so on. A personmoves from stratum 1 (s/p581)to stratum 2 (s/p582)following his or her first
recurrenteventandremainsinstratum2 untilthesecondrecurrenteventtakesplaceor becomesa censorobservation (no more recurrentevent ).Theh/p15/p81(t)in
(13.4.1 )is the stratum-specific underlying hazard. Notice that in (13.4.1 ), the
coefficients are stratum-specific also.
Lett/p81/p16/p58/p37/p58t/p81/p66/p81denote thed/p81ordered distinct failure times in stratum s,
x/p81/p71(t/p81/p71)thecovariatevectorofasubjectinstratum swhofailsattime t/p81/p71,x/p81/p74(t/p81/p71)
the covariate vector of subject lin stratumsat timet/p81/p71, andR(t,s) the set of
persons at risk in stratum sjust prior to time t. Note that the risk set R(t,s)
includes only those persons who have experienced the first s/p571 recurrent
events. Then the partial likelihood for the first model in (13.4.1 )is
L(b)/p58/p147
s/p461/p66/p81/p147
/p71/p14/p16exp[b/p30/p81x/p81/p71(t/p81/p71)]
/p26l/p43R(t/p81/p71,s)exp[b/p30/p81x/p81/p74(t/p81/p71)](13.4.2)
The following example illustrates the construction of the likelihood function
and the necessary data arrangements for using SAS, SPSS, or BMDP to carryout the analysis. 357
Example 13.6 We use the tumor recurrence data from bladder cancer
patients (Andrews and Herzberg, 1985; Wei et al., 1989 )in a clinical trial to
compare three treatments, which was conducted by the Veterans Administra-tion Cooperative Urological Research Group (Byar, 1980 ). All patients had
superficial bladder tumors when they entered the study. These tumors wereremoved and the patients were randomized into three treatment groups:placebo, thiotepa, and pyridoxine. During the follow-up period many patientshad one or more recurrences of tumors and new tumors were removed when
discovered.In this example we use the tumor recurrence data from 86 patientswho received either placebo or thiotepa. Only the first four recurrence timesare considered. The data set, reproduced in Table 13.6, includes treatment(1, placebo; 2, thiotepa )follow-uptime, initial number of tumors (N), initial
tumor size (S)in centimeters, and recurrent time. Each recurrent time of a
patient was measured from the date of first treatment.
In this case, the event of interest is tumor recurrence and the strata are
defined by the number of recurrences (NRs ).To use SAS and other software to
fitdatawiththemodel (13.4.1 ),thedatamustberearrangedinacertainformat
by stratum. To facilitate illustration, we selectsix patients from Table 13.6 andplacethe data of thesesix patientsin Table13.7. The follow-upand recurrencetimes are also shown in Figure 13.1. From the figure we see that stratum 1includes patients 1 (censored at 9 months ),2(censored at 59 months ),3(first
recurrent at 3 months ),4(first recurrent at 12 months ),5(first occurrence at
6 months ), and 6 (first occurrence at 3 months ). The time intervals, (TL, TR],
are(0,9], (0,59], (0,3], (0,12], (0,6], and (0,3], respectively. These intervals are
used to determine the risk set in the stratum-specific likelihood function in(13.4.2 ),andthepatientsinthestratumwereatriskonlyinthesetimeintervals.
To use software packages such as SAS, BMDP, and SPSS, we need torearrangethe data by stratum. Table 13.8 gives the rearranged data. Note thatthe six patients in stratum 1 are arranged in ascending order according to therightendofthetimeinterval.AlsointroducedinthistableareT1 —T4,N1—N4,
and S1—S4, giving the treatment received, initial tumor number, and initial
tumor size of the patients for the four strata, respectively. These variables areset to be zero in the other strata except the stratum they are in. For example,forpatientsinstratum1,T2 —T4,N2—N4,andS2—S4aresettobezerobecause
thesesix patients are in stratum1, not in stratum2, 3, or 4. Stratum 2 includesthose patients who had one recurrence and had either another recurrence orwere censored at end of follow-up. Therefore, stratum 2 has patients 3(censored at 14 months after the first recurrence ),4(second recurrence at 16
months ),5(second recurrence at 12 months ), and 6 (second recurrence at 15
months ). The time intervals between successive recurrences for these four
patients are (3, 14], (12,16], (6,12], and (3,15], respectively. The rearranged
data in order of the right end of the intervals are given in Table 13.8. Strata 3and 4 are constructed in a similar way. Once the data are rearranged exactlyas in Table 13.8, SAS and other software can be used to perform the analysis.
This dataarrangementalsofacilitatesexplanationofthe likelihoodfunction358
Table13.6 TumorRecurrenceDataforPatientswithBladderCancer /p63
Recurrence Time
Treatment Follow-upInitial Initial
GroupTime Number Size 1234
10 1 1
11 1 3
14 2 1
17 1 111 0 5 1
11 0 4 1 6
11 4 1 111 8 1 1
11 8 1 3 5
11 8 1 1 1 2 1 612 3 3 3
12 3 1 3 1 0 1 5
1 23 1 1 3 16 2312 3 3 1 3 9 2 1
1 24 2 3 7 10 16 24
1 25 1 1 3 15 2512 6 1 2
12 6 8 1 1
12 6 1 4 2 2 612 8 1 2 2 5
12 9 1 4
12 9 1 212 9 4 1
13 0 1 6 2 8 3 0
1 30 1 5 2 17 221 3 0 2 1 368 1 2
1 3 1 1 3 1 21 52 4
13 2 1 213 4 2 1
13 6 2 1
13 6 3 1 2 913 7 1 2
1 40 4 1 9 17 22 24
1 4 0 5 1 1 61 92 32 914 1 1 2
14 3 1 1 3
14 3 2 6 61 4 4 2 1 369
1 45 1 1 9 11 20 26
14 8 1 1 1 814 9 1 3
15 1 3 1 3 5
15 3 1 7 1 71 53 3 1 3 15 46 51
15 9 1 1
1 61 3 2 2 15 24 301 64 1 3 5 14 19 27
(Continued overleaf ) 359
Table13.6 Continued
Recurrence Time
Treatment Follow-upInitial Initial
GroupTime Number Size 1234
16 4 2 3 2 8 1 2 1 3
21 1 321 1 125 8 1 5
29 1 2
21 0 1 121 3 1 121 4 2 6
2 1 7 5 3 3135
21 8 5 121 8 1 3 1 721 9 5 1 2
22 1 1 1 1 7 1 9
22 2 1 122 5 1 322 5 1 5
22 5 1 1
2 26 1 1 6 12 1322 7 1 1 622 9 2 1 2
23 6 8 3 2 6 3 5
23 8 1 12 3 9 1 1 2 22 32 73 22 39 6 1 4 16 23 27
2 4 0 3 1 2 42 62 94 0
24 1 3 224 1 1 124 3 1 1 1 2 7
24 4 1 1
2 44 6 1 2 20 23 2724 5 1 224 6 1 4 2
24 6 1 4
24 9 3 325 0 1 12 50 4 1 4 24 47
25 4 3 4
25 4 2 1 3 825 9 1 3
Source:Wei et al (1989 )and StatLib web site: http//lib.stat.cmu.edu/datasets/tumor.
/p63Treatment group: 1, placebo; 2, thioteps. Follow-up time and recurrence time are measured in
months. Initial size is measured in centimeters. Initial number of 8 denotes eight or more initialtumors.360
Figure13.1 GraphicalpresentationofrecurrencetimesofthesixpatientsinTable13.7
(numbers in circle indicate the number of recurrences ).Table13.7 Sixof86BladderCancerPatientsfromthe TumorRecurrenceData /p63
Recurrence Time
Patient Treatment Follow-upInitial Initial
ID GroupTime Number Size 1 2 3 4
119 1 2
20 5 9 1 1
3 1 14 2 6 3
4 0 18 1 1 12 16
5 1 26 1 1 6 12 136 0 5 3 3 1 31 54 65 1
/p63Treatment group: 0, placebo; 1, thiotepa. Following-up time and recurrence time are measured
in months. Initial size is measured in centimeters for the largest initial tumor.
in(13.4.2 ).Weusestratum2toshowthesecondproductin (13.4.2 ).Instratum
2(s/p582),d/p81/p583(there are three uncensored observations: patients 5, 6 and 4,
according to the ordered recurrent times, 12, 15, and 16 months ). Therefore,
the second product is the product of three terms, one for each of these threepatients. Using the notations in (13.4.2 ), we renumber them as patient i/p581, 2,
and 3, respectively. The risk set at the first uncensored time t/p17/p16in stratum 2 361
Table13.8 RearrangedDatafromTable13.7forFittingPWPModelwith
NR-IndexedCoefficients /p63
ID NR TL TR CS T1 T2 T3 T4 N1 N2 N3 N4 S1 S2 S3 S4
3 1031100020006000
6 10310000300010005 1061100010001000
1 1090100010002000
4 10 1 21000010001000
2 10 5 90000010001000————————————————————————————————————————————5 26 1 21010001000100
3 23 1 40010002000600
6 23 1 510000030001004 21 2 1 61000001000100————————————————————————————————————————————5 31 2 1 31001000100010
4 31 6 1 800000001000106 31 5 4 61000000300010————————————————————————————————————————————5 41 3 2 60000100010001
6 44 6 5 11000000030001
/p63ID, patient ID number; NR, number of recurrence, where 1 /p58first recurrence, 2 /p58second
recurrence, and so on; TL and TR, left and right ends of time interval (TL, TR )defined by the
successive rcurrence times and the follow-uptime, where TR denotes either the successive
recurrence time or the follow-uptime; CS, censoring status, where 0 /p58censored, 1 /p58uncensored;
T1 to T4, treatment group; N1 to N4, initial number of tumors; S1 to S4, initial size.
(observed from patient 5 ),o rR(t/p17/p16,2) includes patients in stratum 2, whose
recurrent times, censored or not, are at least 12 ( t/p17/p16) months. Therefore,
R(t/p17/p16,2) includes all four patients in stratum 2. Similarly, the risk set at the
second uncensored time t/p17/p17in stratum 2, R(t/p17/p17,2), includes two patients
(patients 6 and 4 ), andR(t/p17/p18,2) includes only one patient (patient 4 ). Thus,
usingtheIDinTable13.7,let x/p17/p18—x/p17/p21denotethecovariatevectorsforpatients
3—6 in stratum 2, the second product in (13.4.2 )is
/p66/p17/p147
/p71/p14/p16exp[b/p30/p17x/p17/p71(t/p17/p71)]
/p26l/p43R(t/p17/p71,2)exp[b/p30/p17x/p74(t/p17/p71)]
/p58exp(b/p30/p17x/p17/p20)
exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21)
/p59exp(b/p30/p17x/p17/p21)
exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p21)/p59exp(b/p30/p17x/p17/p19)
exp(b/p30/p17x/p17/p19)(13.4.3 )
where thex’s represent the covariate vector (T1, T2, T3, T4, N1, N2, N3, N4,
S1, S2, S3, S4 ). For example, x/p17/p20/p58(0,1,0,0,0,1,0,0,0,1,0,0 ). It is clear that362
Table13.9 RearrangedDatafromTable13.7forFittingPWPModelwithCommon
Coefficients /p63
ID NR TL TR CS TRT N S
31031126
610310315106111111090112
41 0 1 21011
21 0 5 90011————————————————————————————————————————————52 6 1 21111
32 3 1 40126
62 3 1 510314 2 12 16 1 0 1 1————————————————————————————————————————————5 3 12 13 1 1 1 1
4 3 16 18 0 0 1 16 3 15 46 1 0 3 1————————————————————————————————————————————5 4 13 26 0 1 1 1
6 4 46 51 1 0 3 1
/p63TRT, treatment group; N, initial number; S, initial size.inthis modeltheregressioncoefficientsarestratumspecific.Theyrepresentthe
importanceofthecoefficientforpatientsindifferentstrataorpatientswhohaddifferent numbers of recurrent events. If the primary interest is the overallimportance of the covariates, regardless of the number of recurrences or if itcanbeassumedthattheimportanceofcovariatesisindependentofthenumberof recurrences, T1 —T4, N1—N4, and S1—S4 can be combined into a single
variable. As shown in Table 13.9, the three covariates are named TRT, N, andS for the six patients, and coefficients common to all strata can be estimated.
Data sets that have been so rearranged are ready for SAS and other software.
To use SAS and other software, the entire data set in Table 13.6 must first
be rearranged as in Table 13.8 or 13.9. This can also be accomplished using acomputer.
Table 13.10 gives the results from fitting the PWP model to the bladder
tumor data in Table 13.6 with stratum-specific coefficients and commoncoefficients.Noneofthestratum-specificcovariatesissignificantexceptN1,theinitial number of tumors in stratum 1 patients (p/p580.0017 ). There is no
significant difference between the two treatments in any stratum, and the sizeof the initial tumor has no significant effect on tumor recurrence. Whenstratificationisignored,the resultsare similar (thesecondpart ofTable13.10 ).
The number of initial tumors is the only significant prognostic factor, and therisk of recurrence increase would increase almost 13% for every one-tumorincrease in the number of initial tumors. 363
Table13.10 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom
FittedPWPModelswithStratum-specificor CommonCoefficients
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
Model with Stratum-Specific Coefficients
T1 /p570.526 0.316 2.774 0.0958 0.591 0.318 1.097
T2 /p570.504 0.406 1.539 0.2148 0.604 0.273 1.339
T3 0.141 0.673 0.044 0.8345 1.151 0.308 4.305
T4 0.050 0.792 0.004 0.9493 1.052 0.223 4.963N1 0.238 0.076 9.851 0.0017 1.269 1.094 1.472
N2 /p570.025 0.090 0.075 0.7840 0.976 0.818 1.164
N3 0.050 0.185 0.072 0.7887 1.051 0.731 1.511N4 0.204 0.242 0.712 0.3987 1.227 0.763 1.971
S1 0.070 0.102 0.470 0.4931 1.072 0.879 1.308
S2 /p570.161 0.122 1.722 0.1894 0.852 0.670 1.083
S3 0.168 0.269 0.390 0.5321 1.183 0.698 2.005
S4 0.009 0.339 0.001 0.9786 1.009 0.519 1.961
ModelwithCommonCoefficients
TRT /p570.333 0.216 2.380 0.1229 0.716 0.469 1.094
N 0.120 0.053 5.029 0.0249 1.127 1.015 1.251S /p570.008 0.073 0.014 0.9071 0.992 0.860 1.144
In the second PWP model, the follow-uptime starts from the immediately
preceding event or failure time. Analogous to (13.4.1 ), the second PWP model
can be written in terms of a hazard function as
h(t/p34b/p81,x/p71(t))/p58h/p15/p81(t/p57t/p81/p92/p16)exp[b/p30/p81x/p71(t)] (13.4.4 )
wheret/p81/p92/p16denotes the time of the preceding event. The time period between
two consecutive recurrent events or between the last recurrent event time andthe end of follow-upis called the gaptime.
For thelth subject, who fails at time t/p81/p74in stratum s, denote the gaptime as
u/p81/p74/p58t/p81/p74/p57t/p81/p92/p16/p74, wheret/p81/p92/p16/p74is the failure time of the lth subject in the stratum
s/p571. Letu/p81/p7/p16/p8/p58/p37/p58u/p81/p7/p66/p81/p8denote the ordered observed distinct gaptimes in
stratumsandR/p18(u,s)denote the set of subjects at risk in stratum sjust prior
togaptimeu. Again,R/p18(u,s)includesonly thosesubjectswho haveexperienced
the firsts/p571 strata. Then we have the partial likelihoodfor the second model
(13.4.4 ):
L(b)/p58/p147
s/p461/p66/p81/p147
i/p581exp[b/p30/p81x/p81/p71(t/p81/p7/p71/p8)]
/p26l/p43R/p18(u/p81/p7/p71/p8,s)exp(b/p30/p81x/p81/p74(t/p81/p7/p71/p8)](13.4.5 )364
Table13.11 RearrangedDatafromTable13.9for
FittingPWPGapTimeModelwithCommonCoefficients
ID NR GT CS TRT N S
31 3 1 1 2 661 3 1 0 3 151 6 1 1 1 111 9 0 1 1 24 11 21011
2 15 90011————————————————————————————42 4 1 0 1 1
52 6 1 1 1 1
3 21 10126
6 21 21031————————————————————————————53 1 1 1 1 1
43 2 0 0 1 1
6 33 11031————————————————————————————64 5 1 0 3 1
5 41 30111Note that risk sets in (13.4.5 )are defined by the ordered distinct gaptimes in
the strata rather than by the failure times themselves.
Using the notations in Table 13.9, let GT denote the gaptime, then
GT/p58TR—TL. Replacing TR and TL in Tables 13.8 and 13.9 by GT, the data
are ready for SAS and other software. Table 13.11 is the corresponding tablefor the same six patients in Table 13.9 using gap times. Using the notation ofExample 13.6, the second product in (13.4.5 )for stratum 2 is
/p66
/p17/p147
/p71/p14/p16exp[b/p30/p17x/p17/p71(t/p17/p71)]
/p26l/p43R/p18(u/p17/p7/p71/p8,2)exp[b/p30/p17x/p17/p74(t/p17/p74)]
/p58exp(b/p30/p17x/p17/p19)
exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21)
/p59exp(b/p30/p17x/p17/p20)
exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21)/p59exp(b/p30/p17x/p17/p21)
exp(b/p30/p17x/p17/p21)
Note that this is different from (13.4.3 ), due to a different definition of the risk
set.
The results from fitting the PWP gaptime model to all the data in Table
13.6 with stratum-specific coefficients and common coefficients are given inTable 13.12. Again, the number of initial tumors is the only significant 365
Table13.12 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom
theFittedPWPGapTimeModelswithStratum-SpecificorCommonCoefficients
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
Model with Stratum-Speci fic Coef ficients
T1 /p570.526 0.316 2.774 0.0958 0.591 0.318 1.097
T2 /p570.271 0.405 0.448 0.5034 0.763 0.345 1.687
T3 0.210 0.550 0.146 0.7022 1.234 0.420 3.626
T4 /p570.220 0.639 0.119 0.7301 0.802 0.229 2.807
N1 0.238 0.076 9.851 0.0017 1.269 1.094 1.472
N2 /p570.006 0.096 0.004 0.9469 0.994 0.823 1.200
N3 0.142 0.162 0.774 0.3791 1.153 0.840 1.582N4 0.475 0.203 5.492 0.0191 1.609 1.081 2.394
S1 0.070 0.102 0.470 0.4931 1.072 0.879 1.308
S2 /p570.119 0.119 1.003 0.3166 0.888 0.703 1.121
S3 0.278 0.233 1.425 0.2326 1.321 0.836 2.086
S4 0.043 0.290 0.022 0.8822 1.044 0.592 1.842
Model with Common Coef ficients
TRT /p570.279 0.207 1.811 0.1784 0.757 0.504 1.136
N 0.158 0.052 9.258 0.0023 1.171 1.058 1.297S 0.007 0.070 0.011 0.9157 1.007 0.878 1.156
covariates. There are no major differences between the two PWP models for
thisset of data. Itis impossibleto compare the coefficients obtainedin the twomodels. The first model defines time from the beginning of the study andtherefore is recommended if the entire course of recurrent events is of interest.Thesecond model is the choice if the primary interest is to model the gap timebetween events.
Suppose that the text file ‘‘C: /p33EX13d4d1.DAT’’ contains the successive
columns in Table 13.8 for the entire data set in Table 13.6: NR, TL, TR, CS,T1, T2, T3, T4, N1, N2, N3, N4, S1, S2, S3, and S4, and the text file‘‘C:/p33EX13d4d2.DAT’’containsthesevensuccessivecolumnsinTable13.9:NR,
TL, TR, CS, TRT, N, and S. The following SAS code can be used to obtainthe PWP models in Table 13.10.
data w1;
infile ‘c: /p33ex13d4d1.dat’ missover;
input nr tl tr cs t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4;
run;
title ‘‘PWP model with stratified coefficients‘;proc phreg data /p58w1;366
model (tl, tr )*cs(0)/p58t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4 / ties /p58efron;
where tl /p58tr;
strata nr;
run;
data w1;
infile ‘c: /p33ex13d4d2.dat’ missover;
input nr tl tr cs trt n s;
run;
title ‘‘PWP model with common coefficients‘;
proc phreg data /p58w1;
model (tl, tr )*cs(0)/p58trt n s / ties /p58efron;
where tl /p58tr;
strata nr;
run;
Suppose that the text file ‘‘C: /p33EX13d4d3.DAT’’ contains 15 successive
columns similar to Table 13.8 but with gaptime GT. The 15 columns are NR,GT, CS, T1, T2, T3, T4, N1, N2, N3, N4, S1, S2, S3, and S4. The text file‘‘C:qafii0)’07EX13d4d4.DAT’’containsthesuccessivesixcolumnsfromTable13.11:NR,GT, CS, TRT, N, and S. The following SAS, SPSS, and BMDP codes can beused to obtain the PWP gaptime models in Table 13.12.
SAS code:
data w1;
infile ‘c: /p33ex13d4d3.dat’ missover;
input nr gt cs t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4;
run;title ‘‘PWP gaptime model with stratified coefficients’’;proc phreg data /p58w1;
model gt*cs (0)/p58t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4 / ties /p58efron;
strata nr;
run;data w1;
infile ‘c: /p33ex13d4d4.dat’ missover;
input nr gt cs trt n s;
run;
title ‘‘PWP gaptime model with common coefficients‘;proc phreg data /p58w1;
model gt*cs (0)/p58trt n s / ties /p58efron;
strata nr;
run;
SPSS code:
data list file /p58‘c:/p33ex13d4d3.dat’ free
/n rg tc st 1t 2t 3t 4n 1n 2n 3n 4s 1s 2s 3s 4 .
coxreg gt with t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4
/status /p58cs event (1) 367
/strata /p58nr
/print /p58all.
data list file /p58‘c:/p33ex13d4d4.dat’ free
/ nr gt cs trt n s.
coxreg gt with trt n s
/status /p58cs event (1)
/strata /p58nr
/print /p58all.
BMDP 2L code:
/input file /p58‘c:/p33ex13d4d3.dat’ .
variables /p5815.
format /p58free.
/print cova.
Survival.
/variable names /p58nr, gt, cs, t1, t2, t3, t4, n1, n2, n3, n4, s1, s2, s3, s4.
/form time /p58gt.
status /p58cs.
response /p581.
/regress covariates /p58t1, t2, t3, t4, n1, n2, n3, n4, s1, s2, s3, s4.
strata /p58nr.
/input file /p58‘c:/p33ex13d4d4.dat’ .
variables /p586.
format /p58free.
/print cova.
Survival.
/variable names /p58nr, gt, cs, trt, n, s.
/form time /p58gt.
status /p58cs.
response /p581.
/regress covariates /p58trt, n, s.
strata /p58nr.
Anderson--Gill Model
ThemodelproposedbyAndersenandGill (1982 ),theAGmodel,assumesthat
all events are of the same type and are independent. The risk set in thelikelihood function is totally different from that in the PWP models. The riskset of a person at the time of an event would contain all the people who arestill under observation, regardless of how many events they have experiencedbeforethattime.Themultiplicativehazardfunction h(t,x/p71)fortheithpersonis
h(t,x/p71)/p58Y/p71(t)h/p15(t) exp[b/p30x/p71(t)]
whereY/p71(t), an indicator,equals 1 whenthe ith person isunder observation (at
risk)at timetand 0 otherwise and h/p15(t) is an unspecified underlying hazard368
Table13.13 RearrangedDatafromTable13.7for
FittingAGModel
ID TL TR CS TRT N S
10 9 0 1 1 22 0 5 9001130 3 1 1 2 63 3 1 401264 0 1 2101141 2 1 6 1 0 1 1
41 6 1 8 0 0 1 1
50 6 1 1 1 15 6 1 2111151 2 1 3 1 1 1 151 3 2 6 0 1 1 160 3 1 0 3 1
6 3 1 51031
61 5 4 6 1 0 3 164 6 5 1 1 0 3 165 1 5 3 0 0 3 1function. The partial likelihood for nindependent persons is
L(b)/p58/p76/p147
/p71/p14/p16/p147
/p82/p46/p15/p3Y/p71(t) exp (b/p30x/p71)
/p26/p76/p72/p14/p16Y/p72(t) exp (b/p30x/p72)/p4/p66/p71/p7/p82/p8(13.4.6 )
where /afii9829/p71(t)/p581 if theith person has an event at tand/p580 otherwise. Details
of this likelihood function and the estimation of the coefficients can be foundin Fleming and Harrington (1991 )and Andersen et al. (1993 ). Similar to the
PWP models, software packages are available to carry out the computationprovidedthatthedataarearrangedinacertainformat.Thefollowingexampleillustrates the terms in (13.4.6 )and the data format required by SAS.
Example 13.7 We use again the data in Table 13.6 to fit the AG model.
To explain the terms in the likelihood function, we use the data of the sixpeople in Table 13.7. In this model, every recurrent event is considered to beindependent.Therefore,wecanrearrangethedatabypersonandbyeventtime‘‘within’’ an individual. Table 13.13 shows the rearranged data. For example,the person with ID /p584 had two recurrences, at 12 and 16, and the follow-up
timeendedat 18. Thetimeintervals (TL,TR] are (0, 12], (12, 16], and (16,18],
and 12 and 16 are uncensored observations and 18 censored, since there wasno tumor recurrence at 18. For patients with ID /p581 and 2 (i/p581,2), the
respectivesecond product terms in (13.4.6 )are equal to 1 since /afii9829/p71(t)/p580,i/p581,
2, for allt. For patient 3 (i/p583),/afii9829/p71(t)/p581 only att/p583(the first tumor
recurrence time of the patient ). Thus, the respective second product has only 369
Table13.14 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom
theFittedAGModel
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
TRT /p570.412 0.200 4.241 0.0395 0.663 0.448 0.980
N 0.164 0.048 11.741 0.0006 1.178 1.073 1.293
S /p570.041 0.070 0.342 0.5590 0.960 0.836 1.102
one term att/p583 and the denominator of this term sums over all the patients
who are under observation and at risk at time t/p583. From Figure 13.1 it is
easily seen that the sum is over all six patients; that is, the respective secondproduct is
exp(b/p30x/p18)
/p26/p21/p72/p14/p16exp(b/p30x/p72)(13.4.7 )
For patient 4 (i/p584), the second product in (13.4.6 )contains two terms. One
is fort/p5812(the first recurrence time ), and att/p5812, patients 2, 3, 4, 5, and 6
are still under observation, and therefore the denominator of the term sumsover patients 2 to 6. The other term is for t/p5816(the second recurrence time )
and the denominator sums over patients 2, 4, 5, and 6. Patient 3 is no longerunder observation after t/p5814. Thus, the second product term for i/p584i s
exp(b/p30x/p19)
/p26/p21/p72/p14/p17exp(b/p30x/p72)
/p59exp(b/p30x/p19)
exp(b/p30x/p17)/p59/p26/p21/p72/p14/p19exp(b/p30x/p72)(13.4.8 )
Similarly, we can construct each term in (13.4.6 )and the partial likelihood
function.
Using SAS, we obtain the results in Table 13.14. The AG model identifies
treatment and number of initial tumor as significant covariates. Comparedwith placebo, thiotepa does slow down tumor recurrence.
ReaderscanconstructtheSAScodesfortheAGmodelbyusingTable13.13
and by following the codes given in Example 13.6.
Wei et al. Model
By using a marginal approach, Wei, Lin, and Weissfeld (1989 )proposed a
model,theWLW model,forthe analysisof recurrentfailures. Thefailures maybe recurrences of the same kind of event or events of different natures,depending on how the stratification is defined. If the strata are defined by the370
times of repeated failures of the same type, similar to the strata defined in the
PWP models, it can be used to analyze repeatedfailures of the same kind. Thedifference between the PWP models and the WLW model is that the latterconsiders each event as a separate process and treats each stratum-specific(marginal )partial likelihood separately. In the stratum-specific (marginal )
partial likelihood of stratum s, people who have experienced the (s/p571)th
failurecontributeeitheroneuncensoredoronecensoredfailuretimedependingon whether or not they experience a recurrence in stratum s, and the other
subjects contribute only censored times (forced as censored times ). Therefore,
each stratum contains everyone in the study. This is different from the PWPmodels, in which subjects who have not experienced the (s/p571)th failure are
not included in stratum s. If the strata are defined by the type of failure, the
WLW model acts like the competing risks model defined in Section 13.3, andthe type-specific (marginal )partial likelihood for the jth type simply treats all
failures of types other than jin the data as censored.
For thekth stratum of the ith person, the hazard function is assumed to
have the form
h/p73/p71(t)/p58Y/p73/p71(t)h/p73/p15(t) exp (b/p30/p73x/p73/p71),t/p460 (13 .4.9)
whereY/p73/p71(t)/p581, if theith person in the kth stratum is under observation, 0,
otherwise,h/p73/p15(t) is an unspecified underlying hazard function. Let R/p73(t/p73/p71)
denote the risk set with people at risk at the ith distinct uncensored time t/p73/p71in
thekth stratum. Then the specific partial likelihood for the kth stratum is
L/p73(b/p73)/p58/p76/p147
/p71/p14/p16
/p3exp(b/p30/p73x/p73/p71)
/p26l/p43R/p73(t/p73/p71)exp(b/p30/p73x/p73/p74)/p4/p66/p71(13.4.10 )
where /afii9829/p71/p581 if theith observation in the kth stratum is uncensored and 0
otherwise. The coefficients b/p73are stratum specific. In practice, if we are
interested in the overall effect of the covariates, we can assume that thecoefficients from different strata are equal (provided that there are no qualitat-
ive differences among the strata ), combine the strata and draw conclusions
above the ‘‘average effect’’ of the covariates. We again called the coefficients ofthese covariates common coefficients. The event time is from the beginning ofthe study in this model.
Similarto the PWP andAG models,the data must be arranged in a certain
format in order to use available software to carry out estimation of thecoefficients and tests of significance of the covariates. Using the same data asin Examples 13.6 and 13.7, the following example illustrates the terms in thestratum-specific likelihood function and the use of software.
Example 13.8 First, we use the same six patients to illustrate the compo-
nents in the stratum-specific likelihood function in (13.4.10 ). The format the
data have to be in for the available software, such as SAS, SPSS, and BMDP, 371
Table13.15 RearrangedDatafromTable13.7forFittingWLWModelwith
NR-IndexedCoefficients
ID NR TR CS T1 T2 T3 T4 N1 N2 N3 N4 S1 S2 S3 S4
3 1311000200060006 1310000300010005 161100010001000
1 190100010002000
4 11 21000010001000
2 15 90000010001000————————————————————————————————————————————1 290010001000200
5 21 21010001000100
3 21 400100020006006 21 510000030001004 21 610000010001002 25 90000001000100————————————————————————————————————————————1 390001000100020
5 31 310010001000103 31 400010002000604 31 80000000100010
6 34 61000000300010
2 35 90000000100010————————————————————————————————————————————1 490000100010002
3 41 40000100020006
4 41 800000000100015 42 600001000100016 45 110000000300012 45 90000000010001
is similar to that in the PWP and AG models except that all six people are in
each of the four strata (Table 13.15 ). The first stratum (NR/p581)is exactly the
sameasinTable 13.8.Thesix patientsareorderedaccordingto themagnitudeoftheeventtime (censoredornot,TR ).Instratum2(NR /p582),thethreepeople
(with ID /p584, 5, and 6 )whose times to the second tumor recurrence are
uncensored observations. Patients 1 and 2 had censored time at 9 and 59,respectively. Patient 3, who had no second recurrence and was observed until14 months, is considered censored at 14. The other strata are constructed in asimilarmanner.UsingthedataarrangementinTable13.15,wecanseethatforthesecondstratum,the likelihoodfunctionin (13.4.10 )hasthreeterms,onefor
each of persons 5, 6, and 4, whose /afii9829/p71/p581(CS/p581 in the table ). For patient 4,
the risk set at time t/p5816 has two individuals (ID/p584 and 2 ); for patient 5,
the risk set at time t/p5812 contains five individuals (ID/p582, 3, 4, 5, and 6 ); and
for patient 6, the risk set at time t/p5815 has three individuals (ID/p582, 4, and
6). Letx/p17/p72be the covariatevector of the patient with ID /p58jin stratum2; then372
Table13.16 RearrangedDatafromTable13.7for
FittingWLWModelwithCommonCoefficients
ID NR TR CS TRT N S
31 3 1 1 2 661 3 1 0 3 151 6 1 1 1 1
11 9 0 1 1 2
4 11 21011
2 15 90011————————————————————————————12 9 0 1 1 2
5 21 21111
3 21 401266 21 510314 21 610112 25 90011————————————————————————————13 9 0 1 1 2
5 31 311113 31 401264 31 80011
6 34 61031
2 35 90011————————————————————————————14 9 0 1 1 2
3 41 40126
4 41 800115 42 601116 45 110312 45 90011
the likelihood function in (13.4.10 )is
L/p17(b/p17)/p58/p21/p147
/p71/p14/p16/p3exp(b/p30/p17x/p17/p71)
/p26l/p43R/p17(t/p17/p71)exp(b/p30/p17x/p17/p74)/p4/p66/p71/p58exp(b/p30/p17x/p17/p19)
exp(b/p30/p17x/p17/p17)/p59exp(b/p30/p17x/p17/p19)
/p59exp(b/p30/p17x/p17/p20)
exp(b/p30/p17x/p17/p17)/p59exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21)
/p59exp(b/p30/p17x/p17/p21)
exp(b/p30/p17x/p17/p17)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p21)(13.4.11 )
Note that (13.4.11 )is different from (13.4.3 ). The likelihood function for the
other strata and for the entire data set in Table 13.6 can be constructed in asimilar manner. If we ignore the stratum-specific effect and are interested onlyin the average overall effect of the covariates, we combine T1 —T4, N1—N4,
and S1—S4. The rearranged data for the six patients are given in Table 13.16. 373
Table13.17 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom
theFittedWLWModelswithStratum-SpecificorCommonCoefficients
95%
Confidence Interval
Regression Standard Chi-Square Hazards
Variable Coefficient Error Statistic pRatio Lower Upper
Model with Stratum-Speci fic Coef ficients
T1 /p570.526 0.316 2.774 0.0958 0.591 0.318 1.097
T2 /p570.632 0.393 2.588 0.1077 0.531 0.246 1.148
T3 /p570.698 0.460 2.308 0.1278 0.496 0.202 1.225
T4 /p570.635 0.576 1.215 0.2703 0.530 0.171 1.639
N1 0.238 0.076 9.851 0.0017 1.269 1.094 1.472
N2 0.137 0.902 2.229 0.1354 1.147 0.958 1.373
N3 0.174 0.105 2.750 0.0973 1.189 0.969 1.460N4 0.332 0.125 7.112 0.0077 1.394 1.092 1.780
S1 0.070 0.102 0.470 0.4931 1.072 0.879 1.308
S2 /p570.078 0.134 0.337 0.5614 0.925 0.712 1.203
S3 /p570.214 0.183 1.371 0.2416 0.807 0.565 1.155
S4 /p570.206 0.231 0.800 0.3712 0.813 0.517 1.279
Model with Common Coef ficients
TRT /p570.585 0.201 8.460 0.0036 0.557 0.376 0.826
N 0.210 0.047 20.230 0.0001 1.234 1.126 1.352S /p570.052 0.070 0.548 0.4592 0.950 0.828 1.089
The results from fitting the WLW models to the entire data set in Table
13.6 are given in Table 13.17. The model with stratum-specific coefficientssuggests that more initial tumors accelerate tumor recurrence and the acceler-ation is particularly faster for the first recurrence and the third and fourthrecurrences. The signs of the coefficients for T1 —T4 suggest that thiotepa may
slow down tumor growth, but the evidence is not statistically significant. Themodel with common coefficients suggests that thiotepa is significantly moreeffective in prolonging the recurrence time. The results suggest that whenlooking at each stratum independently, there is no strong evidence thatthiotepais more effectivethan placebo.However,thecombinedestimate of thecommon coefficient provides stronger evidence that thiotepa is more effectiveover the course of the study.
13.5 MODELSFORRELATEDOBSERVATIONS
In Cox’s proportional hazards model and other regression methods, a key
assumptionisthat observedsurvivalor eventtimes are independent.However,
inmanypracticalsituations,failuretimesareobservedfromrelatedindividuals374
orfromsuccessiverecurrenteventsorfailuresofthesameperson.Forexample,
in an epidemiological study of heart disease, some of the participants may befrom the same family and therefore are not independent. These families withmultiple participants may be called clusters. In this case, the regression
methodsweintroducedearliermaynotbeappropriate.Severaltypesofmodelsintroduced especially for related observations are discussed by Andersen et al.(1993 ), Liang et al. (1995 ), Klein and Moeschberger (1997 ), and Ibrahim et al.
(2001 ). Details about these models are beyond the scope of this book. In the
following, we introduce briefly the frailty models.
Thefrailty models assume that there is an unmeasured random variable
(frailty )inthehazardfunction.Thisrandomvariableaccountsforthevariation
or heterogeneity among individuals in a cluster. It is also assumed that thefrailtyis independent of censoring.Let nbe the total numberof participants in
the study, some of them related and forming clusters. Let v/p71be the unknown
random variable, frailty, associated with the ith cluster, 1 /p45i/p58n. The frailty
model associated with the proportional hazards model can be written in termsof the log hazard function as
log[h/p71/p72(t;x/p71/p72/p34v/p71)]/p58log[h/p15(t)]/p59v/p71/p59b/p30x/p71/p72(13.5.1 )
for 1/p45j/p45m/p71and 1 /p45i/p58n, wherebdenotes the p/p591 column vector of
unknown regression coefficients, x/p71/p72is the covariate vector of the jth person in
theith cluster,m/p71is the number of individuals in the ith cluster, and h/p15(t)i sa n
unknown underlying hazard function. Compared with the Cox proportionalhazards model, the difference here is the random effect v/p71. Becausev/p71remains
the same in the ith cluster, the association between failure and covariates
within each cluster in this model is assumed to have a symmetric pattern. In afamily study, this model can be used, for example, to model failure timesobserved from siblings by treating each family as a cluster. This model wasproposed by Vaupel et al. (1979 )and developed and discussed by many
researchers, including Clayton and Cuzick (1985 ). The main approach to this
model is to assume that v/p71follows a parametric distribution.
The frailty model in (13.5.1 )can be extended to handle more complicated
situations. For example, the frailty can be a time-dependent variable [replacev/p71byv/p71(t)i n (13.5.1 )]. The frailty model with v/p71(t) can be used to model
successive or recurrent failure time as an alternative to the models in Section13.4. Another example is that there may be more than one type of frailty ineach cluster, and v/p71in(13.5.1 )can be replaced by v/p71/p59u/p71orv/p71/p59u/p71/p59w/p71, and
so on.
Inferences of these frailty models are also based on either a likelihood
functionorapartiallikelihoodfunction.Sincethemodelsinvolveaparametricdistribution,the likelihood or partial likelihood functions are complicatedandare beyond the level of this book.
The frailty models have not been used widely primarily because of the lack
of commercially available software. There are some computer programs 375
available; for example, a SAS macro is available for a gamma frailty model at
the Web site of Klein and Moeschberger (1997 ), and another program is
described by Jenkins (1997 ).
BibliographicalRemarks
Most of the major references for nonproportional hazards models have been
cited in the text of this chapter. Applications of these models include: stratified
models: Vasan et al. (1997 ), Aaronson et al. (1997 ), and Yakovlev et al. (1999 );
frailtymodels :Yashin andIachine (1997 ),Kessinget al. (1999 ),Siegmundet al.
(1999 ),Albert (2000 ),LeeandYau (2001 ),Wienkeelal. (2001 ),andXue (2001 );
competingrisksmodels : Mackenbach et al. (1995 ), Fish et al. (1998 ), Albertsen
et al. (1998 ), Blackstone and Lytle (2000 ), Yan et al. (2000 ), and Tai et al.
(2001 ).
EXERCISES
13.1Consider the cancer-free times from the participants with IDs 15 to 23
in Table 13.1. Follow Example 13.1 to construct the partial likelihoodfunctionbasedon the observed cancer-free times from these nine partici-pants.
13.2Considerthesurvivaltimesfrom30resectedmelanomapatientsinTable
3.1.LetAGEGdenoteagegroup,AGEG /p581ifage /p5845andAGEG /p582
otherwise. Fit the survival times with an AGEG-stratified Cox propor-tional hazards model with the covariates age, gender, initial stage, andtreatmentreceived.Discusstheassociationofthetreatmentreceivedwiththe survival time.
13.3Using the data in Table 12.4, following Example 13.5 and the sample
codes for SAS, SPSS, or BMDP, fit the competing risk model for stroke,CHD,other CVD, or STROKE/CHDseparately, and discuss the resultsobtained.
13.4Using the rearranged data in Tables 13.7 to 13.13 and following
Examples 13.6 to 13.8, complete construction of the remaining terms inthe partial likelihood function based on the PWP model (13.4.2 ), PWP
gaptimemodel (13.4.3 ),and AG model (13.4.9 ), and the remaining three
marginal likelihood functions based on the WLW model (13.4.13 ).376
CHAPTER 14
Identification of Risk Factors
Related to Dichotomous andPolychotomous Outcomes
In biomedical research we are often interested in whether a certain survival-relatedevent willoccur andthe importantfactorsthat influenceits occurrence.Such events may involve two or more possible outcomes; examples are thedevelopment of a given condition and response to a given treatment. If thegiven condition is diabetes and we are only interested in whether someonedevelops the disease (yes or no ), the outcome is binary or dichotomous. If we
are interested in whether the person develops impaired glucose tolerance,diabetes, or remains having normal glucose tolerance, there are three possibleoutcomes,orwesaytheoutcomeis trichotomous .Similarly,responsetoagiven
treatment can have dichotomous (response or no response )orpolychotomous
outcomes (complete response, partial response, or no response ).
To determine whether one is likely to develop a given disease, we need to
know the important characteristics (or factors )related to its development.
High- and low-risk groups can then be defined accordingly. Factors closelyrelated to the development of a given disease are usually called risk factors or
risk variables by epidemiologists. We shall use these terms in a broader sense
to mean factors closely related to the occurrence of any event of interest. Forexample, to find out whether a woman will develop breast cancer because oneof her relatives did, we need to know whether a family history of breast canceris an important risk factor. Therefore, we need to know the following:
1. Of age, race, family history of breast cancer, number of pregnancies,
experience of breast-feeding, and use of oral contraceptives—which aremost important?
2. Can we predict, on the basis of the important risk factors, whether a
woman will develop breast cancer or is more likely to develop breastcancer than another person?
377
In this chapter we introduce several methods for answering these ques-
tions. The general approach is to relate various patient characteristics(or independent variables, or covariates )to the occurrence of an event
(dependent or response variable )on the basis of data collected from
patients in each of the outcome groups. In the case of dichotomous out-comes, there are two outcome groups. For example, to relate variables suchas age, race, and number of pregnancies to the development of breast cancer,we need to collect information about these variables from a group of breast
cancerpatientsaswellasfromagroupofhealthynormalwomen.Foraneventwith polychotomous outcomes, we need to collect data from each outcomegroup.
Often, a large number of patient characteristics deserve consideration.
These characteristics may be demographic variables such as age; geneticvariables such as gene variant or phenotype; behavioral variables such assmoking or drinking behavior and use of estrogen or progesterone medic-ation; environmental variables such as exposure to sun, air pollution, oroccupational dust; or clinical variables such as blood cell counts, weight, andblood pressure. The number of possible risk factors can be reduced throughmedical knowledge of the disease and careful examination of the possible riskfactors individually.
In Section 14.1 we present two methods for examination of individual
variables. One is to compare the distribution of each possible risk variableamong the outcome groups. The other method is the chi-square test for acontingency table. This test is particularly useful when the risk variables arecategorical: for example, dichotomous or trichotomous. In this case, a 2 /p59cor
r/p59ccontingency table can be set up and a chi-square test performed. In
Section 14.2 we discuss logistic, conditional logistic, and other regressionmodels for binary responses and for examining the possible risk variablessimultaneously. Models for multiple outcomes are discussed in Section 14.3.
14.1 UNIVARIATE ANALYSIS
14.1.1 Comparing the Distributions of Risk Variables Among Groups
When the outcome is binary, it is often convenient to call an observation a
success or a failure. Successmay mean that a survival-related event occurred,
andfailurethat it failed to occur. Thus, a success may be a responding
patient, a patient who survives more than five years after surgery, or aperson who develops a given disease. A failure may be a nonrespond-ing patient, a patient who dies within five years after surgery, or a personwho does not develop a given disease. A preliminary examination of thedata can compare the distribution of the risk variables in the success andfailure groups. This method is especially appropriate if the risk variable is378
Table 14.1 Ages of 71 Leukemia Patients (Years)
Responders 20, 25, 26, 26, 27, 28, 28, 31, 33, 33, 36, 40, 40, 45, 45, 50, 50, 53 56,
62, 71, 74, 75, 77, 18, 19, 22, 26, 27, 28, 28, 28, 34, 37, 47, 56, 19
Nonresponders 27, 33, 34, 37, 43, 45, 45, 47, 48, 51, 52, 53, 57, 59, 59, 60, 60, 61, 61,
61, 63, 65, 71, 73, 73, 74, 80, 21, 28, 36, 55, 59, 62, 83
Source:Hart et al. (1977 ). Data used by permission of the author.
continuous. If, for example, the risk factor xis weight and the dependent
variable yis having cardiovascular disease, we may compare the weight
distribution of patients who have developed disease to that of disease-freepatients. If the disease group has significantly higher weights than those of thedisease-freegroup,wemayconsiderweightanimportantriskfactor.Common-ly used statistical methods for comparing two distributions are the t-test for
two independent samples if the assumption of normality holds and theMann —Whitney U-test if the normality assumption is violated and a non-
parametric test is preferred.
Similarly,if there are more than two possibleoutcomes, we canuse analysis
of varianceor the Kruskal —Wallisnonparametrictest to compare the multiple
distributionsofacontinuousvariable.Thefollowingexamplecomparestheagedistribution of responders with that of nonrespondersin a cancer clinical trial.
Example 14.1 Consider the ages of 71 leukemia patients—37 responders
and 34 nonresponders (response is defined as a complete response only )—
given in Table 14.1. Figure 14.1 gives us the estimated age distributions of the
two groups. By using the Mann —Whitney U-test (or Gehan’s generalized
Wilcoxon test ), we find that the difference in age between responders and
nonresponders is statistically significant (p/p580.01). In consequence, a question
may arise as to what age is critical. Can we say that patients under 50 mayhave a better chance of responding than do patients over 50? To answer thisquestion, one can dichotomize the age data and use the chi-square test,discussed next.
14.1.2 Chi-Square Test and Odds Ratio
The chi-square test and the odds ratio are most appropriate when the
independentvariableiscategorical.Iftheindependentvariableisdichotomous,a2/p592 table can be used to represent the data. Any variables that are not
dichotomous can be made so (with a loss of some information )by choos-
ing a cutoff point: for example, age less than 50 years. For multiple-outcome events, 2 /p59corr/p59ctables can be constructed. The independent 379
Figure 14.1 Age distribution of responders and nonresponders.
variables are then examined to find which ones (in some sense )provide the
best risk associations with the dependent variable. We first consider binaryoutcomes and independent variables that have two categories; that is, we setupa2/p592contingencytablesimilartoTable14.2foreachindependentvariable
and look for a high degree of proportionality.
The first step is to calculate the sample proportion of successes in the two
risk groups, a/C/p16andb/C/p17. Further analysis of the table is concernedwith the
precision of these proportions. A standard chi-square test can be used.380
Table 14.2 General Setup of a 2 /p592 Contingency Table
Risk Factor
Present ( E) Absent (E/p16)Total
Dependent variable
Success ab R/p16Failure cd R/p17Total C/p16C/p17N
Proportion of successes (success rate ) a/C/p16b/C/p17
If the rates of success for the two groups EandE/p16are exactly equal, the
expectednumber of patients in the ijth cell (ith row and jth column )is
E/p71/p72/p58N/p59R/p71N/p59C/p72N/p58R/p71/p59C/p72N(14.1.1 )
For example, in the top left cell, the expected number is
E/p16/p16/p58R/p16/p59C/p16N
since the overall success rate is R/p16/Nand there are C/p16individuals in the E
group.Similarexpectednumberscanbe obtainedfor eachof thefourcells. LetO/p71/p72be the number of patients observed in the ijth cell. Then the discrepancies
canbemeasuredbythedifferences( O/p71/p72/p57E/p71/p72). Inaroughsense,thegreaterthe
discrepancies, the more evidence we have against the null hypothesis that thesuccess rates are the same for the two groups. The chi-square test is based on
these discrepancies. Let
X/p17/p58/p17/p26
/p71/p14/p16/p17/p26
/p72/p14/p16(O/p71/p72/p57E/p71/p72)/p17
E/p71/p72
(14.1.2)
Underthenullhypothesis, X/p17followsthechi-squaredistributionwith1 degree
of freedom (df). The hypothesis of equal success rates for groups EandE/p16is
rejected if X/p17/p57/afii9851/p17/p16/p11/p63, where /afii9851/p17/p16/p11/p63is the 100 /afii9825percentage point of the chi-square
distribution with 1 degree of freedom. An alternative way to compute X/p17is
X/p17/p58(ad/p57bc)/p17N
R/p16R/p17C/p16C/p17(14.1.3) 381
Theodds ratio (Cornfield,1951 )is a commonlyusedmeasureofassociation
in2/p592 tables.The odds ratio (OR)is theratiooftwoodds:theoddsof success
when the risk factor is present and the odds of success when the risk factor isabsent. In terms of probabilities,
OR/p58P(success /p34E)/P(failure /p34E)
P(success /p34E/p16)/P(failure /p34E/p16)(14.1.4 )
Using the notation in Table 14.2, P(success /p34E)andP(failure /p34E)may
be estimated by a/C/p16andc/C/p16, respectively. Similarly, P(success /p34E/p16)and
P(failure /p34E/p16)may be estimated, respectively, by b/C/p17andd/C/p17. Therefore,
the numerator and denominator of (14.1.4 )may be estimated, respectively,
by
a/C/p16c/C/p16/p58a
c
and
b/C/p17d/C/p17/p58b
d
Consequently, the OR may be estimated by
OR/p19/p58a/c
b/d/p58ad
bc(14.1.5 )
which is also referred to as the cross-product ratio .
Several methods are available for an interval estimate of OR: for example,
Cornfield (1956 )and Woolf (1955 ). Cornfield’s method, which requires an
iterative procedure, is considered more accurate but more complicated thanWoolf’smethod.WoolfsuggestsusingthelogarithmofOR.Thestandarderrorof log OR
/p19may be estimated by
SE/p19(logOR/p19)/p58/p11
a/p591
b/p591
c/p591
d/p2/p16/p30/p17(14.1.6)
Then a 100 (1/p57/afii9825)%confidence interval (CI)for log OR is
logOR/p19/p60Z/p63/p30/p17SE/p19(logOR/p19)
The confidence interval for OR can be obtained by taking the antilog of the
confidence limits for log OR. If logOR/p51and logOR/p42are the upper and lower382
confidence limits for logOR, elogOR/p51andelogOR/p42are the upper and lower
confidence limits for OR.
Notice that in (14.1.5 ),i fborcis zero, OR /p19is undefined. If any one of the
four cell frequencies is zero, the estimated standard error in (14.1.6 )is also
undefined. Should this occur, some statisticians (Haldane, 1956; Fleiss, 1979,
1981 )suggest that 0.5 be added to each cell before using (14.1.5 )and (14.1.6 )
to solve the computational problem. However, if the cell frequencies are assmall as zero, the addition of 0.5 to each cell will substantially affect the
resulting estimate of OR and its standard error (Mantel, 1977; Miettinen,
1979 ). The estimates so obtained must be interpreted with caution.
An odds ratio of 1 indicates that the odds of success are the same whether
or not the risk factor is present. An odds ratio greater than 1 means that theodds in favor of success is higher when the risk factor is present, and thereforethereis a positiveassociationbetween the riskfactor andsuccess. Similarly,anodds ratio of less than 1 signifies a negative associationbetween the risk factorand success. The interpretation should not be based totally on the pointestimate. A confidence interval is always more meaningful, just as in any otherestimation procedure.
The chi-square statistic in (14.1.2 )may be used to test the null hypothesis
thatthereis no associationbetweentheriskfactorandsuccess,or H/p15:OR /p581.
The following example illustrates the chi-square test and odds ratio.
Example 14.2 In the study of the response rate of 71 leukemia patients
(Example 14.1 ), age is considered one of the possible risk variables. The
following 2 /p592 table is constructed.
Age/p5850 Age /p4650 Total
Response 27 10 37
Nonresponse 12 22 34
Total 39 32 71
The question is whether the response rates in the two age groups differ
significantly or whether age is associated with response.
TheX/p17value according to (14.1.3 )is
X/p17/p58(594/p57120)/p17(71)
(37)(34)(39)(32)/p5810.16
with 1 degree of freedom. Referenceto Table B-2 shows that the probability of 383
gettinga X/p17valueof10.16ifthetworesponseratesareequalinthepopulation
is less than 0.01. Hence the difference between the two response rates issignificant at the 1%level.
The estimate odds ratio, according to (14.1.5 ),i s
OR/p58(27)(22)
(10)(12)/p584.95
The data show that the odds in favor of response are almost five times higher
in patients under 50 years of age than in patients at least 50 years old. Thedifference is significantly different, as indicated by the chi-square test above.
To obtain a confidence interval for OR, we first compute log OR/p19/p581.60.
The estimated standard error of log OR /p19following (14.1.6 )is
SE/p19(logOR/p19)/p58/p11
27/p591
10/p591
12/p591
22/p2/p16/p30/p17/p580.515
A95%confidenceinterval for log OR is 1.60 /p601.96(0.515 ),or(0.59, 2.61 ), and
a 95% confidence interval for OR is ( e/p15/p13/p20/p24,e/p17/p13/p21/p16),o r (1.80, 13.60 ). The wide
interval may be due to the small cell frequencies. Note that the standard errorof log OR/p19is inversely related to the cell frequencies.
In this example, the cutoff point, 50, was chosen arbitrarily. It is often of
interesttotrymorethan onecutoff pointif thenumber ofobservationsineachcell is not too small.
There are cases where the independent variable has c/p572 classes. The
chi-squaretest can be extendedto 2 /p59ctables.The odds ratiomethod canalso
be extended to handle polychotomous independent variables. It is done byselecting one of the classes as the reference class (theE/p16group )and calculating
the measure of association of each of the other classes relative to the referenceclass. Formultiple-outcomeevents,the chi-square testcan be extended to r/p59c
tables. The expected frequencies are computed just as in (14.1.1 ), and compu-
tation of X/p17[chi-square distributed with (r/p571)(c/p571)degrees of freedom] is
thesame asin (14.1.2 )exceptthat the sum is over all r/p59ccells.Fordetails, see
Snedecor and Cocharan (1967, Sec. 9.7 ). The following example illustrates the
procedures.
Example 14.3 Suppose that in the study of response rates of leukemia
patients, another possible risk variable is the marrow absolute leukemic infil-
trate,whichis definedas the percentageofthe total marrowthat is eitherblast
cellsorpromyelocytes.Itisbelievedthatpatientsshouldbeclassifiedintothreeclasses: /p4545%,46 —90%,and /p5790%.The 2 /p593 table is given below. Numbers
in parentheses are expected frequencies. For example, 18.68 /p58(39)(34)/71.384
Marrow Absolute Infiltrate
/p4545% 46 —90% /p5790% Total
Response 4 (8.34)20(20.32 )13(8.34) 37
Nonresponse 12 (7.66)19(18.68 )3(7.66) 34
Total 16 39 16 71
Response rate, (%)25 51 81
OR/p19 1 3.16 13.0
95%CI for OR (0.86, 11.52 )(2.40, 70.46 )
Thequestioniswhetherthedifferencein marrowabsoluteleukemicinfiltrateis
related to response. The value of X/p17is
X/p17/p58(4/p578.34)/p17
8.34/p59(12/p577.66)/p17
7.66/p59/p37/p59(3/p577.66)/p17
7.66/p5810.17
Thenumberofdegreesoffreedomis3 /p571/p582.With X/p17/p5810.17and2degrees
of freedom, the probability that the three absolute infiltrate groups have thesame responserate is less than0.01. The datasuggest that patientswith a highpercentage of marrow absolute infiltrate tend to have a high response rate.Marrow absolute infiltrate may be an important factor in predicting response.
The OR
/p19s given in the table above are calculated using the /p4545% class as
thereference class (or group ). Forexample,for the /p5790%class, the oddsratio
is 13/p5912/4/p593/p5813. The 95% confidence intervals for the ORs are obtained
using (14.1.6 ). Although the odds ratio for the 46 —90%group is larger than 1,
the 95% confidence intervals covers 1. Therefore, the point estimate, 3.16,cannot be taken too seriously. It appears that the major difference is betweenthe/p5790%and /p4545%groups.
Individual examination of each independent variable can provide only a
preliminary idea of how important each variable is by itself. The relativeimportance of all the variables has to be examined simultaneously usingmultivariate methods. In the following section we discuss the linear logisticregression analysis.
14.2 LOGISTIC AND CONDITIONAL LOGISTIC REGRESSION
MODELS FOR DICHOTOMOUS RESPONSES
14.2.1 Logistic Regression Model for Prospective Studies
In a typical prospective study, a random sample of subjects is taken and the
valuesoftheindependentvariablesaremeasuredatagiventime (usuallycalled
baseline measurements ). The subjects are then followed for a given period of 385
time and the outcome (dependent )variable is measured at the end of the
follow-up. Therefore, for a prospective study, the independent variables areregardedasfixedquantitiesduringthefollow-up,buttheoutcomesarerandomand unknown. The purpose of a prospective study is to examine the outcomesandrelate them to the baselinemeasurements.Examples of prospectivestudiesare cohort epidemiologic studies and clinical trials.
Suppose that there are nsubjects andto some of whomthe event of interest
occurred. They are called successes; the others are failures. Let y/p71/p581 if the ith
subject is a success and y/p71/p580 if the ith subject is a failure. Suppose that for
each of the nsubjects, pindependent variables x/p71/p16,x/p71/p17,...,x/p71/p78are measured.
These variables can be either qualitative, such as gender and race, or quanti-tative, such as blood pressure and white blood cell count. The problem is torelate the independent variables, x/p71/p16,...,x/p71/p78, to the dichotomous dependent
variable y/p71.
LetP/p71be the probability of success, P/p71/p58P(y/p71/p581/p34x/p71/p16,...,x/p71/p78), for the ith
subject. The logistic regression model, proposed by Cox (1970 )assumes that
the dependence of the probability of success on independent variables is
P/p71/p58P(y/p71/p581/p34x/p71)/p58exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)
1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)
(14.2.1)
and
1/p57P/p71/p58P(y/p71/p580/p34x/p71)/p581
1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)(14.2.2)
wherex/p71/p58(x/p71/p15/p37x/p71/p78),x/p71/p15/p891, and b/p72are unknowncoefficients.The logarithm
of the ratio of P/p71and 1 /p57P/p71is a simple linear function of the x/p71/p72’s.
Let
/afii9838/p71/p58logP/p711/p57P/p71/p58/p78/p26
/p72/p14/p15b/p72x/p71/p72(14.2.3)
/afii9838/p71/p58log[P/p71/(1/p57P/p71)] iscalledthe logistic transform ofP/p71and(14.2.3 )isalinear
logistic model. Another name for /afii9838/p71islog odds. Thus, the model relates the
independent variables to the logistic transform of P/p71, or log odds. The
probability of success P/p71can then be found from (14.2.3 )or(14.2.1 ). In many
ways (14.2.3 )is the most useful analog for dichotomous response data of the
ordinary regression model for normally distributed data.
To estimate the coefficients b/p72’s, Cox suggests the maximum likelihood
method. Let y/p16,y/p17,...,y/p76be observations with dichotomous values on n
subjects.Thelikelihoodfunctionbasedon the binomialdistributioncontains afactor (14.2.1 )whenever y/p71/p581 and (14.2.2 )whenever y/p71/p580. Thus, the likeli-386
hood function is
L(b/p15,b/p16,...,b/p78)/p58/p76/p147
/p71/p14/p16P/p87/p71/p71(1/p57P/p71)/p16/p92/p87/p71
/p58/p147/p76/p71/p14/p16exp(y/p71/p26/p78/p72/p14/p15b/p72x/p71/p72)
/p147/p76/p71/p14/p16[1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)]/p58exp(/p26/p78/p72/p14/p15b/p72t/p72)
/p147/p76/p71/p14/p16[1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)]
(14.2.4)
where t/p72/p58/p26/p76/p71/p14/p16x/p71/p72y/p71. The log-likelihood function is
l(b/p15,b/p16,...,b/p78)/p58logL/p58/p78/p26
/p72/p14/p15b/p72t/p72/p57/p76/p26
/p71/p14/p16log/p31/p59exp/p1/p78/p26
/p72/p14/p15b/p72x/p71/p72/p2/p4(14.2.5)
The maximum likelihood estimates of b/p72’s that maximize the log-likelihood
function in (14.2.5 )can be obtained by solving the following pequations
simultaneously:
t/p72/p57/p76/p26
/p71/p14/p16x/p71/p72exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)
1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)/p580j/p580, 1,...,p(14.2.6)
This can be done by an iterative procedure such as the Newton —Raphson
procedure. The second derivative of lin(14.2.5 )is
I*/p72/p129/p72/p130/p58/p42/p17l
/p42b/p72/p129/p42b/p72/p130/p58/p57/p76/p26
/p71/p14/p16x/p71/p72/p129x/p71/p72/p130exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)
1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)
j/p16/p580,...,p;j/p17/p580,...,p (14.2.7 )
LetI/p72/p129/p72/p130/p58(/p571)I*/p72/p129/p72/p130. Then the estimated inverse of the Imatrix, I/p19/p92/p16, is the
asymptotic covariance matrix of the b/p72’s. If we use the notation in Section 7.1
and letb/p19/p58(b/p19/p15,b/p19/p16,...,b/p19/p78)/p30denote the MLE of b, the estimated covariance
matrix of the MLE b/p19isV/p19(b/p19)/p58(v/p71/p72)/p58(/p57/p42/p17l(b/p19)//p42b/p42b/p30)/p92/p16 /p58I/p19/p92/p16, where v/p71/p72denotes the ijth element of V/p19(b/p19)or the ijth element of I/p19/p92/p16.
The coefficients so obtained indicate the relationships between the variables
andthelogoddsinfavorofsuccess.Foracontinuousvariable,thecorrespond-ing coefficient gives the change in the log odds for an increase of 1 unit in thevariable.Foracategoricalvariable,thecoefficientisequaltothelogoddsratio(see Section 14.1 ).
An approximate 100 (1/p57/afii9825)%confidence interval for b/p72is
b/p19/p72/p60Z/p63/p30/p17/p40v/p72/p72
(14.2.8 )
where Z/p63/p30/p17is the 100 (1/p57/afii9825/2)percentile of the standard normal distribution. 387
To test the hypothesis that some of the b/p72’s are zero, a likelihood ratio test
can be used. For example, to test H/p15:b/p72/p580, the log-likelihood ratio test
statistic is
X/p42/p58/p572[l(b/p19/p15,b/p19/p16,...,b/p19/p72/p92/p16,0 ,b/p19/p72/p62/p16,...,b/p19/p78)/p57l(b/p19/p15,b/p19/p16,b/p19/p17,...,b/p19/p78)]
(14.2.9 )
where the first term is the maximized log-likelihood subject to the constraint
b/p72/p580. If the hypothesis is true, X/p42is distributed asymptotically as chi-square
with 1 degree of freedom.
An alternative test for the significance of the coefficients is the Wald test,
which can be written as
X/p53/p58b/p19/p17/p72v/p72/p72(14.2.10 )
Under the null hypothesis that b/p72/p580,X/p53has an asymptotic chi-square
distributionwith1degreeoffreedom.AlthoughtheWaldtestis usedbymany,it is less powerful than the likelihood ratio test (Hauck and Donner, 1977;
Jennings, 1986 ). In other words, the Wald test often leads the user to conclude
that the coefficient (consequently, the respective risk factor )is not significant
when, in fact, it is significant.
Similar to earlier discussion of model selection, forward, backward, and
stepwise variable selection methods can be used to select the risk factors thatare significantly associated with a dichotomous response. The independentvariables x/p71/p72in this model do not have to be the original variables. They can
be any meaningful transforms of the original variables: for example, thelogarithm of the original variable, log x/p71/p72, and the deviation of the variable
from its mean, x/p71/p72/p57x/p21/p72.
From (14.2.1 )and (14.2.2 ), the logarithm of the odds ratio for ith and kth
subjects is
logP/p71/(1/p57P/p71)
P/p73/(1/p57P/p73)
/p58/p78/p26
/p72/p14/p16b/p72(x/p71/p72/p57x/p73/p72)( 14.2.11 )
Thus,an estimate ofthe odds ratio canbe obtainedby replacing b/p72in(14.2.11 )
with its MLE, b/p19/p72.
From the estimated regression equation, a predicted probability of success
can be computed by substituting the values of the risk factors in the equation.Using these predicted probabilities, a goodness-of-fit test can be performed totest the hypothesis that the model fits the data adequately. Several such testsare available (Lemeshow and Hosmer, 2000 ): for example, the Pearson chi-
square test, the Hosmer —Lemeshow (Hosmer and Lemeshow,1980 )test, a test
statistic suggested by Tsiatis (1980 ), and the score of Brown (1982 ). In the
following, we introduce the Hosmer—L emeshow test.388
Letp/p71be the estimate of P/p71obtained from the fitted logistic regression
equation for the ith subject, i/p581,...,n. Thep/p71’s can be arranged in ascending
order from smallest to largest. Those probabilities and the correspondingsubjects are then divided into ggroups according to some cutoff points of the
probability.For example, let g/p5810 and the cutoff points of the probability be
equalto k/10,k/p581,2,...,10. Thus,the firstgroupcontains allsubjectswhose
estimatedprobabilitiesare less than or equal to 0.1, the second group containsall subjects whose estimated probabilities are less than or equal to 0.2, and so
on. Let n/p73be the number of subjects in the kth group. The estimated expected
number of successes for the kth group is
E/p73/p58/p76
/p73/p26
/p72/p14/p16p/p72k/p581, 2,...,g
The Hosmer —Lemeshow test statistic is defined as
C/p58/p69/p26
/p73/p14/p16(O/p73/p57E/p73)/p17
n/p73p/p21/p73(1/p57p/p21/p73)(14.2.12 )
where O/p73is the observed number of successes in the kth group and p/p21/p73is the
average estimated probability of the kth group, that is,
p/p21/p73/p581
n/p73/p76/p73/p26
/p72/p14/p16p/p72
Under the null hypothesis that the model is adequate, the distribution of Cin
(14.2.12 )iswellapproximatedbythechi-squaredistributionwith g/p572degrees
of freedom. The test is basically a chi-square test of the discrepancy betweenthe observed and predicted frequencies of success. Thus, a Cvalue larger than
the 100 /afii9825percentage point of the chi-square distribution (orpvalue less than
/afii9825)indicates that the model is inadequate.
Similarto otherchi-squaregoodness-of-fittests,theapproximationdepends
ontheestimatedexpectedfrequenciesbeingreasonablylarge.Ifalargenumber(say, far more than 20% )of the expected frequencies are less than 5, the
approximation may not be appropriate and the pvalue must be interpreted
carefully. If this is the case, adjacent groups may be combined to increase theestimated expected frequencies. However, Hosmer and Lemeshow warn that iffewer than six groups are used to calculate C, the test would be insensitiveand
would almost always indicate that the model is adequate.
Most statistical software packages provide programs for logistic regression
analysis:forexample,SAS (proceduresLOGISTIC,PHREG,andCATMOD ),
BMDP (proceduresLRandPR ),andSPSS (proceduresNOMREG,PROBIT,
PLUM, and LOGISTIC ). Most of them provide estimates of the coefficients
and test statistics, variable selection procedures, and tests of goodness of fit. 389
Example 14.4 In a study of 238 non-insulin-dependent diabetic patients,
10 covariates are considered possible risk factors for proteinuria (the outcome
variable ).Thelogisticregressionmethodis usedtoidentifythemostimportant
risk factors and to predict the probability of proteinuria on the basis of theserisk factors. The 10 potential risk factors are age, gender (1, male; 2, female ),
smoking status (0, no; 1, yes ), percentage of ideal body mass index, hyperten-
sion (0, no; 1, yes ), use of insulin (0, no; 1, yes ), glucose control (0, no; 1, yes ),
duration of diabetes mellitus (DM)in years, total cholesterol, and total
triglyceride. Among the 238 patients, 69 have proteinuria ( y/p71/p581).
Using the stepwise procedure in BMDP, it is estimated that at step 1, the
model contains only b/p15andb/p19/p15/p58/p570.896 and l(b/p19/p15)in(14.2.5 )is/p57143.292. At
step 2, duration of diabetes is added to the model because its maximumlog-likelihood value is the largest among all the covariates. The MLEs of thetwo coefficients are b/p19/p15/p58/p571.467 and b/p19/p16/p58/p570.055, and l(b/p19/p15,b/p19/p16)/p58/p57139.429.
Since
X/p42/p58/p572[l(b/p19/p15)/p57l(b/p19/p15,b/p19/p16)]/p587.726
which is significant (p/p580.005 ), the duration of DM is related significantly to
thechanceofproteinuria.TheHosmer —Lemeshowteststatisticforgoodnessof
fit with only duration of DM in the model, C/p589.814 with 8 degrees of
freedom, gives a pvalue of 0.278.
Atstep3,genderisaddedtothemodelbecauseitsadditionyieldsthelargest
maximum log-likelihood value among all the remaining covariates. The maxi-mumlog-likelihoodvalue, l(b/p19/p15,b/p19/p16,b/p19/p17)/p58/p57137.749,b/p19/p15/p58/p571.453,b/p19/p16/p58/p570.060,
andb/p19/p17/p58/p570.279. To test if gender is significantly related to proteinuria after
duration of DM, we perform the likelihood ratio test
X/p42/p58/p572[l(b/p19/p15,b/p19/p16)/p57l(b/p19/p15,b/p19/p16,b/p19/p17)]/p583.360
which is significant at p/p580.067. The stepwise procedure terminates after the
third step because no other covariates are significant enough to enter theregression model; that is, none of the other covariates have a pvalue less than
0.15, which is set by the program (BMDP ). If any covariate already in the
regressionbecomes insignificantafter some other variables are in, the insignifi-cant variable would be removed. The pvalues for entering and removing a
variable can be determined by the user. The default values for entering andremoving a variable are, respectively, 0.10 and 0.15. Thus, the procedureidentifies duration of DM and gender as the two most important risk factorsbasedonthe datagiven. Aquestionthat may be raisedatthis pointiswhetherone should include gender in the equation since its significance level is largerthan the commonly used 0.05. The recommendation is to include it since it iscloseto0.05andsincethe pvalue shouldnotbe theonly basisfordetermining
whether a covariate should be included in the model. In addition, theHosmer —Lemeshow goodness of fit test statistic C/p585.036, when gender is390
Table 14.3 Estimated Coefficients for a Linear Logistic Regression Model Using Data
from Diabetic Patients
Estimated Standard
Variable Coefficient Error Coefficient/SE exp (coefficient )
Constant /p571.453 0.264 /p575.504 0.234
Duration of DM 0.060 0.020 2.956 1.062
Sex /p570.279 0.152 /p571.836 0.756
included, yields a pvalue of 0.754. Thus, inclusion of this covariate improves
considerably the adequacy of the model. Thus, the final regression equationwith the two significant risk factors is
logP/p711/p57P/p71/p58/p571.453/p590.060(duration of DM )/p570.279 (gender )
Table 14.3 gives the details for the estimated coefficients.
The signs of the coefficients indicate that male patients and patients with a
longer duration of diabetes have a higher chance of proteinuria. Furthermore,for each increase of one year in duration of diabetes, the log odds increase by0.060. Probabilities of proteinuria can be estimated following (14.2.1 ). For
example, the probability of developing proteinuria for a male patient who hashad diabetes for 15 years is
P/p58e/p92/p15/p13/p23/p18/p17
1/p59e/p92/p15/p13/p23/p18/p17
/p580.303
where /p570.832 is obtained by substituting the values of the two covariates in
the fitted equation; that is, /p571.453 /p590.060 (15)/p570.279 (1)/p58/p570.832. Similarly,
for a female patient who has the same duration of diabetes, the probability is0.248.
In addition to individual variables, interaction terms can be included in the
logistic regression model. If the association between an independent variablex/p16and the dependent variable yis not the same in different levels of another
variable, x/p17, there is interaction between x/p16andx/p17. To check if there is
interaction, one can include the product of x/p16andx/p17in the regression model
and test the significance of this new variable. The following example illustratesthe procedure.
Example 14.5 It is well known that adriamycin is effective for treating
certain types of cancer. It is also well known that adriamycin is highly toxic. 391
Some patients develop congestive heart failure (CHF ), but others who receive
asimilardose ofadriamycindo not.In anattempttodetectfactorsthatwouldincrease the risk of developing adriamycin cardotoxicity, various patientcharacteristics of 53 cancer patients were studied. Seventeen of these patientsdeveloped CHF and 36 patients did not. After a careful investigation, it wasfound that the total dose ( z/p16) and percentage decrease in electrocardiographic
QRS voltage ( z/p17) are most closely related to CHF. Table 14.4 shows the data
and some summary statistics. The following linear logistic regression model
with transformed variables z/p16/p58z/p16/p57z/p21/p16andx/p17/p58z/p17/p57z/p21/p17is used:
/afii9838/p58logp
1/p57p
/p58b/p15/p59b/p16x/p16/p59b/p17x/p17/p59b/p18x/p16x/p17
The stepwise procedure selects percentage decrease in QRS as the most
important variable, followed by the total dose (TD)and interaction
(TD/p59QRS ). The logistic regression analysis results are given in Table 14.5.
The stepwise log-likelihood values given in the last column indicate that onlyQRS is significant since 2 (/p5710.185 /p5933.254 )/p5846.138, which yields a pvalue
less than 0.001. Neither the total dose nor the interaction is significant.
Suppose that the last three columnsof Table 14.4, Y, Z1, and Z2, are stored
in a text data file ‘‘C: /p33EX14d2d2.DAT’’, separated by a space. The following
SAS, SPSS, or BMDP codes can be used to obtain the results in Table 14.5.
SAS code:
data w1;
infile ‘c: /p33ex14d2d2.dat’ missover;
input y z1 z2;
x1/p58z1-517.679;
x2/p58z2-26.019;
x12/p58x1*x2;
run;proc logistic data /p58w1 descending;
model y /p58x1 x2 x12/ selection /p58s plcl plrl lackfit;
run;
SPSS code (forward selection method ):
data list file /p58‘c:/p33ex14d2d2.dat’ free
/ y z1 z2.
Compute x1 /p58z1-517.679.
Compute x2 /p58z2-26.019.
Compute x12 /p58x1*x2.
Logistic regression y with x1 x2 x12
/method /p58fstep
/print /p58all.392
Table14.4 TotalDoseandPercentDecreasein QRSof
53 Patients Receiving Adriamycin
Total Percent Decrease
Patient CHF /p63,yDose, z/p16in QRS, z/p17
1 1 435 41
2 1 600 71
3 1 600 51
4 1 540 40
5 1 510 636 1 740 797 1 825 618 1 535 449 1 510 53
10 1 483 27
11 1 460 5312 1 460 6013 1 550 6514 1 540 5815 1 310 41
16 1 500 64
17 1 400 4418 0 440 919 0 600 4220 0 510 1921 0 410 24
22 0 540 /p5724
23 0 575 3924 0 564 3525 0 450 1026 0 570 627 0 480 6
28 0 585 21
29 0 420 1430 0 470 131 0 540 3332 0 585 3333 0 600 4
34 0 570 2
35 0 570 536 0 510 1237 0 470 /p571
38 0 405 4439 0 575 14
40 0 540 /p5710
41 0 500 /p5743
42 0 450 2343 0 520 /p571
(Continued overleaf ) 393
Table 14.4 Continued
Total Percent Decrease
Patient CHF /p63,yDose, z/p16in QRS, z/p17
44 0 495 29
45 0 585 4046 0 450 30
47 0 450 23
48 0 500 12
49 0 540 /p5711
50 0 440 751 0 480 /p5722
52 0 550 2053 0 500 19
Source:Minow et al. (1977 ).
/p631, yes; 0, no; z/p21/p16/p58517.679,z/p21/p17/p5826.019.
Table 14.5 Linear Logistic Regression Analysis Results of Data in Table 14.4
Estimated Standard
Variable Coefficient Error Coefficient/SE Log Likelihood
Constant /p573.757 1.576 /p572.384 /p5733.254
QRS 0.254 0.102 2.480 /p5710.185
TD /p570.024 0021 /p571.160 /p579.225
TD/p59QRS 0.001 0.001 0.677 /p578.803BMDP codes for procedure LR:
/input file /p58‘c:/p33ex14d2d2.dat’ .
variables /p583.
format /p58free.
/variable names /p58y, z1, z2.
/transform x1 /p58z1-517.679.
x2/p58z2-26.019.
x12/p58x1*x2.
/regress depend /p58y.
Interval /p58x1, x2, x12.
Model /p58x1, x2, x12.
Start /p58in, in, in.
Move /p580, 0, 0.394
Method /p58mlr.
/print cell /p58used.
/end
When the independent variables are dichotomous or polychotomous, the
logistic regression coefficients can be linked with odds ratios. Consider thesimplest case, where there is one independent variable, x/p16, which is either 0 or
1. The linear regression model in (14.2.1 )and (14.2.2 )becomes
P(y/p581/p34x/p16)/p58e
b/p15/p59b/p16x/p16
1/p59eb/p15/p59b/p16x/p16
P(y/p580/p34x/p16)/p581
1/p59eb/p15/p59b/p16x/p16
Values of the model when x/p16/p580, 1 are
P(y/p581/p34x/p16/p580)/p58eb/p15
1/p59eb/p15
P(y/p581/p34x/p16/p581)/p58eb/p15/p59b/p16
1/p59eb/p15/p59b/p16
P(y/p580/p34x/p16/p580)/p581
1/p59eb/p15
P(y/p580/p34x/p16/p581)/p581
1/p59eb/p15/p59b/p16
The odds ratio in (14.1.4 )is
OR/p58P(y/p581/p34x/p16/p581)/P(y/p580/p34x/p16/p581)
P(y/p581/p34x/p16/p580)/P(y/p580/p34x/p16/p580)/p58eb/p15/p59b/p16
eb/p15/p58eb/p16
and the log odds ratio is log (OR)/p58log(eb/p16)/p58b/p16[this can also be derived
directly from (14.2.11 )]. Thus, the estimated logistic regression coefficient also
provides an estimate of the odds ratio, that is, OR /p19/p58eb/p19/p16.I f(b/p16/p42,b/p16/p51)is the
confidence interval for b/p16, the corresponding interval for OR is ( eb/p16/p42,eb/p16/p51).
Example 14.6 Consider the age (x)and response (y)data from the 71
leukemia patients presented in Table 14.6 (Examples 14.1 and 14.2 ). The
logistic regression analysis results are given in Table 14.7. Notice thatexp(b/p16)/p58exp(1.5994 )/p584.95, which is equal to theestimateof ORobtainedin
Example 14.2 using (14.1.5 ), and the standard error of b/p19/p16is the same as that
oflogOR/p19exceptforasmallrounding-offerror.TheconfidenceintervalforOR
can also be obtained from the logistic regression analysis results. 395
Table 14.6 Age and Response Data of 71 Leukemia
Patients
Response x/p5850(1) x/p4650(0) Total
Yes(1) 27 10 37
No(0) 12 22 34
Total 39 32 71
Table 14.7 Results of Logistic Regression Analysis of Data in Table 14.6
Estimated Standard
Variable Coefficient Error Coefficient/SE exp (coefficient )
Constant ( b/p19/p15) /p570.7885 0 .3814 /p572.067 0 .45
Age(b/p19/p16) 1.5994 0.5156 3.102 4.95
The relationship between the logistic regression coefficient and odds ratio
can be extended to polychotomous variables by creating dummy variables (or
design variables ). The following example illustrates the procedure.
Example 14.7 Consider the data in Example 14.3. The variable marrow
absolute infiltrate (MAI )has three levels. As in Example 14.3, the /p4545%level
is considered as the reference group. In this case, two design variables will beused and their values are assigned as follows:
D/p16/p58/p71 if MAI /p5846—90%
0 otherwiseD/p17/p58/p71 if MAI /p5790%
0 otherwise
For MAI, the design variable values and the respective number of responders
(Event )and total number of patient in each MAI level (N)are listed in the
following table.
MAI (%)D/p16D/p17Event N
/p454 5 00 41 6
46—90 1 0 20 39
/p5790 0 1 13 16
Usingthese design variables, the logistic regression analysisgives the results in
Table 14.8.396
Table 14.8 Results of Logistic Regression Analysis of Data in Example 14.7 and Two
Design Variables
Estimated Standard
Variable Coefficient Error Coefficient/SE exp (coefficient )
Constant /p571.0986 0.5774 /p571.903 0.33
MAI ( D/p16) 1.1499 0.6603 1.742 3.16
MAI ( D/p17) 2.5649 0.8623 2.974 13.00
The coefficient corresponding to D/p16, 1.1499, is the log odds ratio between
the46 —90%groupandthe /p4545%group.Theoddsratioisexp (1.1499 )/p583.16,
which is exactly equal to the estimate obtained in Example 14.3. Similarly, thecoefficient correspondingto D/p17is the log odds ratio betweenthe /p5790%group
and the /p4545%group. The odds ratio obtained from the regression coefficient,
13.00, is the same as that obtained in Example 14.3.
The estimated standard error for D/p16, 0.6603, is also the standard error of
logOR/p19. A 95%confidence interval for the coefficient is 1.1499 /p601.96(0.6603 ),
or(/p570.1443, 2.4441 ), and consequently, a 95% confidence interval for OR is
(e/p92/p15/p13/p16/p19/p19/p18,e/p17/p13/p19/p19/p19/p16), or (0.86, 11.52 ), which is identical to that obtained in
Example 14.3 using Woolf’s estimate of SE (log OR ).
Suppose that the datain Example14.6are arrangedin fourcolumnsfor D/p16,
D/p17, Event, and Nas in the table above and are saved in a text data file
‘‘C:/p33EX14d2d4.DAT’’. The values of D/p16,D/p17, Event, and Nin each row are
separated by a space. The following SAS, SPSS, or BMDP codes can be usedto obtain the results in Table 14.8.
SAS code:
data w1;
infile ‘c: /p33ex14d2d4.dat’ missover;
input d1 d2 event n;
run;proc logistic data /p58w1 ;
model event/n /p58d1 d2 / plcl plrl;
run;
SPSS code:
data list file /p58‘c:/p33ex14d2d4.dat’ free
/ d1 d2 event n.
Probit event OF n WITH d1 d2
/model /p58logit
/log/p582.718
/print /p58all. 397
BMDP code for procedure LR:
/input file /p58‘c:/p33ex14d2d4.dat’ .
variables /p584.
format /p58free.
/variable names /p58d1, d2, event, n.
/regress count is n.
fcount is event.
Interval /p58d1, d2.
/print cell /p58used.
/end
For a continuous independent variable, the logistic regression coefficient
givesthechangeinlogoddsforanincreaseof1unitinthevariable.Ingeneral,for an increaseof munits in the variable,the log odds ratiois equal to mtimes
the logistic regression coefficient. The derivation is left to the reader as anexercise.
When more than one independent variable is included in the logistic
regressionmodel,eachestimatedcoefficientcanbeinterpretedasanestimateofthe log odds ratio statistically adjusting for all the other variables. Forexample, in Example 14.4, the regression coefficient for the gender variable,/p570.279,is an estimate of the log odds ratio for femalesversus males, adjusting
for duration of diabetes. Or the adjusted odds ratio for females versus males isestimatedas exp (/p570.279 )/p580.76; suggesting that female diabetic patients have
alowerriskofhavingproteinuriathanthatofmalepatients,afteradjustingforduration of diabetes. This interpretation, commonly used by epidemiologists,is appropriate if the linear relationship between the log odds and the indepen-dent variables holds.
Press and Wilson (1978 )compare the logistic regression method to the
discriminant analysis and find that if the independent variables are normalwith identical covariance matrices, discriminant analysis is preferred. Underabnormality, the logistic regression method is preferred. In particular, if theindependentvariablesaredichotomous,wecannotexpecttopredictaccuratelythe probability of success with a discriminant function, even with a largeamountofdata.Theirexamplesshowthat thelogisticregressiongivesa highercorrect classification rate.
14.2.2 Logistic and Conditional Logistic Regression Model for
Retrospective Studies
As mentioned earlier, the logistic regression model defined in (14.2.3 )is
originally designed for prospective studies, where a set of covariates or
independent variables are measured at a baseline examination on a group ofpeople without the disease of interest. These subjects are then followed for aperiodoftimeanddevelopmentofthediseaseamongthemarerecordedduring398
follow-up. The model can be extended to analyze data from retrospective
studies, such as a case —control study. In a case—control study , cases (subjects
with the diseaseof interest )andcontrols (subjectswithoutthe disease )are first
selected and risk factor data such as exposure variables and other covariatesare collected retrospectively. For example, in a case —control study of lung
cancer and cigarette smoking, a group of lung cancer patients and a group ofpeople without lung cancer are selected. Their smoking histories are thencollected along with other risk factors. Therefore, in a case —control study,
participants are selected first based on their disease status, and their history ofrisk factor exposures is collected later. The purpose of a case —control study is
to estimate the association between the risk factors and the disease understudy.Usingprobabilityterms,wearedealingwiththeprobabilitythattheriskfactors take on certain values given that a person is a case or a control. Wedenote this conditional probability by P(x/p34y), wherexdenote the covariates
andythe outcome variable. Using the same notation, the probability of
interest in a prospective study is P(y/p34x). Based on conditional probability
theory, P(x/p34y)can be written as
P(x/p34y)/p58P(y/p34x)P(x)
P(y)
(14.2.13)
where y/p581 for cases and 0 for controls. Thus, P(x/p34y) is a function of P(y/p34x),
P(x), and P(y).
The likelihood function of the logistic regression model for a retrospective
study, similar to (14.2.4 ), is the product of terms in the form of P(x/p34y)in
(14.2.13 )for the cases and controls selected [ (P(x/p34y/p581)from a case and
P(x/p34y/p580)from a control )]. We introduce here two most widely used
approachestothislikelihoodfunction.Oneapproachconsiderstheprobabilityof case/control selection. Since in case —control studies, the cases and controls
areselectedfromthepopulationandthelikelihoodfunctionisbasedonsubjectselection, we introduce an indicator variable to denote whether a person isselected ( s/p581) or is not selected (s/p580). Letn/p16andn/p15be, respectively, the
numbers of selected cases and controls in the study. The likelihood function is
L/p2/p2/p58/p76
/p16/p147
/p71/p14/p16P(x/p71/p34y/p71/p581,s/p581)/p76/p15/p147
/p71/p14/p16P(x/p71/p34y/p71/p580,s/p581) (14 .2.14)
Define
/afii9843/p16/p58P(s/p581/p34y/p581)
to be the probability that a diseased person is selected for the study as a case
and
/afii9843/p15/p58P(s/p581/p34y/p580) 399
to be the probability that a disease-free person is selected for the study as a
control. Assume that the sampling probabilities depend only on disease statusand not on the covariates. Using Bayes’ theorem in probability theory and(14.2.1 ),it canbe shown (thederivationis left to thereaderasan exercise )that
the probability that a person is diseased given that he or she has risk factorsxand was selected for the study can be written as
P(y/p71/p581/p34x/p71,s/p581)/p58exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)
1/p59exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)
(14.2.15 )
where
b*/p15/p58b/p15/p59log/afii9843/p16/afii9843/p15(14.2.16 )
According to the conditional probability in (14.2.13 ), the first term in the
likelihood function in (14.2.14 )is
P(x/p71/p34y/p71/p581,s/p581)/p58exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)
1/p59exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)/p3P(x/p71/p34s/p581)
P(y/p581/p34s/p581)/p4(14.2.17)
Similarly, the second term in (14.2.14 )fory/p71/p580 can be obtained:
P(x/p71/p34y/p71/p580,s/p581)/p581
1/p59exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)/p3P(x/p71/p34s/p581)
P(y/p580/p34s/p581)/p4(14.2.18)
Substituting (14.2.17 )and (14.2.18 )into (14.2.14 ), we obtain the likelihood
function for a case —control study:
L/p2/p2/p58L(b*/p15,b/p16,...,b/p78)/p59/p76/p147
/p71/p14/p16P(x/p71/p34s/p581)
P(y/p34s/p581)(14.2.19)
where n/p58n/p15/p59n/p16,L(b*/p15,b/p16,...,b/p78)is the likelihood function for prospective
studies in (14.2.4 )except that the intercept term b/p15is replaced by b*/p15in
(14.2.16 ).If we assume that the probability distributionof the covariates, P(x),
contains no information about the parameters of interest, or P(x)is indepen-
dent of the coefficients b/p72, and the selection is independent of x, then
maximizing L/p2/p2toobtainestimatesof b/p72isequivalenttomaximizingonly L(b*/p15,
b/p16,...,b/p78) since P(y/p581/p34s/p581)/p58n/p16/nandP(y/p580/p34s/p581)/p58n/p15/n. This im-
plies that we can use the computer program for prospective studies to analyzecase—control study data except that the intercept term cannot be interpreted
meaningfully unless /afii9843/p16and/afii9843/p15are known.
In most practical situations, the assumption made above about P(x)is
reasonable.Historically,in earlyapplicationsof thelogistic regressionmethod,the covariates were assumed to have multivariate normal distribution. Then400
estimationofthecoefficients, b/p72in(14.2.19 ),wouldinvolvethedistribution P(x)
and thus became much more complicated. However, in practice, many of thecovariates are categorical or discrete and are therefore distinctly nonnormal.Thus, it is appropriate to allow P(x)to remain completely arbitrary and use
simply L(b*/p15,b/p16,...,b/p78)in(14.2.19 )in case —control studies.
Another approach to the likelihood function based on (14.2.13 )is to
consider a conditional probability instead. Suppose that n/p16cases and n/p15controls were selected in a case —control study and n/p58n/p16/p59n/p15; letx/p16,
x/p17,...,x/p76be the risk factor sets of the nsubjects without specifying which of
them pertain to the cases and which to the controls. Then the conditionalprobability that the first n/p16x’s are observed from the n/p16cases and the
remainder are from the n/p15controls may be written as
/p147/p76
/p16/p71/p14/p16P(x/p71/p34y/p581)/p147/p76/p71/p14/p76/p16/p62/p16P(x/p71/p34y/p580)
/p26/p43l/p16,...,l/p76/p129/p44(/p147/p76/p16/p71/p14/p16P(x/p74/p71/p34y/p581)/p147/p76/p71/p14/p76/p16/p62/p16P(x/p74/p71/p34y/p580))(14.2.20)
where the summation in the denominator is over the n!/(n/p16!n/p15!)possible ways
of selecting n/p16individuals as cases from the nsubjects, with the remaining n/p15as controls. By using (14.2.13 )and (14.2.1 ),(14.2.20 )reduces to
/p147/p76/p16/p71/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p71/p72)
/p26/p43l/p16,...,l/p76/p129/p44/p147/p76/p16/p71/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p74/p71/p72)(14.2.21 )
Comparing with (12.1.17 ),(14.2.21 )can be considered as a special case of
(12.1.17 )in which there is only one distinct uncensored failure time, say at
t/p581, and all of the n/p16persons failed at t/p581. The remaining n/p15subjects
survive longer than 1 and are censored, say at t/p582, while all nsubjects are at
risk at t/p581. This interpretation of (14.2.21 )permits us to apply the method
andcomputersoftwarefortheCoxproportionalhazardsmodelwithadiscretetime scale and ties to obtain an estimate of bin the logistic regression model
and to perform the corresponding inferences. The procedure is illustrated inExample 14.8. When n/p16andn/p15are large enough, it can be shown that an
analysisbasedonthisconditionalprobabilitywillproduceresultsequivalenttothose based on the likelihood function defined in (14.2.19 )(Efron, 1975;
Farewell, 1979; Breslow and Day, 1980 ). When n/p16andn/p15are large, one may
prefer using (14.2.19 )to(14.2.21 )since the former is easier. However, for a
case—control study with a matched or stratified design, analysis based on the
conditional probability defined in (14.2.21 )is a better choice than (14.2.19 )
(Efron, 1975; Farewell, 1979; Breslow and Day, 1980 ). In the following section
we discuss the application of logistic regression analysis for two widelyaccepted matched designs in case —control studies.
1 : R Matched Design
A widely used case —control design is to have one or more controls matched 401
for each case based on matching variables such as age and gender. Suppose
that for each case there are R(/p461) matched controls. Let x/p71/p72/p73denote the
observed value of the jth covariate ( j/p581,...,p)from the kth subject (k/p581
for the case and k/p582,..., R/p591 for matched controls )in the ith matched set
(i/p581,...,n). The nmatched sets are considered as the samples from the n
different strata defined by the matching variables. Following (14.2.21 )with
n/p16/p581 and n/p15/p58R, the conditionalprobability for the matchedset (1 case and
Rcontrols )in the ith stratum is
exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p16)
exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p16)/p59/p26/p48/p62/p16/p73/p14/p17exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p73)/p581
1/p59/p26/p48/p62/p16/p73/p14/p17exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p73/p57x/p71/p72/p16)]
(14.2.22 )
and thus the conditional likelihood function for all nstrata is the product of
thenterms in (14.2.22 ), that is,
/p76/p147
/p71/p14/p161
1/p59/p26/p48/p62/p16/p73/p14/p17exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p73/p57x/p71/p72/p16)](14.2.23)
When R/p581, that is, a one-to-one pair matching, the conditional likelihood
function obtained from (14.2.23 )reduces to
L(b/p16,...,b/p78)/p58/p76/p147
/p71/p14/p161
1/p59exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p17/p57x/p71/p72/p16)]
/p58/p76/p147
/p71/p14/p16exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p16/p57x/p71/p72/p17)]
1/p59exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p16/p57x/p71/p72/p17)](14.2.24)
Compared with the likelihood function for the ordinary logistic regression in
(14.2.4 ),theconditionallikelihoodfunction (14.2.24 )canbetreatedasaspecial
case of (14.2.4 )withy/p71/p891,b/p15/p890, and x/p71/p72can be replaced by the difference in
x/p71/p72between the case and its matched control. This fact permits the use of
computer programs for ordinary logistic regression in one-to-one matchedcase—control studies. The procedure is as follows:
1. Let nbe the number of case —control pairs.
2. Use x/p71/p72/p16/p57x/p71/p72/p17, the difference between covariates for the case (x/p71/p72/p16)and
its matched control (x/p71/p72/p17), as the independent variable in the model.
3. Let y/p71/p891 for all pairs.
4. Delete the intercept term b/p15from the model.402
n1:n0Matched Design or Stratified Design
Suppose that there are n/p16cases and n/p15controls in the ith stratum and
n/p58n/p16/p59n/p15. Let x/p71/p72/p73denote the observed value of the jth covariate
(j/p581,...,p)from the kth subject (k/p581,2,..., n/p16for the n/p16cases and
k/p58n/p16/p591,...,nfor the n/p15controls )in the ith stratum ( i/p581,...,m). From
(14.2.21 ), the contribution of the ith stratum to the conditional likelihood
function is
/p147/p76/p16/p73/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p73)
/p26/p43k/p16,...,k/p76/p129/p44/p147/p76/p16/p74/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p73/p74)(14.2.25 )
where the summation in the denominator is over all the n!/(n/p16!n/p15!)possible
waystoselect n/p16outofthe nsubjectsascasesandtheremaining n/p15ascontrols.
The term in (14.2.25 )has the same mathematical form as (14.2.21 ). Thus, as
discussed earlier, the computer software for the proportional hazards modelcan be used to estimate the coefficients.
In both 1: Randn/p16:n/p15matched designs, most of the other features in the
ordinary logistic regression model fitting, including the use of design (dummy )
variables and statistical inferences,remain the same. However, the goodness offittestofHosmerandLemeshowisnotapplicabletomatcheddesigns.Readerswho are interested in assessing the logistic regression model in matchedcase—control studies are referred to Pregibon (1984 )and Moolgavkar et al.
(1985 ).
The following example illustrates the basic procedure for the one-to-one
matched design using (14.2.24 ).
Example 14.8 Tostudy theeffect ofobesity, familyhistoryof diabetes,and
level of physical activity to non-insulin-dependent diabetes (NIDDM ),3 0
nondiabeticpersons arematchedwith30 NIDDMpatientsbyage andgender.Obesity is measured by body mass index (BMI ), which is defined as weight in
kilogramsdividedbyheightinmeterssquared.Familyhistoryofdiabetes (FH)
and levels of physical activity (PHY )are binary variables. Table 14.9 gives the
partially fictitious data. Following the procedure given above, the results offitting the three variables using BMDP are given in Table 14.10.
Suppose that text data file ‘‘C: /p33EX14d2d5.DAT’’ contains six successive
columns of data: BMIC, FHC, PHYC, BMIN, FHN, and PHYN, as in Table14.9, separated by a space. The following SAS, SPSS, or BMDP code can beused to generate the results in Table 14.10.
SAS code:
data w1;
infile ‘c: /p33ex14d2d5.dat’ missover;
input bmic fhc phyc bmin fhn phyn;bmi/p58bmic-bmin; 403
Table 14.9 Data of 30 Matched Case--Control Pairs
Case (Diabetic ) Control (Nondiabetic )
Pair BMI FH /p63PHY /p64BMI FH /p63PHY /p64
1 22.1 1 1 26.7 0 1
2 31.3 0 0 24.4 0 1
3 33.8 1 0 29.4 0 0
4 33.7 1 1 26.0 0 0
5 23.1 1 1 24.2 1 06 26.8 1 0 29.7 0 07 32.3 1 0 30.2 0 18 31.4 1 0 23.4 0 19 37.6 1 0 42.4 0 0
10 32.4 1 0 25.8 0 0
11 29.1 0 1 39.8 0 112 28.6 0 1 31.6 0 013 35.9 0 0 21.8 1 114 30.4 0 0 24.2 0 115 39.8 0 0 27.8 1 1
16 43.3 1 0 37.5 1 1
17 32.5 0 0 27.9 1 118 28.7 0 1 25.3 1 019 30.3 0 0 31.3 0 120 32.5 1 0 34.5 1 121 32.5 1 0 25.4 0 1
22 21.6 1 1 27.0 1 1
23 24.4 0 1 31.1 0 024 46.7 1 0 27.3 0 125 28.6 1 1 24.0 0 026 29.7 0 0 33.5 0 027 29.6 0 1 20.7 0 0
28 22.8 0 0 29.2 1 1
29 34.8 1 0 30.0 0 130 37.3 1 0 26.5 0 0
/p631, yes; 0, no.
/p641, physically active; 0, sedentary.
fh/p58fhc-fhn;
phy/p58phyc-phyn;
y/p581;
run;
proc logistic data /p58w1;
model y /p58bmi fh phy / noint plcl plrl lackfit;
run;404
Table 14.10 Results of a Logistic Regression Analysis of Data in Table 11.13
Estimated Estimated
Variable Coefficient Standard Error Coefficient/SE exp (coefficient )
BMI 0.090 0.065 1.381 1.094
FH 0.968 0.588 1.646 2.633PHY /p570.563 0.541 /p571.041 0.569
SPSS code:
data list file /p58‘c:/p33ex14d2d5.dat’ free
/ bmic fhc phyc bmin fhn phyn.
Compute bmi /p58bmic-bmin.
Compute fh /p58fhc-fhn.
Compute phy /p58phyc-phyn.
Compute y /p581.
Logistic regression y with bmi fh phy
/origin/print /p58all.
BMDP LR code:
/input file /p58‘c:/p33ex14d2d5.dat’ .
variables /p586.
format /p58free.
/variable names /p58bmic, fhc, phyc, bmin, fhn, phyn.
/transform bmi /p58bmic-bmin.
fh/p58fhc-fhn.
phy/p58phyc-phyn.
y/p581.
/regress depend /p58y.
Interval /p58bmi, fh, phy.
Model /p58bmi, fh, phy.
Start /p58in, in, in.
Constant /p58out.
Move /p580, 0, 0.
Method /p58mlr.
/print cell /p58used.
/end
Thefollowingexample illustratesthe estimatingproceduresfor the1: Rand
n/p16:n/p15matched case —control designs.
Example 14.9 Table 14.11 lists a subset of simulated data from a case —
control diabetes study that is based on a cohort study of heart disease with a 405
Table 14.11 A Subset of Age-Group and Gender-Matched DM Data in Example 14.9 /p63
AGE AGEG SEX SBP DBP LACR HDL LINSUL SMOKE DMS DM SN
51.8 50 1 148 91 1.35 37 2.81 0 1 0 1
50.9 50 1 116 95 1.04 43 2.74 0 1 0 150.9 50 1 114 85 1.26 64 1.98 1 1 0 150.9 50 1 120 80 1.56 52 2.53 1 1 0 1
54.6 50 1 119 71 1.55 30 2.87 0 2 0 1
50.8 50 1 123 78 1.40 33 3.31 0 2 0 1
53.3 50 1 119 75 1.69 40 2.13 1 3 1 172.0 70 0 129 73 1.06 25 2.69 0 1 0 273.1 70 0 120 68 0.87 30 2.76 0 1 0 272.8 70 0 111 66 2.52 73 3.17 0 2 0 270.3 70 0 115 65 3.16 42 2.96 0 2 0 2
72.1 70 0 140 66 3.18 52 3.48 0 2 0 2
72.8 70 0 136 72 3.36 59 3.22 0 2 0 271.1 70 0 133 85 2.95 73 3.25 0 3 1 256.4 55 1 110 74 0.58 43 2.18 0 1 0 355.7 55 1 122 77 1.18 34 2.76 1 1 0 356.9 55 1 114 74 1.10 25 2.62 0 1 0 3
58.5 55 1 104 74 1.24 23 2.58 0 1 0 3
55.2 55 1 128 77 1.25 43 2.77 0 2 0 355.9 55 1 130 83 1.34 44 2.34 1 2 0 357.4 55 1 116 79 2.23 38 2.46 1 3 1 360.7 60 0 136 85 2.04 42 3.66 0 1 0 462.0 60 0 115 74 1.32 33 2.85 0 1 0 4
64.7 60 0 155 89 2.55 46 3.73 1 1 0 4
64.4 60 0 191 107 3.66 34 2.67 1 2 0 460.5 60 0 109 74 0.89 64 2.88 0 2 0 462.4 60 0 106 72 0.88 35 3.44 1 2 0 462.8 60 0 234 91 7.28 49 2.38 1 3 1 473.2 70 0 119 72 1.00 47 2.53 0 1 0 5
73.0 70 0 128 70 2.87 51 2.62 0 1 0 5
72.2 70 0 124 69 1.43 33 2.49 0 1 0 573.7 70 0 128 72 2.12 38 3.30 0 2 0 571.7 70 0 111 68 3.16 64 3.14 0 2 0 571.2 70 0 104 67 3.00 40 3.30 0 2 0 574.5 70 0 140 82 2.84 46 2.95 1 3 1 5
58.2 55 0 112 77 2.84 71 2.23 0 1 0 6
57.3 55 0 111 77 2.23 41 2.57 0 1 0 658.7 55 0 120 76 1.85 60 2.57 0 1 0 657.2 55 0 120 73 0.55 48 3.10 0 2 0 657.2 55 0 112 73 0.49 45 2.86 0 2 0 655.5 55 0 120 73 /p570.53 49 3.12 0 2 0 6
59.5 55 0 156 76 3.22 59 3.14 1 3 1 6
78.4 75 0 119 75 1.28 53 1.72 1 1 0 777.8 75 0 112 74 1.39 44 1.80 1 1 0 775.5 75 0 123 74 1.41 72 1.93 0 1 0 778.3 75 0 149 84 0.53 40 2.84 0 1 0 7406
Table 14.11 Continued
AGE AGEG SEX SBP DBP LACR HDL LINSUL SMOKE DMS DM SN
76.9 75 0 153 75 3.78 43 2.95 0 2 0 7
75.4 75 0 144 77 4.57 45 2.54 0 2 0 777.9 75 0 156 86 5.38 39 3.34 0 3 1 768.0 65 0 123 70 1.62 48 2.49 0 1 0 8
66.2 65 0 131 72 1.71 56 2.47 0 1 0 8
65.8 65 0 136 80 3.82 56 2.84 1 1 0 8
68.8 65 0 120 66 2.20 45 2.45 0 1 0 868.0 65 0 162 60 2.35 62 3.86 0 2 0 867.9 65 0 115 54 2.33 39 3.92 0 2 0 867.8 65 0 132 79 2.68 42 2.47 0 3 1 863.1 60 1 123 80 1.81 46 1.93 1 1 0 9
61.3 60 1 122 78 1.41 85 1.88 0 1 0 9
60.8 60 1 131 83 2.43 31 1.90 1 1 0 961.8 60 1 109 69 1.27 61 2.35 1 2 0 963.1 60 1 114 73 1.04 38 2.53 0 2 0 960.7 60 1 130 76 0.84 46 2.59 0 2 0 962.2 60 1 133 85 2.10 37 3.04 0 3 1 9
78.7 75 0 147 85 0.48 58 2.62 0 1 0 10
77.4 75 0 167 84 0.92 71 2.66 0 1 0 1078.0 75 0 165 85 0.44 49 3.02 0 1 0 1075.2 75 0 117 70 2.05 42 2.88 0 2 0 1077.8 75 0 151 75 4.27 41 2.76 0 2 0 1078.5 75 0 137 74 2.13 40 2.86 0 2 0 10
78.5 75 0 156 81 5.33 52 327 0 3 1 10
56.5 55 0 108 71 1.58 41 2.27 0 1 0 1158.8 55 0 104 73 2.55 34 2.53 0 1 0 1155.7 55 0 135 77 2.06 106 2.32 1 1 0 1157.8 55 0 110 74 2.59 49 2.37 0 1 0 1157.3 55 0 153 91 0.54 44 3.13 0 2 0 11
55.4 55 0 141 94 1.47 61 3.15 0 2 0 11
56.9 55 0 123 78 3.72 40 3.18 0 3 1 1153.3 50 0 113 74 1.19 46 2.71 0 1 0 1250.8 50 0 143 89 3.45 78 1.84 0 1 0 1250.3 50 0 136 79 /p571.44 48 2.48 1 2 0 12
55.0 50 0 131 77 0.23 49 2.57 1 2 0 12
54.4 50 0 132 77 /p570.08 34 2.45 0 2 0 12
53.3 50 0 135 77 0.48 37 2.93 1 2 0 1250.7 50 0 114 78 2.52 41 4.37 0 3 1 12
/p63AGEG /p5850 if50 /p45age/p5855,/p5855if55 /p45age/p5860,/p5860if 60 /p45age/p5865,/p5865 if 65 /p45age/p5870,/p5870
if 70/p45age/p5875,/p5875 if 75 /p45age/p5880; SEX /p581 if male and /p580 if female; SMOKE /p581 if current
smoker and 0 otherwise; SBP, systolic blood pressure; DBP, diastolic blood pressure; LACR,logarithm of the ratio of urinary albumin and creatinine; HDL, high-density lipoprotein in
cholesterol; LINSUL, logarithm of insuline; DM /p581 if fasting glucose /p46126 mg/dL and /p580
otherwise; DMS, diabetic status defined by ADA fasting glucose criterion: DMS /p581 if normal
fasting glucose, /p582 if impaired fasting glucose, and /p583 if diabetic; SN, stratum number. 407
Table 14.12 Results from the Conditional Logistic Regression Model for the DM Data
in Example 14.9
95%
Confidence Interval
for Odds Ratio
Regression Standard Chi-Square Odds
Variable Coefficient Error Statistic pRatio Lower Upper
AGE 0.327 0.161 4.127 0.0422 1.39 1.01 1.90
DBP 0.046 0.020 5.639 0.0176 1.05 1.01 1.09LACR 0.395 0.133 8.776 0.0031 1.48 1.14 1.93
LINSUL 0.860 0.288 8.900 0.0029 2.36 1.34 4.16
baseline and second examinations (about five years after the baseline examin-
ation ). In this study, 33 persons with diabetes at the second examination are
selected, and for each of these cases, six age group (in five-year interval )- and
gender-matched diabetes-free controls are selected randomly from all partici-pantswithoutdiabetesinthesecondexamination.Thereare33strataandeachstratum contains one case (DM/p581)and its six matched controls (DM/p580),
for a total of 231 participants. The demographic, physical, blood, and urinarydata collected at the baseline examination of the first 12 strata are listed inTable14.11 and arranged by stratum. The conditionallogistic modelbased on(14.2.20 )is used for these 1:6 matched data to identify risk factors for diabetes.
The stepwise selection method is used to select the significant risk factors. Theresults are shown in Table 14.12. AGE, DBP, LACR, and LINSUL aresignificantriskfactorsfordiabetes.ThelargerthevaluesofAGE,DBP,LACR,and LINSUL, the higher is the risk of being diabetic.
Asnotedearlier,for (14.2.21 ),thecomputersoftwareforaCoxproportional
model with discrete time scale can be used to obtain an estimate of parameterb. Suppose that the text data file ‘‘C: /p33EX14d2d6.DAT’’ contains 12 successive
columns,separatedbyaspace,withdataasinTable14.11:AGE,AGEG,SEX,SBP, DBP, LACR, HDL, LINSUL, SMOKE, DMS, DM, and SN. Thefollowing SAS code shows how the SAS procedure for the proportionalhazards model with discrete time scale can be used to obtain an estimate ofparameter bfor the conditional logistic regression model in a matched
case—controlstudy (theresultsaregivenin Table14.12 ).In theSAScode,first,
we define a nominal variable (for survival time ), TIME, and let TIME /p581i fa
case (DM/p581), and /p582 if a control (DM/p580). This is accomplished by the
statement ‘‘time /p582-dm;’’. Second, DM is also used to indicate censoring
status, DM /p580 meaning censored, and /p581 uncensored. Thus,a case will have
an uncensored time 1 and a control will have a censored time 2.
data w1;
infile ‘c: /p33ex14d2d6.dat’ missover;408
Table 14.13 Data from Example 14.9 If Stratified by Age
Group Only
Number of Number of
Stratum AGEG DMs Non-DMs Total
15 0 —54 9 54 63
25 5 —59 8 48 56
36 0 —64 4 24 28
46 5 —69 4 24 28
57 0 —74 5 30 35
67 5 —79 3 18 21—— — ——Total 33 198 231
Table 14.14 Results from the Conditional Logistic Regression Model for the Data in
Example 14.9 with Strata Defined by Age Groups
95%
Confidence Interval
for Odds Ratio
Regression Standard Chi-Square Odds
Variable Coefficient Error Statistic pRatio Lower Upper
DBP 0.050 0.020 6.304 0.0120 1.05 1.01 1.09
LACR 0.391 0.124 9.876 0.0017 1.48 1.16 1.89
HDL /p570.037 0.019 4.034 0.0446 0.96 0.93 1.00
LINSUL 0.694 0.292 5.668 0.0173 2.00 1.13 3.55input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;
time/p582-dm;
run;proc phreg data /p58w1 noprint;
model time*dm (0)/p58age sbp dbp lacr hdl linsul smoke
/ ties /p58discrete selection /p58s;
strata sn;
run;
Using the same data, if we stratify by age group only, Table 14.13 lists the
number of cases and controls in each age group (stratum ). This can be
consideredasanexampleofastratifieddesignwithadifferentnumbersofcasesand controls in each stratum: stratum 1 has a 9:54 match, stratum 2 an 8:48match, and so on. The results from the conditional logistic regression modelbased on this new stratification are given in Table 14.14. 409
The following SAS code can be used to generate the results in Table 14.14.
The code can be modified to perform conditional logistic regression analysisfor data from any n/p16:n/p15matched design or a stratified design.
data w1;
infile ‘c: /p33ex14d2d6.dat’ missover;
input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;time/p582-dm;
run;
proc phreg data /p58w1 noprint;
model time*dm (0)/p58age sex sbp dbp lacr hdl linsul smoke
/ ties /p58discrete selection /p58s;
strata ageg;
run;
14.2.3 Other Models for Dichotomous Outcomes
In the logistic regression model (14.2.3 ), the left side is a function of the
probability of success, P/p71, and the right side is the linear combination of
covariates.The function,calleda link function ,definesthe relationshipbetween
the covariates and P/p71. In general, a link function represents the underlying
biological, physical, or epidemiological relationship between the mean of thedependentvariable (probabilityofsuccess )andthecovariates x/p16,x/p17,...,x/p78.I n
the logistic regressionmodel the link function,say g, is the logit functionof P/p71,
that is,
g(P/p71)/p58logit (P/p71)/p58logP/p711/p57P/p71
andg(P/p71)is assumed to be linearly related to the covariates, that is,
logP/p711/p57P/p71/p58/p78/p26
/p72/p14/p15b/p72x/p71/p72
or
P/p71/p58exp(/p26/p46/p72/p14/p15b/p72x/p71/p72)
1/p59exp(/p26/p46/p72/p14/p15b/p72x/p71/p72)(14.2.26)
Twootherformsoflinkfunction g(P) that assumealinearrelationshipwith
the covariates have been proposed and used in the literature. In the following,weintroducethesetwolinkfunctionsandthecorrespondingregressionmodel.
1.T he probit (or normit)function. This is the link function defined by the
inverse of the cumulative standard normal distribution function, /afii9818/p92/p16(·):
g(P/p71)/p58/afii9818/p92/p16(P/p71) (14 .2.27)410
The corresponding model is
/afii9818/p92/p16(P/p71)/p58/p78/p26
/p72/p14/p15b/p72x/p71/p72
or
P/p71/p58/afii9818/p1/p46/p26
/p72/p14/p16b/p72x/p71/p72/p2(14.2.28 )
2.The complementary log-log link function . This function is defined by
g(P/p71)/p58log[/p57log(1/p57P/p71)] (14.2.29 )
The corresponding model is
log[/p57log(1/p57P/p71)]/p58/p78/p26
/p72/p14/p15b/p72x/p71/p72
or
P/p71/p581/p57exp/p3/p57exp/p1/p46/p26
/p72/p14/p15b/p72x/p71/p72/p2/p4(14.2.30)
The logistic regression model in (14.2.26 )is for binary outcomes such as
diseased versus nondiseased. The model (14.2.28 )can be thought of as an
alternative model for binary outcome. In addition, it can be used to modelthose binary outcomes that are defined by a cutoff point on the basis of anormallydistributed variable.For example,the cutoff point may be defined bythelastquintileorquartileofacontinuousmeasurementinanepidemiologicalstudy. When the binary outcomes are defined by a cutoff point in anasymmetricdistribution,themodelin (14.2.30 )maybeappropriate.Themodel
(14.2.30 )can also be considered as a version of the Cox proportional hazards
model for grouped survival times (Kalbfleisch and Prentice, 1973 ).
Lety/p16,y/p17,...,y/p76be the observations with dichotomous values on the n
subjects: y/p71/p581 for success and y/p71/p580 for failure. Similar to (14.2.4 ), the
likelihood functions for the models in (14.2.28 )and (14.2.30 )can be obtained
by replacing the corresponding P/p71in the following formula:
L(b/p15,b/p16,...,b/p78)/p58/p76/p147
/p71/p14/p16P/p87
/p71/p71(1/p57P/p71)/p16/p92/p87/p71
The MLEs of the coefficients and the asymptotic likelihood inferences are
similar to those given in Section 14.2.1 for the ordinary logistic regressionmodel except that interpretation of the odds ratio is not possible for the lattertwo models. The procedure LOGISTIC in SAS provides options for all threemodels. 411
Table 14.15 Asymptotic Partial Likelihood Inference from the Regression Models with
Different Link Functions for the Data in Example 14.9
95%
Confidence Interval
for Odds Ratio
Regression Standard Chi-Square Odds
Variable Coefficient Error Statistic pRatio Lower Upper
Model with L ogit L ink Function
INTERCPT /p578.419 1.792 22.061 /p580.0001
DBP 0.044 0.018 5.673 0.0172 1.05 1.01 1.08
LACR 0.343 0117 8.627 0.0033 1.41 1.13 1.79
LINSUL 0.870 0.287 9.191 0.0024 2.39 1.38 4.27
Hosmer —Lemeshow test statistic 18.9460 0.0152
Model with Inverse Normal Link Function
INTERCPT /p574.532 0.953 22.597 /p580.0001
DBP 0.023 0.010 5.302 0.0213
LACR 0.186 0.065 8.060 0.0045
LINSUL 0.445 0.156 8.146 0.0043
Hosmer —Lemeshow test statistic 7.386 0.4956
Model with Log-Log Link Function
INTERCPT /p577.740 1.530 25.589 /p580.0001
DBP 0.038 0.016 5.919 0.0150
LACR 0.305 0.096 10.153 0.0014
LINSUL 0.785 0.241 10.592 0.0011
Hosmer —Lemeshow test statistic 17.415 0.0261Example 14.10 Consider the data in Example 14.9 as nonstratified data,
Table 14.15 gives the results from the regression models defined in (14.2.26 ),
(14.2.28 ), and (14.2.30 )by using the stepwise selection method. Based on
Hosmer —Lemeshow test statistics, the regression model with the inverse
normal link function gives a good fit to the data (p/p580.4956 ), whereas the
othertwomodelsdonot (p/p580.0152and p/p580.0261 ).Allthreemodelsidentify
DBP, LACR, and LINSUL as significant covariates for the development ofdiabetes.
The following SAS, SPSS, and BMDP codes may be used to generate the
results in Table 14.15.
SAS code:
data w1;
infile ‘c: /p33ex14d2d6.dat’ missover;
input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;
run;412
title ‘‘Regression model with the logit link function-generalized logistic regression’’;
proc logistic data /p58w1 descending;
model dm /p58age sex sbp dbp lacr hdl linsul smoke
/ selection /p58s lackfit link /p58logit;
run;title ‘‘Regression model with the inverse normal link function‘;proc logistic data /p58w1 descending;
model dm /p58age sex sbp dbp lacr hdl linsul smoke
/ selection /p58s lackfit link /p58probit;
run;
title ‘‘Regression model with the log-log link funtion‘;proc logistic data /p58w1 descending;
model dm /p58age sex sbp dbp lacr hdl linsul smoke
/ selection /p58s lackfit link /p58cloglog;
run;
SPSS code for the model in (14.2.28 )with the forward selection method:
data list file /p58‘c:/p33ex14d2d6.dat’ free
/ age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn.
Logistic regression dm with age sex sbp dbp lacr hdl linsul smoke htn
/method /p58fstep
/print /p58all.
BMDP code for procedure LR and the model in (14.2.28 ):
/input file /p58‘c:/p33ex14d2d6.dat’ .
variables /p5812.
format /p58free.
/variable names /p58age, ageg, sex, sbp, dbp, lacr, hdl, linsul, smoke,
dms, dm, sn.
Use/p58age, sex to smoke.
/regress depend /p58dm.
Interval /p58age, sex to smoke.
Method /p58mlr.
/print cell /p58used.
/end
14.3 MODELS FOR POLYCHOTOMOUS OUTCOMES
TheregressionmodelsinSection14.2canbeextendedtohandleoutcomesthat
have more than two categories. Thesecategoriesmay be nominal, forexample,different types of heart disease or psychological conditions; or ordinal, forexample,differentlevelsofglucoseintoleranceordifferentseverityofcommuni-cationdisorders.Anoutcomevariablewithmorethantwopossibilitiesiscalledpolychotomous orpolytomous . In this section we discuss first the model for 413
nominalpolychotomousoutcomes (generalizedlogisticregressionmodel ),then
the model for ordinal polychotomous outcomes (ordinal regression model ).
Details regarding these models can be found in Aitchison and Silvey (1957 ),
McCullagh (1980 ), Green (1984 ), McCullagh and Nelder (1989 ), Hosmer and
Lemeshow (1989, 2000 ), Cox and Snell (1989 ), Afifi and Clark (1990 ), Agresti
(1990 ), Collett (1991 ), and Ananth and Kleinbaum (1997 ).
14.3.1 Models for Nominal Polychotomous Outcomes:
Generalized Logistic Regression Models
LetY/p71denote the outcome for individual i. The outcome can be one of the m
nominalcategories,suchasdifferentcelltypesoflungcancer.Let Y/p71/p58kdenote
thatY/p71belongs to the kth category and k/p581,2,...,m. Suppose that for each
ofnsubjects, pindependent variables x/p71/p58(x/p71/p16,x/p71/p17,...,x/p71/p78)/p30are measured.
These variables can be either qualitative or quantitative. Let P(Y/p71/p58k/p34x/p71)be
the probability that Y/p71/p58kgiven the pmeasured covariates x/p71; then
/p26/p75/p73/p14/p16P(Y/p71/p58k/p34x/p71)/p581. Without loss of generality,using the last catalogas the
reference, the generalized logistic regression model
logP(Y/p71/p58k/p34x/p71)
P(Y/p71/p58m/p34x/p71)/p58a/p73/p59/p78/p26
/p72/p14/p16b/p73/p72x/p71/p72k/p581, 2,...,m/p571 (14 .3.1)
can be used to study the association of the covariates xto the outcome. To
simplifythe notation,let u/p73/p71/p58a/p73/p59/p26/p78/p72/p14/p16b/p73/p72x/p71/p72. Similarto (14.2.1 )and (14.2.2 ),
the model in (14.3.1 )assumes that the dependence on the covariates of the
probability of being in the kth category is
P(Y/p71/p58k/p34x/p71)/p58/p7exp(u/p73/p71)
1/p59/p26/p75/p92/p16/p72/p14/p16exp(u/p72/p71)k/p581, 2,...,m/p571
1
1/p59/p26/p75/p92/p16/p72/p14/p16exp(u/p72/p71)k/p58m(14.3.2)
This model reduces to the logistic regression model in (14.2.1 )and (14.2.2 )
when m/p582.
Letk/p16,...,k/p76be the outcomes observed for the nsubjects. Then the
log-likelihood function based on the noutcomes observed is the logarithm of
the product of all P(Y/p71/p58k/p71/p34x/p71)’s from the nsubjects, that is,
l(a/p16,a/p17,...,a/p75/p92/p16,b/p16,b/p17,...,b/p75/p92/p16)/p58logL/p58log/p3/p76/p147
/p71/p14/p16P(Y/p71/p58k/p71/p34x/p71)/p4(14.3.3 )
where P(Y/p71/p58k/p71/p34x/p71)is given in (14.3.2 )andb/p73/p58(b/p73/p16,...,b/p73/p78)/p30,k/p581,
2,...,m/p571. There are a total of (m/p571)(p/p591)unknown coefficients. The
estimation and hypothesis testing procedures for the coefficients are similar to414
those in the logistic regression model for dichotomous outcomes. Strictly
speaking, the models in (14.3.1 )are not logistic regression models if m/p572.
Therefore, the interpretation of the coefficients in these models needs to beclarified. Let us consider modeling the relationship between gender andcardiovascular disease status, NORMAL, STROKE, and CHD (coronary
heart disease ). Let the outcome variable Ybe defined as Y/p581 if CHD, /p582i f
STROKE, and /p583 if NORMAL, and the covariate SEX defined as SEX /p581
if male and /p580 if female. Then the two models according to (14.3.1 )are
logP(Y/p71/p581/p34SEX/p71)
P(Y/p71/p583/p34SEX/p71)
/p58a/p16/p59b/p16·SEX/p71
logP(Y/p71/p582/p34SEX/p71)
P(Y/p71/p583/p34SEX/p71)/p58a/p17/p59b/p17·SEX/p71
It is clear that neither of them is a logistic regression model. In the following,
we show how to interpret the coefficients b/p16andb/p17in these models. From the
first model,
logP(Y/p581/p34SEX /p581)/P(Y/p583/p34SEX /p581)
P(Y/p581/p34SEX /p580)/P(Y/p583/p34SEX /p580)
/p58logP(Y/p581/p34SEX /p581)
P(Y/p583/p34SEX /p581)/p57logP(Y/p581/p34SEX /p580)
P(Y/p583/p34SEX /p580)
/p58(a/p16/p59b/p16)/p57a/p16
/p58b/p16
and thus
P(Y/p581/p34SEX /p581)/P(Y/p583/p34SEX /p581)
P(Y/p581/p34SEX /p580)/P(Y/p583/p34SEX /p580)/p58exp(b/p16) (14 .3.4)
Now let us cast the data into a 3 /p592 contingency table as in Table 14.16. The
left side of (14.3.4 )can be estimated by
(f/n/p16)/(b/n/p16)
(e/n/p15)/(a/n/p15)/p58fa
be
However, if only the data from the normal and CHD participants are used,
fa
be/p58[f/(b/p59f)]/[b/(b/p59f)]
[e/(e/p59a)]/[a/(e/p59a)] 415
Table 14.16 Nominal Cross-Classification of
Cardiovascular (CVD) Status by Gender
SEX
CVD
Status (Y) Female (0) Male (1)
NORMAL (3) ab
STROKE (2) cd
CHD (1) ef——Total n/p15n/p16
which is an estimate of
P(CHD /p34Male )/[1/p57P(CHD /p34male )]
P(CHD /p34female )/[1/p57P(CHD /p34female )]
or the ratio of the odds of a male having CHD to the odds of a female having
CHD. Therefore, the exp( b/p19/p16)obtained from the first model can be interpreted
as an estimate of the ratio of the odds of a male having CHD to the odds ofa female having CHD if only the data from the normal and CHD participantsareused. Similarly,exp (b/p19/p17)obtainedfromthe secondmodel canbe interpreted
as an estimate of the ratio of the odds of a male having STROKE to the oddsof a female having STROKE if only the data from the normal and STROKEparticipants are used. The same interpretation also holds for coefficients ofcontinuous covariates in the models of (14.3.1 ); that is, an exponentiated
coefficient for a continuous covariate is the odds ratio of a 1-unit increase inthe covariate assuming that other covariates are the same.
Example 14.11 We use the data in Example 14.9 and assume that DM
(Y/p581), IFG (Y/p582), and NFG (Y/p583)are three nominal categories. Let the
referent category be NFG. For simplicity, only two covariates, systolic bloodpressure (SBP )and log insulin (LINSUL ), are included. Table 14.17 gives the
results from fitting these covariates to the model (14.3.1 ).
logP(ith participant is DM )
P(ith participant is NFG )
/p58logP(Y/p71/p581/p34x/p71)
P(Y/p71/p583/p34x/p71)
/p58/p577.648 /p590.026SBP/p71/p591.047LINSUL/p71
logP(ith participant is IFG )
P(ith participant is NFG )/p58logP(Y/p71/p582/p34x/p71)
P(Y/p71/p583/p34x/p71)
/p58/p574.949 /p590.011SBP/p71/p590.876LINSUL/p71416
Table 14.17 Asymptotic Partial Likelihood Inference from the Generalized Logistic Regression Model for Example 14.11
95%
Confidence Interval
Order of for Odds Ratio
Coefficients Regression Standard Chi-Square Odds
byS A S kVariable Coefficient Error Statistic pRatio Lower Upper
DM vs. NFG
b/p161 INTERCP /p577.648 1.648 21.530 /p580.0001
b/p181 SBP 0.026 0.010 6.300 0.0121 1.03 1.01 1.05
b/p201 LINSUL 1.047 0.304 11.870 00006 2.85 1.57 5.17
----------------------------------------------------------------------------------------------------------
IFG vs. NFG
b/p172 INTERCP /p574.949 1.427 12.020 00005
b/p192 SBP 0.011 0.010 1.410 0.2346 1.01 0.99 1.03
b/p212 LINSUL 0.876 0.262 11.150 0.0008 2.40 1.44 4.01
----------------------------------------------------------------------------------------------------------
DM vs. IFG
INTERCP /p572.699 1.850 2.130 0.1445
SBP 0.015 0.012 0.480 0.2239 1.02 0.99 1.04LINSUL 0.171 0.333 1.270 0.6066 1.19 0.62 2.28
----------------------------------------------------------------------------------------------------------
H/p15:b/p18/p58b/p191.48 0 .2239
H/p15:b/p20/p58b/p210.27 0 .6066
417
Consequently,
logP(ith participant is DM )
P(ith participant is IFG )/p58logP(Y/p71/p581/p34x/p71)
P(Y/p71/p582/p34x/p71)
/p58logP(Y/p71/p581/p34x/p71)
P(Y/p71/p583/p34x/p71)/p57logP(Y/p71/p582/p34x/p71)
P(Y/p71/p583/p34x/p71)
/p58(/p577.648 /p594.949 )/p59(0.026 /p570.011 )SBP/p71
/p59(1.047 /p570.876 )LINSUL/p71
/p58/p572.699 /p590.015SBP/p71/p590.171LINSUL/p71
Thus, the odds ratio is 1.03 [exp (0.026 )] times (or 3% higher )for a 1-unit
increase in SBP, and 2.85 [exp (1.047 )] times (or 185% higher )for a 1-unit
increase in LINSUL from the model for DM vs. NFG. The odds ratio is 2.40[exp (0.876 )] times (or 140%higher )for a 1-unit increase in LINSULfrom the
modelforIFGversusNFG.SBPis notsignificantinthemodelforIFGversusNFG (p/p580.2346 ). Neither SBP nor LINSUL is significant in the model for
DMversus IFG (p/p580.2239and p/p580.6066, respectively ).One canalso follow
the examples in Chapter 7, 9, 11, and 12 to perform additional statisticalinferences.Forinstance,wecantestwhetherthecoefficientsforSBPinthefirsttwo models are equal (whether the odds ratio for a 1-unit increase of SBP in
the model for DM versus NFG is equal to that in the model for IFG versusNFG ), that is, H/p15:b/p18/p57b/p19/p580(where the subscripts 3 and 4 are the orders of
the coefficients given by SAS ). From (11.2.13 ), under H/p15, Wald’s statistic,
X/p53/p58(b/p19/p18/p57b/p19/p19)/p17/(v/p18/p18/p59v/p19/p19/p572v/p18/p19), has an asymptotic chi-square distribution
with 1 degree of freedom, where v/p18/p18andv/p19/p19are the estimated variance of b/p18andb/p19, respectively, and v/p18/p19is the estimated covariance of b/p18andb/p19. From
Table 14.17, the hypothesis is not rejected (p/p580.2239 ). Similarly, the hypoth-
esisH/p15:b/p20/p57b/p21/p580 is not rejected (p/p580.6066 ); that is, there is insufficient
evidence to say that the change in odds ratio for a 1-unit increase in LINSULin the model for DM versus NFG is not equal to that in the model for IFGversus NFG.
The following SAS, SPSS, and BMDP codes can be used to obtain the
results in Table 14.17.
SAS code:
data w1;
infile ‘c: /p33ex14d2d6.dat’ missover;
input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;y/p584-dms;
run;
title ‘‘Generalized logistic regression model’’;
proc catmod data /p58w1;
direct sbp linsul;418
model y /p58sbp linsul
/ ml covb;
contrast ‘Equal coefficients for SBP’ all —parms 0 0 1 /p57100 ;
contrast ‘Equal coefficients for LINSUL’ all —parms00001 /p571;
run;
SPSS code:
data list file /p58‘c:/p33ex14d2d6.dat’ free
/ age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn.
Compute y /p584-dms.
nomreg y with sbp linsul
/print /p58fit history parameter lrt.
BMDP PR code:
/input file /p58‘c:/p33ex14d2d6.dat’ .
variables /p5812.
format /p58free.
/variable names /p58age, ageg, sex, sbp, dbp, lacr, hdl, linsul, smoke,
dms, dm, sn.
Use/p58age, sex to smoke.
/transform y /p584-dms.
/group codes (y)/p581, 2, 3.
Names (y)/p58DM, IFG, NFG.
/regress depend /p58y.
Level /p583.
Type /p58nom.
Interval /p58age, sex to smoke.
enter /p58.05, .05.
remove /p58/p58.05, .05.
/print cell /p58model.
/end
14.3.2 Model for Ordinal Polychotomous Outcomes:
Ordinal Regression Models
If the outcomes involve a rank ordering, that is, the outcome variable is
ordinal,severalmultivaluedregressionmodelsareavailable.Readersinterestedin these models are referred to McCullagh and Nelder (1989 ), Agresti (1990 ),
Ananth and Kleinbaum (1997 ), and Hosmer and Lemeshow (2000 ). In the
followingdiscussion,weintroducethemostfrequentlyusedmodel,thepropor-tionaloddsmodel.Inthismodel,theprobabilityofanoutcomebeloworequalto a given ordinal level, P(Y/p45k), is compared to the probability that it is
higher than the level given, P(Y/p57k).
LetY/p71betheoutcomeofthe ithsubject.Assumethat Y/p71canbeclassifiedinto
mordinal levels. Let Y/p71/p58kifY/p71is classified into the kth level and 419
k/p581,2,...,m. Suppose that for each of nsubjects, pindependent variables
x/p71/p58(x/p71/p16,x/p71/p17,...,x/p71/p78)/p30are measured. These variables can be either qualitative
orquantitative.Ifthe logit link function definedinSection14.2.3isused,similar
to the logistic regression model (14.2.3 ), we consider the following models:
logit (P(Y/p71/p45k/p34x/p71))/p58logP(Y/p71/p45k/p34x/p71)
1/p57P(Y/p71/p45k/p34x/p71)/p58a/p73/p59/p78/p26
/p72/p14/p16b/p72x/p71/p72
k/p581, 2,...,m/p571 (14.3.5 )
or, equivalently, let u/p73/p71/p58a/p73/p59/p26/p78/p72/p14/p16b/p72x/p71/p72,
P(Y/p71/p45k/p34x/p71)/p58exp(a/p73/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)
1/p59exp(a/p73/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)/p58exp(u/p73/p71)
1/p59exp(u/p73/p71)
k/p581, 2,...,m/p571 (14 .3.6)
Therefore,
P(Y/p71/p58k/p34x/p71)/p58P(Y/p71/p45k/p34x/p71)/p57P(Y/p71/p45k/p571/p34x/p71)
/p58/p7exp(u/p16/p71)
1/p59exp(u/p16/p71)k/p581
exp(u/p73/p71)
1/p59exp(u/p73/p71)/p57exp(u/p73/p92/p16/p71)
1/p59exp(u/p73/p92/p16/p71)k/p582,...,m/p571
1/p57exp(u/p75/p92/p16/p71)
1/p59exp(u/p75/p92/p16/p71)k/p58m(14.3.7)
Ifm/p582, that is, there are only two outcome levels, (14.3.7 )reduces to the
logistic regression model in (14.2.3 ). The models in (14.3.5 )can be thought of
as having only two outcomes [( Y/p45k) versus ( Y/p57k)] and therefore are
logistic regression models. Thus, interpretation of the coefficients, b/p72, such as
the exponentiated coefficient [exp (b/p72)] for a discrete or a continuous covariate
is similar to that in a logistic regression model.
Letk/p16,...,k/p76beobservedoutcomesfrom nsubjects.Thenthelog-likelihood
function based on the noutcomes observed is the logarithm of the product of
allP(Y/p71/p58k/p71/p34x/p71)’s from the nsubjects, that is,
l(a/p16,a/p17,...,a/p75/p92/p16,b/p16,b/p17,...,b/p78)/p58logL/p58log/p3/p76/p147
/p71/p14/p16P(Y/p71/p58k/p71/p34x/p71)/p4(14.3.8 )
where P(Y/p71/p58k/p71/p34x/p71)is as givenin (14.3.7 ). Themaximumlikelihoodestimation
and hypothesis-testing procedures for the coefficients are similar to thosediscussed previously. If the probit link function in(14.2.27 )is used, the models420
and formula corresponding to (14.3.5 )—(14.3.7 )are
/afii9818/p92/p16(P(Y/p71/p45k/p34x/p71))/p58a/p73/p59/p78/p26
/p72/p14/p16b/p72x/p71/p72k/p581, 2,...,m/p571
P(Y/p71/p45k/p34x/p71)/p58/afii9818(u/p73/p71)k/p581, 2,...,m/p571
P(Y/p71/p58k/p34x/p71)/p58P(Y/p71/p45k/p34x/p71)/p57P(Y/p71/p45k/p571/p34x/p71)
/p58/p7/afii9818(u/p16/p71) k/p581
/afii9818(u/p73/p71)/p57/afii9818(u/p73/p92/p16/p71)k/p582,...,m/p571
1/p57/afii9818(u/p75/p92/p16/p71) k/p58m
If the complementary log-log link function in(14.2.29 )is used, the models and
formula corresponding to (14.3.5 )—(14.3.7 )are
log[/p57log(1/p57P(Y/p71/p45k/p34x/p71))]/p58a/p73/p59/p78/p26
/p72/p14/p16b/p72x/p71/p72k/p581, 2,...,m/p571
P(Y/p71/p45k/p34x/p71)/p581/p57exp[/p57exp(u/p73/p71)] k/p581, 2,...,m/p571
P(Y/p71/p58k/p34x/p71)/p58P(Y/p71/p45k/p34x/p71)/p57P(Y/p71/p45k/p571/p34x/p71)
/p58/p71/p57exp[/p57exp(u/p16/p71)] k/p581
exp[/p57exp(u/p73/p92/p16/p71)]/p57exp[/p57exp(u/p73/p71)] k/p582,...,m/p571
exp[/p57exp(u/p75/p92/p16/p71)] k/p58m
The log-likelihood function based on these two models can be obtained by
replacing P(Y/p71/p58k/p71/p34x/p71)in(14.3.8 )with the respective expressions above.
Example 14.12 Now consider the NFG, IFG, and DM categories in
Example14.9 that represent three levels of severity in glucoseintolerance. DM(diabetes )is defined as fasting plasma glucose (FPG )/p46126 mg/dL, IFG
(impaired fasting glucose )as FPG between 110 and 125 mg/dL, and NFG
(normal fasting glucose )as FPG /p58110 mg/dL. Thus, it is reasonable to
consider the outcome variable as ordinal. Let the outcome variable Y/p581i f
DM, 2 if IFG, and 3 if NFG. We fit the models in (14.3.5 )using the SAS
procedure LOGISTIC with all the covariates. The SAS program allows usersto use a variable selection method (forward, backward, and stepwise ). In this
case,weuse thestepwiseselectionmethod, andthe resultsare given in the firstpart of Table 14.18. The stepwise method identifies SBP and LINSUL assignificant independent variables. For k/p581 [i.e., we compare diabetes with 421
Table 14.18 Asymptotic Partial Likelihood Inference from the Ordinal Regression Model with Different Link Functions for
the Diabetic Status Data in Example 14.9
95%Confidence
Interval for Odds Ratio
Regression Standard Chi-Square Odds
k Variable Coefficient Error Statistic p Ratio Lower Upper
Model with Logit Link Function
1 INTERCP1 /p576.753 1.183 32.571 0.0001
2 INTERCP2 /p575.485 1.151 22.708 0.0001
SBP 0.019 0.007 6.114 0.0134 1.02 1.00 1.03LINSUL 0.925 0.213 18.803 0.0001 2.52 1.67 3.90
Log-likelihood ratio statistic for H/p15:b/p16/p58b/p17/p580/p6326.831 0 .0001
Model with Inverse Normal Link Function
1 INTERCP1 /p573.971 0.677 34.415 0.0001
2 INTERCP2 /p573.240 0.664 23.790 0.0001
SBP 0.011 0.004 6.311 0.0120
LINSUL 0.530 0.123 18.674 0.0001
Log-likelihood ratio statistic for H/p15:b/p16/p58b/p17/p5802 6 .261 0 .0001
Model from Complementary Log-Log Link Function
1 INTERCP1 /p575.626 0.915 37.813 0.0001
2 INTERCP2 /p574.562 0.894 26.025 0.0001
SBP 0.014 0.006 5.721 0.0168
LINSUL 0.715 0.162 19.534 0.0001
Log-likelihood ratio statistic for H/p15:b/p16/p58b/p17/p5802 5 .835 0 .0001
/p63b/p16andb/p17are coefficients for SBP and LINSUL, respectively.
422
nondiabetes (NFG /p59IFG)] the estimated model in (14.3.5 )is
logP(Y/p71/p451/p34x/p71)
1/p57P(Y/p71/p451/p34x/p71)/p58logP(participant iis diabetic )
P(participant iis nondiabetic )
/p58/p576.753 /p590.019SBP/p71/p590.925LINSUL/p71
Fork/p582, the estimated model in (14.3.5 )is
logP(Y/p71/p452/p34x/p71)
1/p57P(Y/p71/p452/p34x/p71)/p58logP(participant iis either DM or IFG )
P(participant iis NFG )
/p58/p575.485 /p590.019SBP/p71/p590.925LINSUL/p71
Accordingto (14.3.7 ), we can estimatetheprobabilityof developingDM, IFG,
or remaining NFG. For example, the probability of developing IFG is
P(Y/p71/p582/p34x/p71)/p58P(participant iis IFG )
/p58exp(/p575.485 /p590.019SBP/p71/p590.925LINSUL/p71)
1/p59exp(/p575.485 /p590.019SBP/p71/p590.925LINSUL/p71)
/p57exp(/p576.753 /p590.019SBP/p71/p590.925LINSUL/p71)
1/p59exp(/p576.753 /p590.019SBP/p71/p590.925LINSUL/p71)
Thus, for a person whose systolic blood pressure is 140 mmHg and whose log
insulin is 3, the probability of developing IFG can be obtained by pluggingthese values into the preceding equation. The result is
P(participant is IFG )/p580.951
1/p590.951/p570.268
1/p590.268
/p580.276
As noted earlier, the coefficients in these models can be interpreted as those
in the ordinary logistic regressionmodel for binary outcomes. In this example,the higher SBP and LINSUL are, the higher the odds of having DM than ofnot having DM, or the higher the odds of having either DM or IFG than ofbeing NFG. The odds ratio is 1.02 [exp (0.019 )] times (or 2% higher )for a
1-unit increase in SBP assuming that LINSUL is the same, and 2.52 times (or
152%higher )for a 1-unit increase in LINSULassuming that SBP is the same.
From the table, SBP and LINSUL are related significantly to the diabeticstatus in all models with different link functions.
SAS and SPSS can also be used for the other two link functions:the inverse
ofthecumulativestandardnormaldistributionandthecomplementarylog-log 423
link functions introduced in Section 14.2.3. Table 14.18 includes the results
frommodelswiththesetwolinkfunctions.Theresultsareverysimilartothoseobtained using the logit link function.
The following SAS, SPSS, and BMDP codes can be used to obtain the
results in Table 14.18.
SAS code:
data w1;
infile ‘c: /p33ex14d2d6.dat’ missover;
input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;
run;title ‘‘Ordinal regression model with logic link function’’;proc logistic data /p58w1 descending;
model dms /p58age sex sbp dbp lacr hdl linsul smoke
/ selection /p58s lackfit link /p58logit;
run;title ‘‘Ordinal regression model with inverse normal link function‘;proc logistic data /p58w1 descending;
model dms /p58age sex sbp dbp lacr hdl linsul smoke
/ selection /p58s lackfit link /p58probit;
run;title ‘‘Ordinal regression model with complementary log-log link function’’;proc logistic data /p58w1 descending;
model dms /p58age sex sbp dbp lacr hdl linsul smoke
/ selection /p58s lackfit link /p58cloglog;
run;
SPSS code:
data list file /p58‘c:/p33ex14d2d6.dat’ free
/ age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn.
Compute y /p584-dms.
plum y with sbp linsul
/link /p58logit
/print /p58fit history parameter.
plum y with sbp linsul
/link /p58probit
/print /p58fit history parameter.
plum y with sbp linsul
/link /p58cloglog
/print /p58fit history parameter.
BMDP PR code for the logit link function only:
/input file /p58‘c:/p33ex14d2d6.dat’ .
variables /p5812.
format /p58free.424
/variable names /p58age, ageg, sex, sbp, dbp, lacr, hdl, linsul, smoke,
dms, dm, sn.Use/p58age, sex to smoke.
/transform y /p584-dms.
/group codes (y)/p581, 2, 3.
Names (y)/p58DM, IFG, NFG.
/regress depend /p58y.
Level /p583.
Type /p58ord.
Interval /p58age, sex to smoke.
enter /p58.05, .05.
remove /p58/p58.05, .05.
/print cell /p58used.
/end
Note that the model for ordinal polychotomous outcomes in BMDP PR is
defined as
logP(Y/p71/p57k/p34x/p71)
1/p57P(Y/p71/p57k/p34x/p71)/p58/afii9825/p23/p73/p59/p78/p26
/p72/p14/p16/afii9826/p18/p72x/p71/p72/p58u/p23/p73/p71k/p581, 2,...,m/p571
Compared with (14.3.5 ),/afii9825/p23/p73/p58/p57a/p73,k/p581, 2,...,m/p571;/afii9826/p18/p72/p58/p57b/p72,j/p581,
2,...,p.
Bibliographical Remarks
The linear logistic regression method is discussed extensively in Cox (1970 ),
Cox and Snell (1989 ), Collett (1991 ), Kleinbaum (1994 ), and Hosmer and
Lemeshow (2000 ). Cox’s book provides the theoretical background, and
Hosmer and Lemeshow discuss broad application of the method, includingmodel-building strategies and interpretation and presentation of analysisresults. In addition to the papers and books cited in this chapter, other workson the subject include Anderson (1972 ), Mantel (1973 ), Prentice (1976 ),
Prentice and Pyke (1979 ), Holford et al. (1978 ), and Breslow and Day (1980 ).
Applications of the logistic regression model can easily be found in variousbiomedical journals.
EXERCISES
14.1Consider the study presented in Example 3.5 and the data for the 40
patients in Table 3.10.(a)Construct a summary table similar to Table 3.11.
(b)Construct a table similar to Table 3.12.
(c)Usethechi-squaretesttodetectanydifferencesinretinopathyrates
among the subgroups obtained in part (b). 425
(d)On the basis of these 40 patients, identify the most important risk
factors using a linear logistic regression method.
14.2Consider the data for the 33 hypernephroma patients given in Exercise
Table 3.1. Let ‘‘response’’ be defined as stable, partial response, orcomplete response.(a)Compare each of the five skin test results of the responders with
those of the nonresponders.
(b)Use alinearlogisticregressionmethodtoidentifythemostimport-
ant risk factors related to response.
(i)Consider the five skin tests only.
(ii)Consider age, gender, and the five skin tests.
14.3Consider all nine risk variables (age, gender, family history of
melanoma, and six skin tests )in Exercise 3.3 and Exercise Table 3.3.
Identify the most important prognostic factors that are related toremission. Use both univariate and multivariate methods.
14.4Consider the data of 58 hypernephroma patients given in Exercise
Table 3.2. Apply the logistic regression method to response (defined as
complete response, partial response, or stable disease ). Include gender,
age, nephrectomy treatment, lung metastasis, and bone metastasis asindependent variables.
(a)Identify the most significant independent variables.
(b)Obtain estimates of odds ratios and confidence intervals when
applicable.
14.5Consider the case where there is one continuous independent variable
X/p16. Show that the log odds ratio for X/p16/p58x/p16/p59mversus X/p16/p58x/p16is
mb/p16, where b/p16is the logistic regression coefficient.
14.6Using the data in Table 12.4, define the index function CVD as
CVD /p581i fd g /p461, and CVD /p580 otherwise, and fit a logistic re-
gression model for CVD by using the stepwise selection method toselectrisk factorsamongthe samefactorsas thosenotedatthe bottomof Table 12.7. Compare the results obtained with those in Table 12.7.
14.7Assumingthat P(apersonissampled /p34y,x)/p58P(apersonissampled /p34y),
that is, the sampling probability is independent of the risk factors x,
derive (14.2.15 ).
14.8By using (14.2.14 )and (14.2.1 ), show that (14.2.20 )reducesto (14.2.21 ).
14.9Derive (14.3.2 ).426
14.10Consider the data in Table 12.4. Fit the generalized logistic regression
modelin (14.3.1 )for DG with covariates AGE, SEX, LACR, and LTG
by using the SAS CATMOD, SPSS NOMREG, or BMDP PRprocedure. Select risk factors among those noted at the bottom ofTable 12.7 using the stepwise selection method in the BMDP PRprocedure. Compare the results with those given in Table 13.5.
14.11UsingthesamenotationanddataasinTable14.11, (1)fit theoutcome
variable Ywiththegeneralizedlogisticregressionmodelin (14.3.1 )with
SEXasthecovariate; (2)fitalogisticregressionforthebinaryoutcome
DM versus NFG, with SEX as the covariate, by using the data fromDM and NFG participants only; (3)fit a logistic regression for the
binary outcome IFG versus NFG, with SEX as the covariate, by usingthe data from IFG and NFG participants only; (4)compare the
coefficients obtained from (2)and (3)with the coefficients obtained
from (1), and (5)report what you have found.
14.12Perform the same analyses as in Exercise 14.11 but use SBP as the
covariate, and discuss your findings. 427
APPENDIX A
Newton- -Raphson Method
The Newton —Raphson method (Ralston and Wilf, 1967;Carnahan et al., 1969 )
is a numerical iterative procedure that can be used to solve nonlinearequations. An iterative procedure is a technique of successive approximations,and each approximation is called an iteration. If the successive approximations
approach the solution very closely, we say that the iterations converge. The
maximum likelihood estimates of various parameters and coefficients discussedin Chapters 7, 9, and 11 to 14 can be obtained by using the Newton —Raphson
method. In this appendix we discuss and illustrate the use of this method, firstconsidering a single nonlinear equation and then a set of nonlinear equations.
Let f(x)/p580 be the equation to be solved for x. The Newton —Raphson
method requires an initial estimate of x, say x/p24/p15, such that f(x/p24/p15) is close to zero
preferably, and then the first approximate iteration is given by
x/p24/p16/p58x/p24/p15/p57f(x/p24/p15)
f/p30(x/p24/p15)
(A.1)
where f/p30(x/p24/p15) is the first derivative of f(x) evaluated at x/p58x/p24/p15. In general, the
(k/p591)th iteration or approximation is given by
x/p24/p73/p62/p16/p58x/p24/p73/p57f(x/p73)
f/p30(x/p73)(A.2)
where f/p30(x/p24/p73)is the first derivative of f(x) evaluated at x/p58x/p24/p73. The iteration
terminates at the kth iteration if f(x/p24/p73)is close enough to zero or the difference
between x/p24/p73and x/p24/p73/p92/p16is negligible. The stopping rule is rather subjective.
Acceptable rules are that f(x/p24/p73)o r d/p58x/p24/p73/p57x/p24/p73/p92/p16is in the neighborhood of
10/p92/p21or 10 /p92/p22.
Example A.1 Consider the function
f(x)/p58x/p18/p57 x/p592
428
Figure A.1 Graphical presentation of the Newton —Raphson method for Example A.1.
We wish to find the value of xsuch that f(x)/p580 by the Newton —Raphson
method. The first derivative of f(x)i s
f/p30(x)/p583x/p17/p57 1
Since f(/p571)/p582 and f(/p572)/p58/p57 4, graphically (Figure A.1 ), we see that the
curve cuts through the xaxis [ f(x)/p580] between /p571 and /p572. This gives us a
good hint of an initial value of x. Suppose that we begin with x/p24/p15/p58/p57 1;
f(x/p24/p15)/p582 and f/p30(x/p24/p15)/p582. Thus, the first iteration, following (A.1), gives
x/p24/p16/p58/p57 1/p572
2/p58/p57 2
and f(x/p24/p16)/p58/p57 4 and f/p30(x/p24/p16)/p5811. Following (A.2), we obtain the following:
Second iteration:
x/p24/p17/p58/p57 2/p594
11/p58/p57 1.6364
f(x/p24/p17)/p58/p57 0.7456 f/p30(x/p24/p17)/p587.0334 - 429
Third iteration:
x/p24/p18/p58/p57 1.6364/p590.7456
7.0334/p58/p57 1.5304
f(x/p24/p18)/p58/p57 0.054 f/p30(x/p24/p18)/p586.0264
Fourth iteration:
x/p24/p19/p58/p57 1.5304/p590.054
6.0264/p58/p57 1.52144
f(x/p24/p19)/p58/p57 0.00036 f/p30(x/p24/p19)/p585.9443
Fifth iteration:
x/p24/p20/p58/p57 1.52144 /p590.00036
5.9443/p58/p57 1.52138
f(x/p24/p20)/p580.0000017
At the fifth iteration, for x/p58/p57 1.52138, f(x) is very close to zero. If the
stopping rule is that f(x)/p4510/p92/p21, the iterative procedure would terminate after
the fifth iteration and x/p58/p57 1.52138 is the root of the equation x/p18/p57 x/p592/p580.
Figure A.1 gives the graphical presentation of f(x) and the iteration.
It should be noted that the Newton —Raphson method can only find the real
roots of an equation. The equation x/p18/p57 x/p592/p580 has only one real root, as
shown in Figure A.1;the other two are complex roots.
The Newton —Raphson method can be extended to solve a system of
equations with more than one unknown. Suppose that we wish to find valuesofx/p16,x/p17,...,x/p78such that
f/p16(x/p16,...,x/p78)/p580
f/p17(x/p16,...,x/p78)/p580
/p36
f/p78(x/p16,...,x/p78)/p580
Let a/p71/p72be the partial derivative of f/p71with respect to x/p72;that is, a/p71/p72/p58/p42f/p71//p42x/p72.430 -
The matrix
J/p58a/p16/p16/p37 a/p16/p78
a/p17/p16/p37 a/p17/p78
/p36/p36
a/p78/p16/p37 a/p78/p78
is called the Jacobian matrix . Let the inverse of J, denoted by J/p92/p16,b e
J/p92/p16 /p58b/p16/p16/p37 b/p16/p78
b/p17/p16/p37 b/p17/p78
/p36/p36
b/p78/p16/p37 b/p78/p78
Let x/p73/p16,x/p73/p17,...,x/p73/p78be the approximate root at the kth iteration;let f/p73/p16,...,f/p73/p78be the corresponding values of the functions f/p16,...,f/p78, that is,
f/p73/p16/p58f/p16(x/p73/p16,...,x/p73/p78)
/p36
f/p73/p78/p58f/p78(x/p73/p16,...,x/p73/p78)
and let b/p73/p71/p72be the ijth element of J/p92/p16evaluated at x/p73/p16,...,x/p73/p78. Then the next
approximation is given by
x/p73/p62/p16/p16/p58x/p73/p16/p57(b/p73/p16/p16f/p73/p16/p59b/p73/p16/p17f/p73/p17/p59/p37/p59b/p73/p16/p78f/p73/p78)
x/p73/p62/p16/p17/p58x/p73/p17/p57(b/p73/p17/p16f/p73/p16/p59b/p73/p17/p17f/p73/p17/p59/p37/p59b/p73/p17/p78f/p73/p78)( A.3)
/p36
x/p73/p62/p16/p78/p58x/p73/p78/p57(b/p73/p78/p16f/p73/p16/p59b/p73/p78/p17f/p73/p17/p59/p37/p59b/p73/p78/p78f/p73/p78)
The iterative procedure begins with a preselected initial approximate x/p15/p16,
x/p15/p17,...,x/p15/p78, proceeds following (A.3), and terminates either when f/p16,f/p17,...,f/p78are close enough to zero or when differences in the xvalues at two consecutive
iterations are negligible.
Example A.2 Suppose that we wish to find the value of x/p16and x/p17such that
x/p17/p16/p59x/p16x/p17/p572x/p16/p571/p580 x/p18/p16/p57x/p16/p59x/p17/p572/p580 - 431
In this case, p/p582:
f/p16/p58x/p17/p16/p59x/p16x/p17/p572x/p16/p571 f/p17/p58x/p18/p16/p57x/p16/p59x/p17/p572
Since /p42f/p16//p42x/p16/p582x/p16/p59x/p17/p572,/p42f/p16//p42x/p17/p58x/p16,/p42f/p17//p42x/p16/p583x/p17/p16/p571, and /p42f/p17//p42x/p17/p581,
the Jacobian matrix is
J/p58/p32x/p16/p59x/p17/p572
3x/p17/p16/p571x/p161/p4(A.4)
Let the initial estimates be x/p15/p16/p580,x/p15/p17/p581,f/p15/p16/p58/p57 1, and f/p15/p17/p58/p57 1:
J/p58/p3/p5710
/p5711/p4J/p92/p16 /p58 /p3/p5710
/p5711/p4
Iteration 1. Following (A.3), we obtain
x/p16/p16/p580/p57[(/p571)(/p571)/p590(/p571)]/p58/p57 1 x/p16/p17/p581/p57[(/p571)(/p571)/p591(/p571)]/p581
With these values, f/p16/p16/p581,f/p16/p17/p58/p57 1, and
J/p58/p3/p573/p571
21 /p4J/p92/p16 /p58 /p3/p571/p571
23 /p4
Iteration 2. From (A.3)we obtain
x/p17/p16/p58/p57 1/p57[(/p571)(1)/p59(/p571)(/p571)]/p58/p57 1 x/p17/p17/p581/p57[(2)(1) /p59(3)(/p571)]/p582
With these values, f/p17/p16/p580 and f/p17/p17/p580. Therefore, the iteration procedure
terminates and the solution of the two simultaneous equations is x/p16/p58/p57 1,
x/p17/p582.
The number of iterations required depends strongly on the initial values
chosen. In Example A.2, if we use x/p15/p16/p580,x/p15/p17/p580, it requires about 11 iterations
to find the solution. Interested readers may try it as an exercise.432 -
APPENDIX B
Statistical Tables
433
Table B-1 Normal Curve Areas
Source:Abridgedfrom Table 1 of Statistical Tables and Formulas ,by A. Hald, JohnWiley &Son s,
1952. Reproduced by permissionof JohnWiley & Son s.
434
Table B-2 Percentage Points of the /afii98512-Distribution
Source:‘‘Tables of the Percentage Points of the /afii9851/p17-Distribution,’’ by Catherine M. Thompson,
Biometrika , Vol. 32, pp. 188 —189(1941 ). Reproduced by permissionof the editor of Biometrika.
435
Table B-3 5 %Points of the F-Distribution
436
oa RT|giabee
eaa aT Lee
437
Table B-3 2.5 %Points of the F-Distribution
438
a)835 E2285 S385 SES82 SEESE ERERE EREE
3)-283 S822 22552 25558 RESEE S8ELE SEg83
§8868 C2555 SERAE RSAES Skee SERRE E5523
¢|<S8e SEEES ZESRE £2298 BORES BECES Bezeg|
|S RSE StF SERS GERI ENING ESkae GEES
Z|5]£38SE52§ S208 S658 FREES SgnzR ROLES5[8]gageSec8SSekE22R28CERESESEEEGezee
£|.)829SESEE G2E23 £LG82 SESE Loans foggy
E[k| Bpre Coc8S Bass GEnkS SREL8 ALES GEER
i puueib ae
|gueFUG8ESEGREELERELEGGEZEEEBRESTBee SEE05 SSEK3 SS555 S8083
o|s8o2 258% EESES SELIG REESE SESE EuEas
gers SS598 BEE55 FLERE RESES SERSR EASES
VATerwmenusS2unySEESRANEERARERASSESTamang Fo moped foPaaog
439
Table B-3 1 %Points of the F-Distribution
440
|sez se8s2 BEES ©S288 S28E8 FE
2
i] Sonse SEQMsees ay 22seoes
.
|| BRE SRESS FREES GENRE Skee ded ARES
oldEe
ageSP2352gHESGEFRESESLVRPSSz! |HaeESSEUEEEEE282EDEEESERREES
3
evs RESSZ EEGZE EPSES SRESE SEREY SESRE
:
|~82% 22522 SEES SELES ZEERE TEES!
iB
[7S SReeee
441
Table B-3 0.5 %Points of the F-Distribution
442
Source:‘‘Tables of Percentage Points of the Inverted Beta (F)Distribution,’’ by Maxine Merrington and Catheri ne M. Thompson,
Biometrika,V o l .3 3 ,p p .7 3 —88(1943 ). Reproduced by permissionof the editor of Biometrika.
443
Table B-4 Upper Tail Probabilities for the Null Distribution of the Kruskal--Wallis H
Statistic: k/p583,n1/p581(1)5, n2/p58n1(1)5, 2/p45n3/p58n2(1)5
444
Table B-4 ( continued )
445
Table B-4 ( continued )
446
Table B-4 ( continued )
447
Table B-4 ( continued )
448
Table B-4 ( continued )
449
Table B-4 ( continued )
450
Table B-4 ( continued )
451
Table B-4 ( continued )
452
Table B-4 ( continued )
453
Table B-4 ( continued )
454
Table B-4 ( continued )
455
Table B-4 ( continued )
456
Table B-4 ( continued )
457
Table B-4 ( continued )
Source:Table F of A Nonparametric Introduction to Statistics , by C. H. Kraft and C van Eedan,
Macmillan, New York, 1968. Reproduced by permission of the Macmillan Publishing Company.
458
Table B-5 Selected Critical Values for All Treatments: Multiple Comparisons Based on
Kruskal--Wallis Rank Sums
Source:‘‘Rank Sum Multiple Comparisons in One- and Two-Way Classification,’’ by B. J.
McDonald and W. A. Thompson, Biometrika , Vol. 54, pp. 487 —497 (1967 ). Reproduced by
permissionof the editor of Biometrika . The starred values are from ‘‘Distribution-Free Multiple
Comparisons,’’ Ph.D. thesis (1963 ), P. Nemenyi, Princeton University, with permission of the
author. 459
Table B-6 Selected Critical Values for the Range of kIndependent N(0, 1) Variables:
k/p582(1)20(2)40(10)100
For a given kand/afii9825, the tabled entry is q(/afii9825,k,/p45).
Source:‘‘TableofRange and StudentizedRange,’’by H.L. Harter, Ann. Math. Statist. , Vol.31,pp.
1122—1147 (1960 ). Reproduced by permissionof the editor of the Annals of Mathematical
Statistics.
460
Table B-7Percentage Points of the t-Distribution
Source:‘‘Table of Percentage Points of the t-Distribution,’’ by Maxine Merrington, Biometrika ,
Vol. 32, p. 300 (1941 ). Reproduced by permissionof the editor of Biometrika .
461
Table B-8 Coefficients ( aiandbi) of the Best Estimates of the Mean ( /afii9839) and Standard
Deviation ( /afii9846) in Censored Samples Up to n/p5820 from A Normal Population
462
Table B-8 ( continued )
463
Table B-8 ( continued )
464
Table B-8 ( continued )
465
Table B-8 ( continued )
466
2|8
eae
iuHaE
=[88 53 58 RS
3jiga88588
iB HEEE
|:HERB EHHs\RR EH
;URERNREHH
s(QHHRREEHES
PEER GEEEESs
;HHBRHHEER AE
SHRHNEERHE GEE
SHRERHEHRHBUEE
sHHEGHERHREERH ES PEEeee SeeREae
/HERERHRERHEEHAS FESReeeeenaeas i]tleaes
467
Table B-8 ( continued )
468
469
Table B-8 ( continued )
470
=le285?
5(88232
aleg atgeceefce22as
PRS EGEEER
471
Table B-8 ( continued )
Source:‘‘Estimation of Location and Scale Parameters by Order St atistics from Singly and Doubly Censored Samples, Parts I a nd II,’’ by A. E.
S a r h a na n dB .G .G r e e n b e r g ,Ann. Math. Statist. , Vol. 27, pp. 427 —451 (1956 ). Reproduced by permissionof the editor of the Annals of
Mathematical Statistics.
472
Table B-9 Variances and Covariances of the Best Linear Est imates of the Mean ( /afii9839/p24) and Standard Deviation ( /afii9846/p24)f o r
Censored Samples Up to Size 20 from a Normal Population
473
Table B-9 ( continued )
Source:Up to n/p5815 of this table is reproduced from A. E. Sarhanan d B. G. G reen berg, ‘‘Estimationof Locationan d Scale Parameters by O rder Statistics
from Singly and Censored Samples, Parts I and II,’’ Ann. Math. Statist ., Vol. 27, pp. 427 —451(1956),a ndV o l .2 9 ,p p .7 9 —105(1958), with permissionof the
editor of the Annals of Mathematical Statistics . The rest of the table is produced from A. E. Sarhan and B.G. Greenberg, ‘‘Estimation of Location and Scale
Parameters by Order Statistics from Singly and D oubly Censored Samples, Part III,’’ Tech. Rep. 4-OOR, Proje ct 1597, U.S. Army Research Office.
474
Table B-10 1 /(1/p57R) and /afii9828/p24for the Estimation of the Parameters of the Gamma
Distribution When There Are No Censored Observations
Source:‘‘Estimationof Parameters of the Gamma DistributionUsin g Order Statistics,’’ by M. B.
Wilk, R. Gnanadesikan, and Marilyn J. Huyett, Biometrika , Vol. 49, pp. 525 —545 (1962 ).
Reproduced by permissionof the editor of Biometrika .
475
Table B-11 /afii9828/p24(P,S) and /afii9839/p24(P,S) for Various Values of n/r:n/r/p581.0
ForP/p450.52 read Sfrom the left-hand margin, and for P/p460.56 read Sfrom the right-hand
margin. Note that the figures in region 2 are printed in bold roman type and those in region 3 inbold italic type; the remainder of the table (outside of regions 2 and 3 )is region1.
476
Table B-11 ( continued )
477
Table B-11 ( continued )
478
Table B-11 ( continued )
479
Table B-11 ( continued )
480
Table B-11 ( continued )
481
Table B-11 ( continued )
482
Table B-11 ( continued )
483
Table B-11 ( continued )
484
Table B-11 ( continued )
Source:‘‘Estimationof Parameters of the Gamma DistributionUsin g Order Statistics,’’ by M. B.
Wilk, R. Gnanadesikan, and Marilyn J. Huyett, Biometrika , Vol. 49, pp. 525 —545 (1962 ).
Reproduced by permissionof the editor of Biometrika.
485
Table B-12 Percentage Points l/afii9825Such That P(/afii9828/p241//afii9828/p242/p58l/afii9825)/p581/p57/afii9825
Source:‘‘Two Sample Test in the Weibull Distribution,’’ by D. R. Thoman and L. J. Bain,
Technometrics , Vol. 11, pp. 805 —815 (1969 ). Reproduced by permissionof the editor of Techno-
metrics.
486
Table B-13 Percentage Points z/afii9825Such That P(G/p58z/afii9825)/p581/p57/afii9825
Source:‘‘Two Sample Test in the Weibull Distribution,’’ by D. R. Thoman and L. J. Bain,
Technometrics , Vol. 11, pp. 805 —815 (1969 ). Reproduced by permissionof the editor of Techno-
metrics.
487
References
Aaronson, K. D., Schwartz, J. S., Chen, T. M., Wong, K. L., Goin, J. E, and Mancini,D.
M.(1997 ). Development and Prospective Validation of a Clinical Index to Predict
Survival in Ambulatory Patients Referred for Cardiac Transplant Evaluation.Circulation ,95, 2660—2667.
Abramowitz, M., and Stegun, I. A. (1964 ).Handbook of Mathematical Functions with
Formulas, Graphs, and Mathematical Tables. Applied Mathematics Series 55.
National Bureau of Standards, Washington, DC.
Afifi, A. A., and Clark, V. (1990 ).Computer-Aided Multivariate Analysis , 2nd ed.
Lifetime Learning Publications, Belmont, CA.
Agresti, A. (1990 ).Categorical Data Analysis. Wiley, New York.
Aitchison, J. (1970 ). Statistical Problems of Treatment Allocation. Journal of the Royal
Statistical Society, Series A ,133, 206—238
Aitchison, J., and Brown, J. A. C. (1957 ).The Lognormal Distribution. Cambridge
University Press, Cambridge.
Aitchison,J., and Silvey, S. D. (1957 ). The Generalizationof ProbitAnalysisto the Case
of Multiple Responses. Biometrika ,44, 131—140.
Aitkin, M., Laird, N., and Francis, B. (1983 ). A Reanalysis of the Stanford Heart
Transplant Data (with discussion ).Journal of the American Statistical Association ,
78, 264—292.
Akaike, H. (1969 ). Fitting Autoregressive Models for Prediction. Annals of the Institute
of Statistical Mathematics ,21, 243—247.
Akaike, H. (1974 ). A New Look at the Statistical Model Identification. IEEE Transac-
tions on Automatic Control, AC-19, 716—723.
Albert, I. J. (2000 ). The Use of Frailty Models in Genetic Studies: Application to the
Relationship between End-Stage Renal Failure and Mutation Type in AlportSyndrome. European Community Alport Syndrome Concerted Action Group(ECASCA ).Journal of Epidemiology and Biostatistics ,5(3), 169—175.
Albertson, P. C., Hanley, J. A., Gleason, D. F., and Barry, M. J. (1998 ). Competing Risk
Analysis of Men Aged 55 to 74 Years at Diagnosis Managed Conservatively for
ClinicallyLocalizedProstate Cancer. Journal of the American Medical Association ,
280(11),97 5—980.
488
Alioum, A., and Commenges, D. (1996 ). A Proportional Hazards Model for Arbitrarily
Censored and Truncated Data. Biometrics ,52, 512—524.
Altshuler, B. (1970 ). Theory for Measurement of Competing Risks in Animal Experi-
ments. Mathematical Biosciences ,6,1—11.
Ananth, C. V., and Kleinbaum, D. G. (1977 ). Regression Models for Ordinal Data: A
Review of Methods and Application. InternationalJournalof Epidemiology ,26,
1323—1333.
Andersen,P. K. (1982 ). Testing Goodness of Fit of Cox’sRegression Model. Biometrics ,
38,6 7—77.
Andersen, P. K. (1992 ). Repeated Assessment of Risk Factors in Survival Analysis.
Statistical Methods in Medical Research ,1,2 97—315
Andersen, P. K., Borgan, O., Gill, R. D., and Keiding, N. (1993 ).Statistical Models
Based on Counting Processes. Springer-Verlag, New York.
Andersen, P. K., and Gill, R. D. (1982 ). Cox’s Regression Model Counting Process: A
Large Sample Study. Annals of Statistics ,10, 1100—1120
Anderson,J. A. (1972 ). SeparateSampleLogisticDiscrimination, Biometrika ,59,1 9—35.
Andrews,D. F., and Herzberg,A. M. (1985 ).Data: A Collection of Problems from Many
Fields for the Student and Research Worker . Springer-Verlag, New York.
ARIC Investigators. (1989 ). The Atherosclerosis Risk in Communities (ARIC )Study:
Design and Objectives. American Journal of Epidemiology ,129, 687—702.
Arjas, E. (1988 ). A Graphical Method for Assessing Goodness of Fit in Cox’s
Proportional Hazards Model. Journal of the American Statistical Association ,83,
204—212.
Armitage, P. (1959 ). The Comparison of Survival Curves. Journal of the Royal
Statistical Society, Series A ,122, 279—300.
Armitage, P. (1971 ).Statistical Methods in Medical Research. Blackwell Scientific
Publications, Oxford.
Armitage, P. (1981 ). Importance of Prognostic Factors in the Analysis of Data from
Clinical Trials. Controlled Clinical Trials ,1, 347—353.
Armitage, P., and Gehan, E. A. (1974 ). Statistical Methods for the Identification and
Use of Prognostic Factors. International Journal of Cancer ,13,1 6—35.
Asal, N. R., Geyer, J. R., Risser, D. R., Lee, E. T., Kadamani, S., and Cherng, N. (1988a ).
Risk Factors in Renal Cell Carcinoma, Part I. Methodology, Demographics,Tobacco, Beverage and Obesity. Cancer Detection and Prevention ,11, 359—377.
Barnard, G. A. (1963 ). Some Aspects of the Fiducial Argument. Journal of the Royal
Statistical Society, Series B ,34, 216—217.
Bartholomew, D. J. (1957 ). A Problem in Life Testing. Journal of the American
Statistical Association ,52, 350—355.
Bartholomew, D. J. (1963 ). The Sampling Distribution of an Estimate Arising in Life
Testing. Technometrics ,5, 361—374.
Baumgartner, R. N., Roche, A. F., et al. (1987 ). Fatness and Fat Patterns: Associations
withPlasma Lipidsand Blood Pressure in Adults,18 to 57 Years of Age. American
Journal of Epidemiology ,126, 614—628. 489
Beale, E. M. L., Kendall, M. G., and Mann, D. W. (1976 ). The Discarding of Variable
in Multivariate Analysis. Biometrika ,54, 357—366.
Berkson, J. (1942 ). The Calculation of Survival Rates, in Carcinoma and Other
Malignant Lesions of the Stomach , edited by W. Walters, H. K. Gray, and J. T.
Priestley. W.B. Saunders, Philadelphia.
Berkson, J., and Gage, R. R. (1950 ). Calculation of Survival Rates for Cancer.
Proceedings of Staff Meetings, Mayo Clinic ,25, 250.
Birnbaum, Z. W., and Saunders, S. C. (1958 ). A Statistical Model for Life-Length of
Materials. Journal of the American Statistical Association ,53, 151—160.
Blackstone, E. H., and Lytle, B. W. (2000 ). Competing Risks after Coronary Bypass
Surgery: The Influence of Death on Reintervention. Journal of Thoracic and
Cardiovascular Surgery ,119(6), 1221—1230.
Bliwise, D. L., Kutner, N. G., Zhang, R., and Parker, K. P. (2002 ). Survival by Time of
Day of Hemodialysis in an Elderly Cohort. Journal of the American Medical
Association ,286(21), 2690—2694.
Boag, J. W. (1949 ). Maximum Likelihood Estimates of Proportion of Patients Cured
by Cancer Therapy. Journal of the Royal Statistical Society, Series B ,11, 15.
Bolard, P., Quantin, C. P., Esteve, J., Faivre, J., and Abrahamowicz, M. (2001 ).
Modeling Time-Dependent Hazard Ratios in Relative Survival: Application toColon Cancer. Journal of Clinical Epidemiology ,54(10)986—996.
Bonadonna, G., et al. (1976 ). Combination Chemotherapy as an Adjuvant Treatment
in Operable Breast Cancer. New England Journal of Medicine ,294, 405—410.
Brancato, G., Pezzotti, P., Rapiti, E., Perucci, C. A., Abeni, D., Babbalacchio, A., and
Rezza, G. (1997 ). Multiple Imputation Method for Estimating Incidence of HIV
Infection: The Multicenter Prospective HIV Study. International Journal of Epi-
demiology ,26(5), 1107—1114.
Breslow, N. (1970 ). A Generalized Kruskal —Wallis Test for Comparing KSamples
Subject to Unequal Pattern of Censorship. Biometrika ,57, 579—594.
Breslow, N. (1974 ). Covariance Analysis of Survival Data under the Proportional
Hazards Model. International Statistical Review ,43,4 3—54.
Breslow, N. E. (1975 ). Analysis of Survival Data under the Proportional Hazards
Model. International Statistical Review ,43,4 5—48.
Breslow, N. E., and Crowley, J. (1974 ). A Large Sample Study of the Life Table and
Product Limit Estimates under Random Censoring. Annals of Statistics ,2,
437—453.
Breslow, N. E., and Day, N. E. (1980 ).Statistical Methods in Cancer Research ,
Vol. 1, The Analysis of Case-Control Studies . International Agency for Research
on Cancer, Lyon, France.
Breslow,N., and Powers, W. (1978 ). Are ThereTwo LogisticRegressionsfor Retrospec-
tive Studies? Biometrics ,34, 100—105.
Breslow, N. E., Day, N. E., Halvorsen, K. T., Prentice, R. L., and Sabai, C. (1978 ).
Estimationof Multiple Relative Risk Functions in Matched Case-Control Studies.
American Journal of Epidemiology ,108,2 99—307.
Broadbent, S. (1958 ). Simple Mortality Rates. Journal of Applied Statistics ,7, 86.490
Broderick, A., Mori, M., Nettleman, M. D., Streed, S. A., and Wenzel, R. P. (1990 ).
Nosocomial Infections: Validation of Surveillance and Computer Modeling toIdentify Patients at Risk. American Journal of Epidemiology ,131, 734—742.
Brookmeyer, R., and Goedert, J. J. (1989 ). Censoring in an Epidemic with an
Application to Hemophilia-Associated AIDS. Biometrics ,45, 325—335.
Brown, B. W., and Hollander, M. (1977 ).Statistics: A Biomedical Introduction . Wiley,
New York.
Brown, C. C. (1982 ). On a Goodness-of-FitTest for the Logistic Model Based on Score
Statistics. Communications in Statistics ,11, 1087—1105.
Brown, G. W., and Flood, M. M. (1947 ). Tumbler Mortality. Journal of the American
Statistical Association ,42, 562—574.
Burdette, W. J., and Gehan, E. A. (1970 ).PlanningandAnalysisof ClinicalStudies.
Charles C. Thomas, Springfield, IL.
Buzdar, A. U., Gutterman, J. U., Blumehscein, G. R., Hortobagiji, G. H., Tashima, C.
K., Smith, T. L, Hersh, E. M., Freiriech, E. J., and Gehan, E. A. (1978 ). Intensive
Postoperative Chemoimmunotherapy for Patients with Stage II and Stage IIIBreast Cancer. Cancer,41, 1064—1075.
Byar, D. P. (1974 ). Selecting Optimum Treatment in Clinical Trials Using Covariate
Information. Presented at the 1974 Annual Meeting of the American Statistical
Association, August 28.
Byar, D. P. (1980 ). The Veterans Administration Study of Chemoprophylaxis for
Recurrent Stage I Bladder Tumors: Comparisons of Placebo, Pyridoxine, andTopical Thiotepa, In Bladder Tumors and Other Topics in Urological Oncology ,
edited by M. Pavone-Macaluso, P. H. Smith, and F. Edsmyn. Plenum Press, NewYork, pp. 363—370.
Byar,D.P.,Huse,R., andBailar,J. C.III, andthe VeteransAdministrationCooperative
Urological Research Group (1974 ). An Exponential Model Relating Censored
Survival Data and Concomitant Information for Prostatic Cancer Patients.Journal of the National Cancer Institute ,52, 321—326.
Byers, R. H. Jr., Morgan, W. M., Darrow, W. W., Doll, L., Jaffe, H. W., Rutherford, G.,
Hessol, N., and O’Malley, P. M., (1988 ). Estimating AIDS Infection Rates in the
San Francisco Cohort. AIDS,2(3), 207—210.
Carbone, P., Kellerhouse, L., and Gehan, E. (1967 ). Plasmacytic Myeloma: A Study of
the Relationship of Survival to Various Clinical Manifestations and Anomalous
Protein Type in 112 Patients. American Journal of Medicine ,42,93 7—948.
Carnahan, B., Luther, H. A., and Wilkes, J. O. (1969 ).Applied Numerical Methods.
Wiley, New York.
Carter, S. K., Oleg, S., and Slavik, M. (1977 ). Phase I Clinical Trials, in Methods of
Development of New Anticancer Drugs . National Cancer Institute Monograph 45.
U.S. Department of Health, Education, and Welfare Publication (NIH )76—1037.
National Cancer Institute, Bethesda, MD.
Chernoff, H., and Leiberman, G. J. (1954 ). Use of Normal Probability Paper. Journal
of the American Statistical Association ,49, 778—785.
Chiang, C. L. (1961 ). Standard Error of the Age-Adjusted Death Rate. Vital Statistics:
Special Reports, Selected Studies ,47, 9. U.S. Department of Health, Education,
and Welfare, Washington, DC. 491
Chiang, C. L. (1968 ).Introduction to Stochastic Processes in Biostatistics. Wiley, New
York.
Chiasson, M. A., Stoneburner, R. L., et al. (1990 ). Risk Factors for Human Immunode-
ficiency Virus Type 1 (HIV-1 )Infection in Patients at a Sexually Transmitted
DiseaseClinicin NewYorkCity. American Journal of Epidemiology ,131,208—220.
Clayton, D., and Cuzick, J. (1985 ). The Em algorithm for Cox’s regression model using
GLIM.AppliedStatistics ,34, 148—156.
Cochran, W. G., and Cox, G. M. (1957 ).Experimental Designs , 2nd ed. Wiley, New
York.
Cohen, A. C., Jr. (1951 ). Estimating Parameters of Logarithmic-Normal Distributions
by Maximum Likelihood. Journal of the American Statistical Association ,46,
206—212.
Cohen, A. C., Jr. (1959 ). Simplified Estimators for the Normal Distribution When
Samples Are Singly Censored or Truncated. Technometrics ,1(3), 217—237.
Cohen, A. C., Jr. (1961 ). Table for Maximum Likelihood Estimates: Singly Truncated
and Singly Censored Samples. Technometrics ,3, 535—541.
Cohen, A. C., Jr. (1963 ). Progressively Censored Sample in Life Testing. Technometrics ,
5, 327—339.
Cohen, A. C., Jr. (1976 ). Progressively Censored Sampling in the Three Parameter
Log-Normal Distribution. Technometrics ,18.
Cohen, J., and Cohen, P. (1975 ).Applied Multiple Regression/Correlation Analysis for
the Behavioral Sciences. Lawrence Erlbaum Associates, Hillsdale, NJ.
Collett, D. (1991 ).Modelling Binary Data . Chapman & Hall, London.
Collins, J. A., Garner, J. B., Wilson, E. H., Wrixon, W., and Casper, R. F. (1984 ).A
Proportional Hazards Analysis of the Clinical Characteristics of Infertile Couples.American Journal of Obstetrics and Gynecology ,148, 527—532.
Connelly, R. R., Cutler, S. J., and Baylis, P. (1966 ). End Result in Cancer of the Lung:
ComparisonofMaleand FemalePatients. Journal of the National Cancer Institute ,
36, 277—287.
Cornfield, J. (1951 ). A Method of Estimating Comparative Rates from Clinical Data:
Applications to Cancer of the Lung, Breast and Cervix. Journal of the National
Cancer Institute ,11, 1269—1275.
Cornfield, J. (1956 ). A Statistical Problem Arising from Retrospective Studies, in
Proceedings of the 3rd Berkeley Symposium on Mathematical Statistics andProbability , Vol. 4, edited by J. Neyman. University of California Press, Berkeley,
CA, 135—148.
Cornfield, J. (1962 ). Joint Dependence of Risk of Coronary Heart Disease in Serum
Cholesterol and Systolic Blood Pressure: A Discriminant Function Analysis.Federation Proceedings ,21,5 8—61.
Correa, P., Pickle, L. W., Fortham, E., et al. (1983 ). Passive Smoking and Lung Cancer.
Lancet,2,5 95—597.
Cox, D. R. (1961 ). Tests of Separate Families of Hypotheses. Proc.FourthBerkeley
SymposiuminMathematicalStatistics , I, Berkeley: University of California Press,
105—123.492
Cox, D. R. (1962 ). Further Results on Tests of Separate Families of Hypotheses. J.R.
Stat.Soc.B ,24, 406—424.
Cox, D. R. (1953 ). Some Simple Tests for Poisson Variates. Biometrika ,40, 354—360.
Cox, D. R. (1959 ). The Analysis of Exponentially Distributed Life-Times with Two
Types of Failures. Journal of the Royal Statistical Society, Series B ,21, 411—421.
Cox, D. R. (1962 ).Renewal Theory . Methuen, London.
Cox, D. R. (1964 ). Some Applications of Exponentially Distributed Life-Times with
Two Types of Failures. Journal of the Royal Statistical Society, Series B ,26,
103—110.
Cox, D. R. (1970 ).Analysis of Binary Data. Methuen, London.
Cox, D. R. (1972 ). Regression Models and Life Tables. Journal of the Royal Statistical
Society, Series B ,34, 187—220.
Cox, D.R., and Hinkley,D. V. (1974 ).TheoreticStatistics ,Chapmanand Hall,London.
Cox, D. R., and Oakes, D. (1984 ).Analysis of Survival Data . Chapman & Hall, New
York.
Cox, D. R., and Snell, E. J. (1968 ). A General Definition of Residuals. Journal of the
Royal Statistical Society, Series B ,30, 248—275.
Cox, D. R., and Snell, E. J. (1989 ).The Analysis of Binary Data, 2nd ed . Chapman &
Hall, London.
Crawford, S. L., Tennstedt,S. L., and McKinlay, J. B. (1995 ). A Comparison of Analytic
Methods for Non-random Missingness of Outcome Data. J.ClinEpidemiol ,48,
209—219.
Crist, W., Boyett, J., and Jackson, J., et al. (1989 ). Prognostic Importance of the
Pre-B-CellImmunophenotypeand Other PresentingFeatures in B-Lineage Child-hood Acute Lymphoblastic Leukemia: A Pediatric Oncology Group Study. Blood,
74, 1252—1259.
Crowley,J.,and Hu,M. (1977 ).CovarianceAnalysisofHeartTransplantSurvivalData.
Journal of the American Statistical Association ,72,2 7—36.
Crowley, J., and Thomas, D. R. (1975 ). Large Sample Theory for the Log Rank Test.
Technical Report 415 . Department of Statistics, University of Wisconsin, Madison,
WI.
Cutler, S. J., and Ederer, F. (1958 ). Maximum Utilization of the Life Table Method in
Analyzing Survival. Journal of Chronic Diseases ,8,6 99—712.
Cutler, S. J., Griswold, M. H., and Eisenberg, H. (1957 ). An Interpretation of Survival
Rates: Cancer of the Breast. Journal of the National Cancer Institute ,19, 1107—
1117.
Cutler, S. J., Ederer, F., Griswold, M. H., and Greenberg, R. A. (1959 ). Survival of
Breast-Cancer Patients in Connecticut, 1935 —54.Journal of the National Cancer
Institute,23, 1137—1156.
Cutler, S. J., Ederer, F., Griswold, M. H., and Greenberg, R. A. (1960a ). Survival of
Patients with Uterine Cancer, Connecticut, 1935 —54.Journal of the National
Cancer Institute ,24, 519—539.
Cutler, S. J., Ederer, F., Griswold, M. H., and Greenberg, R. A. (1960b ). Survival of
Patients with Ovarian Cancer, Connecticut, 1935 —54.Journal of the National
Cancer Institute ,24, 541—549. 493
Cutler, S. J., Axtell, L., and Heise, H. (1967 ). Ten Thousand Cases of Leukemia:
1940—62.Journal of the National Cancer Institute ,39,993—1026.
Daniel, C. (1959 ). Use of Half-Normal Plots in Interpreting Factorial Two-Level
Experiments. Technometrics ,1, 311—341.
Daniel, W. W. (1987 ).Biostatistics: A Foundation for Analysis in the Health Sciences.
Wiley, New York.
Davis, D. J. (1952 ). An Analysis of Some Failure Data. Journal of the American
Statistical Association ,47, 113—150.
Davis, H. T., and Feldstein, M. L. (1979 ). The Generalized Pareto Law as a Model for
Progressively Censored Survival Data. Biometrika ,66,2 99—306.
Dawber, T. R. (1980 ).The Framingham Study. Harvard University Press, Cambridge,
MA.
Dawber, T. R., Meadors, G. F., and Moore, F. E. Jr. (1951 ). Epidemiological Ap-
proaches to Heart Disease: The Framingham Study. American Journal of Public
Health,41, 279—286.
Delong, D. M., Guirguis, G. H., and So, Y. C. (1994 ). Efficient Computation of Subset
Selection Probablilities with Application to Cox Regression. Biometrica. 81
607—611.
Dharmalingam, A., Pool, I., and Dickson, J. (2000 ). Biosocial Determinants of Hyster-
ectomy in New Zealand. American Journal of Public Health ,90(9), 1455—1458.
Dixon,W. J.,Brown, M.B., Engelman,L.,Hill,M. A., and Jennrich,R.I. (1990 ).BMDP
Statistical Software Manual . University of California Press, Berkeley, CA.
Draper, N. R., and Smith, H. (1966 ).Applied Regression Analysis . Wiley, New York.
Drenick, R. F. (1960 ). The Failure Law of Complex Equipment. Journal of Social and
Industrial Applied Mathematics ,8, 680.
Dunn, O. J. (1964 ). New Table for Multiple Comparisons with a Control. Biometrics ,
20, 482—491.
Ederer, F., Axtell, L. M., and Cutler, S. J. (1961 ). The Relative Survival Rate: A
Statistical Methodology. National Cancer Institute Monographs ,6, 101—121.
Efron, B. (1975 ). The Efficiency of Logistic Regression Compared to Normal Dis-
criminant Analysis. Journal of the American Statistical Association ,70,8 92—898.
Efron, B. (1977 ). The Efficiency of Cox’s Likelihood Function for Censored Data.
Journal of the American Statistical Association ,72, 557—565.
Efron, B. (1994 ). Missing Data, Imputation, and the Bootstrap. JournaloftheAmerican
StatisticalAssociation ,89, 463—475.
Eisenberger, M., Krasnow, S., Ellenberg, S., et al. (1989 ). A Comparison of Carboplatin
Plus Methotrexate versus Methotrexate Alone in Patients with Recurrent and
Metastatic Head and Neck Cancer. Journal of Clinical Oncology ,7, 1341—1345.
Elaad,E., andBen-Shakhar,G. (1989 ).Effects of MotivationandVerbalResponseType
on Psychophysiological Detection of Information. Psychophysiology ,26, 442—451.
Elandt-Johnson, R. C., and Johnson, N. L. (1980 ).Survival Models and Data Analysis .
Wiley, New York.
Enas, G. G., Dornseit, B. E., Sampson, C. B., Rockhold, F. W., and Wuu, J. (1989 ).
Monitoring versus Interim Analysis of Clinical Trials: A Perspective from thePharmaceutical Industry. Controlled Clinical Trials ,10,5 7—70.494
Epstein, B. (1958 ). The Exponential Distribution and Its Role in Life Testing. Industrial
Quality Control ,15,2—7.
Epstein, B. (1960a ). Estimation of the Parameters of Two Parameter Exponential
Distribution from Censored Samples. Technometrics ,2, 403—406.
Epstein, B. (1960b ). Estimation from Life Test Data. Technometrics ,2, 447—454.
Epstein, B., and Sobel, M. (1953 ). Life Testing. Journal of the American Statistical
Association ,48, 486—502.
Farewell, V. T. (1979 ). Some Results on the Estimation of Logistic Models Based on
Retrospective Data. Biometrika ,66,2 7—32.
Farrington, C. P. (2000 ). Residuals for Proportional Hazards Models with Interval-
Censored Survival Data. Biometrics ,56(2), 473—482.
Feigl, P., and Zelen, M. (1965 ). Estimation of Exponential Survival Probabilities with
Concomitant Information. Biometrics ,21, 826—838.
Feinleib, M. (1960 ). A Method of Analyzing Log-Normally Distributed Survival Data
with Incomplete Follow-up. Journal of the American Statistical Association ,55,
534—545.
Feinleib, M., and MacMahon, B. (1960 ). Variation in the Duration of Survival of
Patients with Chronic Leukemias. Blood,17, 332—349.
Feskanich, D., Singh, V., Willett, W. C., and Colditz, G. A. (2002 ). Vitamin A Intake
and Hip Fractures among Postmenopausal Women. Journal of the American
Medical Association ,287(1)47—54.
Fish, E. B., Chapman, J. A. and Link, M. A. (1998 ). Competing Causes of Death for
Primary Breast Cancer. Annals of Surgical Oncology ,5(4), 368—375.
Fisher, R. A. (1922 ). On the Mathematical Foundation of Theoretical Statistics.
Philosophical Transactions of the Royal Society of London, Series A ,222.
Fisher, R. A. (1936 ). The Use of Multiple Measurements in Toxonomic Problems.
Annals of Eugenics ,7, 312—330.
Fleiss, J. L. (1979 ). Confidence Intervals for the Odds Ratio in Case-Control Studies:
The State of the Art. Journal of Chronic Diseases ,32,6 9—82.
Fleiss, J. L. (1981 ).Statistical Methods for Rates and Proportions. Wiley, New York.
Fleming, T. R., and Harrington, D. P. (1979 ). Non-parametric Estimation of the
Survival Distribution in Censored Data. Unpublished manuscript.
Fleming, T. R., and Harrington, D. P. (1991 ).Counting Processes and Survival Analysis .
Wiley, New York.
Fleming, T. R., O’Fallon, J. R., O’Brian, P. C., and Harrington, D. P. (1980 ). Modified
Kolmogorov—Smirnov Test Procedures with Application to Arbitrarily Right
Censored Data. Biometrics ,36, 607—626.
Fleming, T. R., Harrington, D. P., and O’Brien, P. C. (1984 ). Designs for Group
Sequential Tests. Controlled Clinical Trials ,5, 348—361.
Florin, V., and Ronghui, X. (2000 ). Proportional Hazards Model with Random Effects.
Statistics in Medicine ,19(24), 3309—3324.
Fraser, D. A. S. (1968 ).The Structure of Inference. Wiley, New York.
Freedman, L. S. (1982 ). Tables of the Number of Patients Required in Clinical Trials
Using the Log Rank Test. Statistics in Medicine ,1, 121—129. 495
Frei, E., et al. (1961 ). Studies of Sequential and Combination Antimetabolite Therapy
in Acute Leukemia: 6 —Mercaptopurine and Methotrexate. Blood,18, 431—454.
Freireich, E. J., Gehan, E. A., Frei, E., et al. (1963 ). The Effect of 6-Mercaptopurine on
the Duration of Steroid-Induced Remissions in Acute Leukemia: A Model for
Evaluation of Other Potential Useful Therapy. Blood,21(6),6 99—716.
Freireich, E. J., Gehan, E. A., Rall, D. P., Schmidt, L. H., and Skipper, H. E. (1966 ).
Quantitative Comparison of Toxicity of Anticancer Agents in Mouse, Rat,Hamster, Dog, Monkey, and Man. Cancer Chemotherapy Report ,50,4 .
Freireich, E. J., Gehan, E. A., Bodey, G. P., Hersh, E. M., Hart, J. S., Gutterman, J. U.,
and McCredie, K. B. (1974 ). New Prognostic Factors Affecting Response and
Survival in Adult Leukemia. Transactions of the Association of American Phys-
icians,87,2 98—305.
Friedman, L. M., Furberg, C. D., and DeMets, D. L. (1985 ).Fundamentals of Clinical
Trials, 2nd ed. PSG Publishing, Littleton, MA.
Gaddum, J. H. (1945a ). Log Normal Distributions. Nature, London ,156, 463.
Gaddum, J. H. (1945b ). Log Normal Distributions. Nature, London ,156, 747.
Gail, M., and Gart, J. J. (1973 ). The Determination of Sample Sizes for Use with the
Exact Conditional Test in 2 /p592 Comparative Trials. Biometrics ,29, 441—448.
Gail, M. H., Lubin, J. H., and Rubinstein, L. V. (1981 ). Likelihood Calculations for
Matched Case-Control Studies and Survival Studies with Tied Death Times.
Biometrika ,68, 703—707.
Gajjar, A. V., and Khatri, C. G. (1969 ). Progressively Censored Samples from Log-
Normal and Logistic Distributions. Technometrics ,11,7 93—803.
Garside, M. J. (1965 ). The Best Sub-set in Multiple Regression Analysis. Applied
Statistics ,14,1 96—200.
Gehan, E. A. (1965a ). A Generalized Wilcoxon Test for Comparing Arbitrarily
Singly-Censored Samples. Biometrika ,52, 203—223.
Gehan, E. A. (1965b ). A Generalized Two-Sample Wilcoxon Test for Doubly-Censored
Data. Biometrika ,52, 650—653.
Gehan, E. A. (1970 ). Unpublished notes on survival time studies. The University of
Texas M. D. Anderson Cancer Center, Houston, Texas.
Gehan, E. A. (1969 ). Estimating Survival Function from the Life Table. Journal of
Chronic Diseases ,21, 629—644.
Gehan, E. A., and Thomas, D. G. (1969 ). The Performance of Some Two-Sample Tests
in Small Samples with and without Censoring. Biometrika ,56, 127—132.
Gelenberg, A. J., Kane, J. M., Keller, M. B., et al. (1989 ). Comparison of Standard and
Low Serum Levels of Lithium for Maintenance Treatment of Bipolar Disorder.
New England Journal of Medicine ,321, 1489—1493.
George, S. L., Fernback, D. J., et al. (1973 ). Factors Influencing Survival in Pediatric
Acute Leukemia: The SWCCSG Experience, 1959 —1970. Cancer,32, 1542—1553.
Gertsbakh, I. B. (1989 ).Statistical Reliability Theory. Marcel Dekker, New York.
Gill, R., and Schumacher, M. (1987 ). A Simple Test of the Proportional Hazards
Assumption. Biometrika ,74, 289—300.496
Gillum, R. F., Fortmann, S. P., Prineas, R. J., and Kottke, T. E. (1984 ). International
Diagnostic Criteria for Acute Myocardial Infarction and Acute Stroke. American
Heart Journal ,108, 150—158.
Glasser, M. (1967 ). Exponential Survival with Covariance. Journal of the American
Statistical Association ,62, 561—568.
Gompertz, B. (1825 ). On the Nature of the Function Expressive of the Law of Human
Mortality and on the New Mode of Determining the Value of Life Contingencies.Philosophical Transactions ,513.
Gore, S. M. (1983 ). Graft Survival after Renal Transplantation: Agenda for Analysis.
KidneyInt. ,24, 516—525.
Grambsch, P M., Therneau, T. M. (1994 ). Proportional Hazards Tests in Diagnostics
Based on Weighted Residuals. Biometrika ,81, 515—526.
Gray, R. J. (1990 ). Some Diagnostic Methods for Cox Regression Models through
Hazard Smoothing. Biometrics ,46,93—102.
Green, P. J. (1984 ). Iteratively Reweighted Least Squares for Maximum Likelihood
Estimation, and Some Robust and Resistant Alternatives (with discussion ).Jour-
nal of the Royal Statistical Society ,46(2), 149—192.
Greenwood, J. A., and Durand, D. (1960 ). Aids for Fitting the Gamma Distribution by
Maximum Likelihood. Technometrics ,2,5 5—65.
Greenwood, M. (1926 ). The Natural Duration of Cancer. Reports on Public Health and
Medical Subjects , Her Majesty’s Stationary Office, London, 33,1—26.
Griswold, M. H., and Cutler, S. J. (1956 ). The Connecticut Cancer Register: Seventeen
Years of Experience. Connecticut Medical Journal ,20, 366—372.
Griswold, M. H., Wilder, C. S., Cutler, S. J., and Pollack, E. S. (1955 ).Cancer in
Connecticut, 1935 —1951. Monograph. Connecticut State Department of Health,
Hartford, CT.
Grizzle, J. E. (1967 ). Continuity Correction in the /afii9851/p17-Test for 2 /p592 Tables. American
Statistician ,21,2 8—32.
Gross, A. J., and Clark, V. A. (1975 ).Survival Distributions: Reliability Applications in
the Biomedical Sciences . Wiley, New York.
Grove, R. D., and Hetzel, A. M. (1963 ).Vital Statistics Rates in the United States,
1940—1960.National Center for Health Statistics, Washington, DC.
Gupta, A. K. (1952 ). Estimation of the Mean and Standard Deviation of a Normal
Population from a Censored Sample. Biometrika ,39, 260—273.
Gupta, S. S. (1960 ). Order Statistics from the Gamma Distribution. Technometrics ,2,
243—262.
Hagar,H. W.,and Bain,L. J. (1970 ). InferentialProceduresforthe GeneralizedGamma
Distribution. Journal of the American Statistical Association ,65, 1601—1609.
Hahn, G. J., and Shapiro, S. S. (1967 ).StatisticalModelsinEngineering . Wiley, New
York.
Haldane, J. B. S. (1956 ). The Estimation and Significance of the Logarithm of a Ratio
of Frequencies. Annals of Human Genetics ,20, 309—311.
Halperin, M. (1952 ). Maximum Likelihood Estimation in Truncated Samples. Annals
of Mathematical Statistics ,23, 226—238. 497
Halperin, M., Blackwelder, W. C., and Verter, J. I. (1971 ). Estimation of the Multivari-
ate Logistic Risk Function: A Comparison of the Discriminant Function andMaximum Likelihood Approaches. Journal of Chronic Diseases ,24, 125—158.
Hammond, I. W., Lee, E. T., Davis, A. W., and Booze, C. F. (1984 ). Prognostic Factors
Related to Survival and Complication-Free Times in Airmen Medically Certified
after Coronary Bypass Surgery. Aviation, Space, and Environmental Medicine ,
April, pp. 321—331.
Hannan, E. J. (1979 ). The Determination of the Order of an Autoregression. Journal of
the Royal Statistical Society, Series B ,41,1 90—195.
Hanson, B. S., Isacsson, S-O., Janzon, L., and Lindell, S. E. (1989 ). Social Network and
Social Support Influence Mortality in Elderly Men. American Journal of Epi-
demiology ,130, 100—111.
Harrison,J.D., Jones,J.A.,and Morris,D.L. (1990 ).TheEffectofthe GastrinReceptor
Antagonist Proglumide on Survival in Gastric Carcinoma. Cancer,66, 1449—1452.
Hart, J. S., George, S. L., Frei, E., Bodey, G. P., Nickerson, R. C., and Freireich, E. J.
(1977 ). Prognostic Significance of Pretreatment Proliferative Activity in Adult
Acute Leukemia. Cancer,39, 1603—1617.
Harter, H. L., and Moore, A. H. (1965 ). Maximum Likelihood Estimation of the
Parameters of Gamma and Weibull Populations from Complete and from Cen-sored Samples. Technometrics ,7, 639—643.
Harter, H. L., and Moore, A. H. (1966 ). Local Maximum Likelihood Estimation of the
Parameters of Three-Parameter Log-Normal Population from Complete and
Censored Sample. Journal of the American Statistical Association ,61, 842—851.
Harter, H. L., and Moore, A. H. (1967 ). Asymptotic Variance and Covariances of
Maximum Likelihood Estimators, from Censored Samples, of the Parameters ofWeibull and Gamma Parameters. Annals of Mathematical Statistic s,38, 557—570.
Hastings, N. A. J., and Peacock, J. B. (1974 ).Statistical Distributions . Butterworth,
London.
Hauck, W. W., Jr., and Donner, A. (1977 ). Wald’s Test as Applied to Hypotheses in
Logit Analysis. Journal of the American Statistical Association ,72, 851—853.
Haughton,D. M. A. (1988 ). On the Choice of a Model to Fit Data froman Exponential
Family. Annals of Statistics ,16, 342—355.
Heitjan, D. F. (1997 ). Annotation: What Can be Done About Missing Data? Ap-
proaches to imputation, Am.J.PublicHealth ,87, 548—550.
Hemstreet, G. P., Yin, S., Ma, Z., Bonner, R. B., Bi, W., Rao, J. Y., Zang, M., Zheng,
Q., Bane, B., Asal, N., Li, G., Feng, P., Hurst, R. E., and Wang, W. (2001 ).
Biomarker Risk Assessment and Bladder Cancer Detection in a Cohort Exposedto Benzidine,JournaloftheNationalCancerInstitute ,93, 427—436.
Hill, A. B. (1960a ).Controlled Clinical Trials. Blackwell Scientific, Oxford
Hill, A. B. (1960b ).Statistical Methods in Clinical and Preventive Medicine. Oxford
University Press, Oxford.
Hill, A. B. (1971 ).Principles of Medical Statistics. Oxford University Press, New York.
Hirayama, T. (1981 ). Non-smoking Wives of Heavy Smokers Have a Higher Risk of
Lung Cancer: A Study from Japan. British Medical Journal ,282, 183—185.498
Hoel, D. G., Sobel, M., and Weiss, G. H. (1975 ). A Survey of Adaptive Sampling for
Clinical Trials. Perspectives in Biometrics ,1,2 9—61.
Holford, T. R., White, C., and Kelsey, J. L. (1978 ). Multivariate Analysis for Matched
Case-Control Studies. American Journal of Epidemiology ,107, 245—256.
Hollander, M., and Proschan, F. (1979 ). Testing to Determine the Underlying Distribu-
tion Using Randomly Censored Data. Biometrics ,35,3 93—401.
Hollander, M., and Wolfe,D. A. (1973 ).Nonparametric Statistical Methods. Wiley, New
York.
Horner, R. D. (1987 ). Age at Onset of Alzheimer’s Disease: Clue to the Relative
ImportanceofEtiologicFactors? American Journal of Epidemiology ,126,409—414.
Hosmer, D. W., and Lemeshow, S. (1980 ). A Goodness-of-Fit Test for the Multiple
Logistic Regression Model. CommunicationsinStatistics ,A10, 1043—1069.
Hosmer,D.W., and Lemeshow,S. (1999 ).Applied Survival Analysis , 2nd ed.Wiley, New
York.
Hosmer, D. W., and Lemeshow, S. (2000 ).Applied Logistic Regression. Wiley, New
York.
Howell, D. W. (1987 ).Statistical Methods for Psychology. Duxbury Press, Boston.
Hung, C. T., Lim, J. K. C., and Zoest, A. R. (1988 ). Optimization of High-Performance
Liquid Chromatographic Analysis for Isoxazolye Penicillins Using FactorialDesign. Journal of Chromatography ,425, 331—341.
Ibrahim, J. G., Chen, M. H., and Sinha, D. (2001 ).Bayesian Survival Analysis.
Springer-Verlag, New York
Ingram, D. D., and Kleinman, J. C. (1989 ). Empirical Comparisons of Proportional
Hazards and Logistic Regression Models. Statistics in Medicine ,8, 525—538.
Irwin, J. O. (1949 ). The Standard Error of an Estimate of Expectational Life. Journal
of Hygiene ,47, 188—189.
Jenkins, S. P. (1997 ). Discrete Time Proportional Hazards Regression. Stata Technical
Bulletin,39,1 7—32.
Jennings, D. E. (1986 ). Judging Inference Adequacy in Logistic Regression. Journal of
the American Statistical Association ,81, 471—476.
Johnson, N. L., and Kotz, S. (1970a ).Distributions in Statistics: Continuous Univariate
Distributions (Vol. 1 )Houghton Mifflin, Boston.
Johnson, N. L., and Kotz, S. (1970b ).Distributions in Statistics: Continuous Univariate
Distributions. (Vol. 2 )Houghton Mifflin, Boston
Johnson, P., and Pearce, J. M. (1990 ). Recurrent Spontaneous Abortion and Polycystic
Ovarian Disease: Comparison of Two Regimens to Induce Ovulation. British
Medical Journal ,300, 154—156.
Kahn,H. A. (1983 ).An Introduction to Epidemiologic Methods. OxfordUniversityPress,
New York.
Kalbfleisch, J. D. (1974 ). Some Extensions and Applications of Cox’s Regression and
Life Model. Presented at the joint meeting of the Biometric Society and theAmerican Statistical Association, Tallahassee, FL, March 20 —22.
Kalbfleisch, J. D., and Prentice, R. L. (1973 ). Marginal Likelihoods Based on Cox’s
Regression and Life Table Model. Biometrika ,60, 267—278. 499
Kalbfleisch, J. D., and Prentice, R. L. (1980 ).The Statistical Analysis of Failure Time
Data. Wiley, New York.
Kao, J. H. K. (1958 ). Computer Methods for Estimating Weibull Parameters in
ReliabilityStudies. I.R.E. Transactions on Reliability and Quality Control ,PGRQC
13,1 5—22.
Kaplan, E. L., and Meier, P. (1958 ). Nonparametric Estimation from Incomplete
Observations.Journalof theAmericanStatisticalAssociation ,53, 457—481.
Kay, R. (1979 ). Proportional Hazard Regression Models and the Analysis of Censored
Survival Data. Applied Statistics ,26, 227—237.
Kay, R. (1984 ). Goodness of Fit Methods for the Proportional Hazards Model: A
Review. Revue Epidemiologie et de Santé Publique ,32, 185—198.
Kelsey, J. L., Thompson, W. D., and Evans, A. S. (1986 ).Methods in Observational
Epidemiology. Oxford University Press, New York.
Kessing, L.V., Olsen, E. W., and Andersen, P. K. (1999 ). Recurrence in Affective
Disorder:AnalyseswithFrailty Models. American Journal of Epidemiology ,149(5),
404—411.
King, J. R. (1971 ).Probability Charts for Decision Making. Industrial Press, New York.
King, M., Bailey, D. M., Gibson, D. G., Pitha, J. V., and McCay, P. B. (1979 ). Incidence
andGrowth ofMammaryTumors Induced by7,12-Dimethylbenz (/afii9825)antheaceneas
Related to the Dietary Content of Fat and Antioxidant. Journal of the National
Cancer Institute ,63, 656—664.
Kitagawa, E. M. (1964 ). Standardized Comparisons in Population Research. Demogra-
phy,1,2 96—315.
Klein, J. P., and Moeschberger, M. L. (1997 )Survival Analysis. Springer-Verlag, New
York.
Kleinbaum, D. G. (1994 ).LogisticRegression:ASelf-LearningText. Springer-Verlag,
New York.
Kleinbaum, D. G., Kupper, L. L., and Muller, K. E. (1988 ).Applied Regression Analysis
and Other Multivariate Methods, 2nd ed. PWS-Kent, Boston.
Kodlin, D. (1967 ). A New Response Time Distribution. Biometrics ,23, 227—239.
Krishna, I. P. V. (1951 ). A Non-parametric Method of Testing kSamples. Nature,167,
33.
Kruskal, W. H., and Wallis, W. A. (1952 ). Use of Ranks in One-Criterion Variance
Analysis. Journal of the American Statistical Association ,47, 583—621.
Kuzma, J. W. (1967 ). A Comparison of Two Life Table Methods. Biometrics ,23,5 1—64.
Lagakos, S. W. (1980 ). The Graphical Evaluation of Explanatory Variables in Propor-
tional Hazard Regression Models. Biometrika ,68,93—98.
Lan, K. K. G., and DeMets, D. L. (1983 ). Discrete Sequential Boundaries for Clinical
Trials. Biometrika ,70, 659—663.
Lan, K. K. G., and DeMets, D. L. (1989 ). Changing Frequency of Interim Analysis in
Sequential Monitoring. Biometrics ,45, 1017—1020.
Lassare, S. (2001 ). Analysis of Progress in Road Safety in Ten European Countries.
Accident Analysis and Prevention ,33(6), 743—751.
Lawless, J. F. (1982 ).Statistical Methods and Model for Lifetime Data . Wiley, New
York.500
Lawless, J. F. (1983 ). Statistical Methods in Reliability. Technometrics ,25, 305—316.
Lee, A. H., and Yau, K. K. (2001 )Determining the Effects of Patient Case Mix on
Length of Hospital Stay: A Proportional Hazards Frailty Model Approach.Methods of Information in Medicine ,40(4), 288—292.
Lee, E. T. (1980 ).Statistical Methods for Survival Data Analysis. Lifetime Learning
Publications, Belmont, CA.
Lee, E. T. (1992 ).StatisticalMethodsforSurvivalDataAnalysis , second edition, Wiley,
New York.
Lee, E. T., and Thomas, D. R. (1980 ). Confidence Interval for Comparing Two Life
Distributions. IEEE Transactions on Reliability ,R-29,5 1—56.
Lee, E. T., Desu, M. M., and Gehan, E. A. (1975 ). A Monte-Carlo Study of the Power
of Some Two-Sample Tests. Biometrika ,62, 425—432.
Lee, E. T., Ishmael, D. R., Bottomley, R. H., and Murray, J. L. (1982 ). An Analysis of
Skin Tests and Their Relationship to Recurrence and Survival in Stage III andStage IV Melanoma Patients. Cancer,49, 2336—2341.
Lee, E. T., Yeh, J. L., Cleves, M. A., and Shafer, D. (1988 ). Vascular Complications in
Noninsulin Dependent Diabetic Oklahoma Indians. Diabetes,37(Suppl. 1 ).
Lee, E. T., Lee, V. S., Lu, M., et al. (1992 ). Incidence and Risk Factors of Diabetic
Retinopathy in Oklahoma Indians with NIDDM. Diabetes Care ,15, 1620—1627.
Lee, E. T., Russell, D., Jorge, N., Kenny, S., and Yu, M. (1993 ). A Follow-up Study of
Diabetic Oklahoma Indians: Mortality and Causes of Death. Diabetes Care ,16,
300—305.
Leenen, F. H. H., Balfe, J. A., Pelech, A. N., et al. (1987 ). Postoperative Hypertension
after Repair of Coarctation of Aorta in Children: Protective Effect of Propranolol.
American Heart Journal ,113, 1164—1173.
Lehmann,E. L. (1953 ). The Power of Rank Tests. Annals of Mathematical Statistics ,24,
23—43.
Leitner, L. M., Roumy, M. Ruckebusch, M., Sutra, J. F. (1986 ). Monoamines and Their
Catabolites in the Rabbit Carotid Body. Effets of reserpine, sympathectomy andcarotid sinus nerve section, EuropeanJournalofPhysiology ,406, 552—556.
Lemeshow, S., and Hosmer, D. W. (1982 ). A Review of Goodness-of-Fit Statistics for
Use in the Development of Logistic Regression Models. American Journal of
Epidemiology ,115,92—106.
Leyland-Jones, B., Donnelly, H., Groshen, S., Myskowski, P., Donner, A. L., Fanucchi,
M., Fox, J., and the Memorial Sloan-Kettering Antiviral Working Group (1986 ).
2/p30-Fluror-5-Iodoarabinosylcytosine, A New Potent Antiviral Agent: Efficacy in
ImmunosuppressedIndividuals with Herpes Zoster. Journal of Infectious Diseases ,
154, 430—436.
Liang, K. Y., Self, S. G., and Liu, X. (1990 ). The Cox Proportional Hazards Model with
Change Point: An Epidemiologic Application. Biometrics ,46, 783—793.
Liang, K. Y., Self, S. G., Bandeen-Roche, K. J., and Zeger, S. L. (1995 ). Some Recent
Developments for Regression Analysis of Multivariate Failure Time Data. Life-
time Data Analysis, 1, 403—415.
Lieblein, J., and Zelen, M. (1956 ). Statistical Investigation of the Fatigue Life of
Deep-Grove Ball Bearings. Journal of Research of the National Bureau of Stan-
dards,57, 273—316. 501
Lilliefors, H. W. (1971 ). Reducing the Bias of Estimators of Parameters for the Erlang
and Gamma Distribution. Unpublished manuscript.
Lindley, D. V. (1968 ). The Choice of Variables in Multiple Regression. Journal of the
Royal Statistical Society, Series B ,30,3 1—53.
Linka, A. Z., Sklenar, J., Wei, K. I., Jayaweera,A. R., Skyba, D. M., and Kaul, S. (1998 ).
Assessment of Transmural Distribution of Myocardial Perfusion with ContrastEchocardiography. Circulation 3 ;98(18); 1912—1920.
Little, R. J., and Rubin, D. B. (1987 ).StatisticalAnalysiswithMissingData , John Wiley
& Sons, New York.
Liu, P. Y., and Crowley, J. (1978 ). Large Sample Theory of the MLE Based on Cox’s
Regression Model for Survival Data. Technical Report 1 . Wisconsin Clinical
Cancer Center (Biostatistics ), University of Wisconsin, Madison, WI.
Lubin, J. H. (1981 ). A Computer Program for the Analysis of Matched Case-Control
Studies. Computers and Biomedical Research ,14, 138—143.
McAlister, D. (1879 ). The Law of the Geometric Mean. Proceedings of the Royal
Society,29, 367.
McCracken, D. D., and Dorn, W. S. (1964 ).Numerical Methods and Fortran Pro-
gramming. Wiley, New York.
McCullagh, P. (1980 ). Regression Model for Ordinal Data. Journal of the Royal
Statistical Society ,42(2), 109—142.
McCullagh, P., and Nelder, J. A. (1989 ).Generalized Linear Models . Chapman & Hall,
London.
McFadden, D. (1976 ). A Comment on Discriminant Analysis ‘‘versus’’ Logit Analysis.
Annals of Economic and Social Measurement ,5, 511—523.
Mackenbach, J. P., Kunst, A. E., Lautenbach, H., Bijlsma, F., and Oei, Y.B. (1995 ).
Competing Causes of Death: An Analysis Using Multiple-Cause-of-Death Data
from The Netherlands. 141(5), 466—475.
Mafart, P., Couvert, O., Gaillard, S., and Leguerinel, I. (2002 ). On Calculating Sterility
inThermalPreservationMethods:Applicationofthe WeibullFrequencyDistribu-tion Model. International Journal of Food Microbiology ,72(12); 107—113.
Mann, H. B., and Whitney, D. R. (1947 ). On a Test of Whether One of Two Random
Variables Is Stochastically Larger Than the Other. Annals of Mathematical
Statistics ,18,5 0—60.
Mann, N. R. (1970 ). Estimators and Exact Confidence Bounds for Weibull Parameters
Based on a Few Ordered Observations. Technometrics ,12, 345—361.
Mann, N. R., Schafer, R. E., and Singpurwalla, N. D. (1974 ).Methods for Statistical
Analysis of Reliability and Life Data. Wiley, New York.
Manninen,O. (1988 ). Changes in Hearing, CardiovascularFunctions, Haemodynamics,
Upright Body Sway, Urinary Catecholamines and Their Correlates after Pro-longed Successive Exposure to Complex Environmental Conditions. International
Archives of Occupational and Environmental Health ,60, 249—272.
Mantel, N. (1966 ). Evaluation of Survival Data and Two New Rank Order Statistics
Arising in Its Consideration. Cancer Chemotherapy Reports ,50, 163—170.
Mantel,N. (1967 ). RankingProceduresfor Arbitrarily RestrictedObservations. Biomet-
rics,23,6 5—78.502
Mantel, N. (1970 ). Why Stepdown Procedures in Variable Selection. Technometrics ,12,
621—625.
Mantel, N. (1973 ). Synthetic Retrospective Studies and Related Topics. Biometrics ,29,
479—486.
Mantel, N. (1977 ). Test and Limits for the Common Odds Ratio of Several 2 /p592
Contingency Tables: Methods in Analogy with the Mantel —Haenszel Procedure.
Journal of Statistical Planning Information ,1, 179—189.
Mantel, N., and Haenszel, W. (1959 ). Statistical Aspects of the Analysis of Data from
Retrospective Studies of Disease. Journal of the National Cancer Institute ,22,
719—748.
Mantel,N., and Hankey,B. F. (1978 ). A Logistic RegressionAnalysisof Response-Time
Data Where the Hazard Function Is Time Dependent. Communications in Statis-
tics A: Theory and Methods ,7, 333—347.
Mantel, N., and Myers, M. (1971 ). Problems of Convergence of Maximum Likelihood
Iterative Procedures in Multiparameter Situation. Journal of the American Statis-
tical Association ,66, 484—491.
Mantel, N., and Stark, C. R. (1968 ). Computation of Indirect Adjusted Rates in the
Presence of Confounding. Biometrics ,24,997—1005.
Marascuilo, L. A., and McSweeney, M. (1977 ).Nonparametric and Distribution-Free
Methods for the Social Sciences. Brooks/Cole, Monterey, CA.
Marubini, E., and Valsecchi, M. G. (1995 ).Analyzing Survival Data from Clinical Trials
and Observational Studies. Wiley, New York.
Matthews, D. E., and Farewell, V. (1985 ).Using and Understanding Medical Statistics.
S. Karger, New York.
Mausner, J. S., and Kramer, S. (1985 ).Epidemiology: An Introductory Text. W.B.
Saunders, Philadelphia.
Meier, P. (1975a ). Statistics and Medical Experimentation. Biometrics ,31, 511—529.
Meier, P. (1975b ). Estimation of a Distribution Function from Incomplete Observa-
tions, in Perspectives in Probability and Statistics , edited by J. Gaui. Applied
Probability Trust, Sheffield, England.
Meisinger, C., Thorand, B., Schneider, A., Stieber, J., Doring, A., and Lowel, H. (2002 )
Sex Differences in Risk Factors for Incident Type 2 Diabetes Mellitus: theMONICA Augsburg Cohort Study. ArchInternMed ,162,8 2—89.
Miettinen, O. S. (1979 ). Comments on ‘‘Confidence Intervals for the Odds Ratio in
Case-Control Studies: The State of the Art,’’ by J. L. Fleiss. Journal of Chronic
Diseases,32,8 0—82.
Miller, R. G., Jr. (1966 ).Simultaneous Statistical Inference. McGraw-Hill, New York.
Miller, R. G. (1981 ).Survival Analysis. Wiley, New York.
Minow, R. A., Benjamin, R. S., Lee, E. T., and Gottlieb, J. A. (1977 ). Adriamycin
Cardiomyopathy: Risk Factors. Cancer,39, 1397—1402.
Molloy, D. W., Guyatt, G. H., Wilson, D. B., et al. (1991 ). Effect of Tetrahydroaminoac-
ridine on Cognition, Function and Behaviour in Alzheimer’s Disease. Canadian
Medical Association Journal ,144,2 9—34.
Montaner, J. S. G., Lawson, L. M., Levitt, N., et al. (1990 ). Costicorsteroids Prevent
Early Deterioration in Patients with Moderately Severe Pneumocystis Carinii 503
Pneumonia and the Acquired Immunodeficiency Syndrome (AIDS ).Annals of
Internal Medicine ,113,1 4—20.
Moolgavkar, S., Lustbader, E., and Venzon, D. J. (1985 ). Assessing the Adequacy of the
Logistic Regression Model for Matched Case-Control Studies. Statistics in Medi-
cine,4, 425—435.
Moreau, T., O’Quigley, J., and Mesbah, M. (1985 ). A Global Goodness-of-Fit Statistic
for the Proportional Hazards Model. Applied Statistics ,34, 212—218.
Morrison, D. F. (1967 ).Multivariate Statistical Methods. McGraw-Hill, New York.
Morrison,R. S., and Siu, A.L. (2000 ).Survivalin End-StageDementiaFollowingAcute
Illness. Journal of the American Medical Association ,284(1),4 7—52.
Myers, M., Hankey, B. F., and Mantel, N. (1973 ). A Logistic-Exponential Model for
Use with the Response-Time Data Involving Regressor Variables. Biometrics ,29,
257—269.
Myers, M. H. (1969 ). A Computing Procedure for a Significance Test of the Difference
between Two Survival Curves. Methodological Note 18 in Methodological Notes.
End Results Sections, National Cancer Institute, National Institutes of Health,Bethesda, MD.
Nadas, A. (1970 ). On Proportional Hazard Functions. Technometrics ,12, 413—416.
National Cancer Institute (1970 ).Proceedings of the Symposium on Statistical Aspects
of Protocol Design , San Juan, Puerto Rico, December 9 —10.
Natrella, M. G. (1963 ).Experimental Statistics . National Bureau of Standards Hand-
book 91. U.S. Government Printing Office, Washington, DC, Tables A-25, A-26.
Nelson, W. (1972 ). Theory and Applications of Hazard Plotting for Censored Failure
Data. Technometrics ,14,94 5—966.
Nelson, W. (1982 ).Applied Life Data Analysis. Wiley, New York.
Nemenyi, P. (1963 ). Distribution-Free Multiple Comparisons. Ph.D. dissertation, Prin-
ceton University.
Neter, J., and Wasserman, W. (1974 ).Applied Linear Statistical Models. Richard D.
Irwin, Homewood, IL.
Nie, N. H., Hull, C. H., Jenkins, J. G., Steinbrenner, K., and Bent, D. H. (1975 ).SPSS:
Statistical Package for the Social Sciences. McGraw-Hill, New York.
O’Brien, P. C., and Fleming, T. R. (1979 ). A Multiple Testing Procedure for Clinical
Trials. Biometrics ,35, 549—556.
Osgood, E. W. (1958 ). Methods for Analyzing Survival Data, Illustrated by Hodgkin’s
Disease. American Journal of Medicine ,24,4 0—47.
Parker, R. L., Dry, T. J., Willius, F. A., and Gage, R. P. (1946 ). Life Expectancy in
Angina Pectoris. Journal of the American Medical Association ,131,95—100.
Parzan, E. (1974 ). Some Recent Advances in Time Series Modeling. IEEE Transactions
on Automatic Control. AC-19, 723—730.
Pearson, E. S., and Hartely, N. O. (1958 ).Biometrika Tables for Statisticians , Vol. 1.
Cambridge University Press, Cambridge.
Pearson, K. (1922, 1957 ).Tables of the Incomplete /afii9772-Function. Cambridge University
Press, Cambridge.
Pershagen,G. (1986 ). Reviewof Epidemiologyin Relationto PassiveSmoking. Archives
of Toxicology ,9(Suppl. ),6 3—73.504
Pershagen, G., Hrubec, Z., and Svensson, C. (1987 ). Passive Smoking and Lung Cancer
in Swedish Women. American Journal of Epidemiology ,125,1 7—24.
Peto, R., and Lee, P. N. (1973 ). Weibull Distributions for Continuous Carcinogenesis
Experiments. Biometrics ,29, 457—470.
Peto, R., and Peto, J. (1972 ). Asymptotically Efficient Rank Invariant Procedures.
Journal of the Royal Statistical Society, Series A ,135, 185—207.
Peto, R., Lee, P. N., and Paige, W. S. (1972 ). Statistical Analysis of the Bioassay of
Continuous Carcinogens. British Journal of Cancer ,26, 258—261.
Peto, R., Pike, M. C., Armitage, P., Breslow, N. E., Cox, D. R., Howard, S. V., Mantel,
N., McPherson, K., Peto, J., and Smith, P. G. (1976, 1977 ). Design and Analysis
of Randomized Clinical Trials Requiring Prolonged Observation of Each Patient.British Journal of Cancer , Part I, 34, 585—612, 1976; Part II, 35,1—39, 1977.
Pierce,M., Borges, W. H., Heyn,R., Wolfe, J., and Gilbert,E. S. (1969 ). Epidemiological
Factors and Survival Experience in 1770 Children with Acute Leukemia. Cancer,
23, 1296—1304.
Pike, M. C. (1966 ). A Method of Analysis of a Certain Class of Experiments in
Carcinogenesis. Biometrics ,22, 142—161.
Piper, J. M., Matanoski, G. M., and Tonascia, J. (1986 ). Bladder Cancer in Young
Women. American Journal of Epidemiology ,123, 1033—1042.
Pregibon, D. (1984 ). Data Analytic Methods for Matched Case-Control Studies.
Biometrics ,40, 639—651.
Prentice, R. L. (1973 ). Exponential Survivals with Censoring and Explanatory Vari-
ables. Biometrika ,60, 279—288.
Prentice, R. L. (1974 ). A Log-Gamma Model and Its Maximum Likelihood Estimation.
Biometrica. 61539—544.
Prentice, R. L. (1976 ). Use of the Logistic Model in Retrospective Studies. Biometrics ,
32,5 99—606.
Prentice, R. L., and Gloeckler, L. A. (1978 ). Regression Analysis of Grouped Survival
Data with Application to Breast Cancer Data. Biometrics ,34,5 7—67.
Prentice, R. L., and Kalbfleisch, J. D. (1979 ). Hazard Rate Models with Covariates.
Biometrics ,35,2 5—39.
Prentice, R. L., and Marek, P. (1979 ). A Quantitative Discrepancy between Censored
Data Rank Tests. Biometrics ,35, 861—867.
Prentice, R. L., and Pyke, R. (1979 ). Logistic Disease Incidence Models and Case-
control Studies.Biometrica ,73, 403—411.
Prentice, R. L., Williams, B. J., and Peterson, A. V. (1981 ). On the Regression Analysis
of Multivariate Failure Time Data. Biometrica ,68, 373—379.
Press, S. J. (1972 ).Applied Multivariate Analysis. Holt, Rinehart & Winston,New York.
Press, S. J., and Wilson, S. (1978 ). Choosing between Logistic Regression and Dis-
criminant Analysis. Journal of the American Statistical Association ,73,6 99—705.
Ralston, A., and Wilf, H. (1967 ).Mathematical Methods for Digital Computers. Wiley,
New York.
Rao, C. R. (1952 ).Advanced Statistical Methods in Biometric Research. Wiley, New
York. 505
Rao, C. R. (1973 ).Linear Statistical Inference and Its Application , 2nd ed. Wiley, New
York.
Riffenburgh, R. H., and Johnstone, P. A. (2001 ). Survival Patterns of Cancer Patients.
Cancer,91(12), 2469—2475.
Rissanen, J. (1986 ). A Predictive Least-SquaresPrinciple. IMA Journal of Mathematical
Control of Information ,3, 211—222.
Rowe-Jones, D. C., Peel, A. L. G., Kingston, R. D., Shaw, J. F. L., Teasdale, C., and
Cole, D. S. (1990 ). Single Dose Cefotaxime Plus Metronidazole versus Three Dose
Cefuroxime Plus Metronidazole as Prophylaxis against Wound Infection in
Colorectal Surgery: Multicentre Prospective Randomised Study. British Medical
Journal,300,1 8—22.
Sacher, G. A. (1956 ). On the Statistical Nature of Mortality, with Special Reference to
Chronic Radiation Mortality. Radiology ,67, 250—257.
Sarhan, A. E., and Greenberg, B. G. (1956 ). Estimation of Location and Scale
Parameters by Order Statistics from Singly and Doubly Censored Samples, PartI, The Normal Distribution up to Samples of Size 10. Annals of Mathematical
Statistics ,27, 427—451.
Sarhan, A. E., and Greenberg, B. G. (1957 ). Estimation of Location and Scale
Parameters by Order Statistics from Singly and Doubly Censored Samples, Part
III.Technical Report 4-OOR, Project 1597. U.S. Army Research Office.
Sarhan, A. E., and Greenberg, B. G. (1958 ). Estimation of Location and Scale
Parameters by Order Statistics from Singly and Doubly Censored Samples, PartII.Annals of Mathematical Statistics ,29,7 9—105.
Sarhan,A. E., and Greenberg,B. G. (1962 ).Contribution to Order Statistics. Wiley, New
York.
Sacks,H., Chalmers,T. C., and Smith, H. (1982 ). Randomizedversus HistoricalControl
for Clinical Trials. American Journal of Medicine ,72, 233—240.
SAS Institute. (2000 ).SAS/STAT User ’s Guide, Version 8.1. SAS Institute, Cary, NC.
Sasieni, P. D. (1996 ). Proportional Excess Hazards. Biometrika ,83(1), 127—141.
Savage, I. R. (1956 ). Contributions to the Theory of Rank Order Statistics: The Two
Sample Case. Annals of Mathematical Statistics ,27,5 90—615.
Saw, J. G. (1959 ). Estimation of the Normal Population Parameters Given a Singly
Censored Sample. Biometrika ,46, 150—159.
Schade,D.S., Mitchell,W. J., and Griego,G. (1987 ).AdditionofSulfonylureato Insulin
Treatment in Poorly Controlled Type II Diabetes. Journal of the American
Medical Association ,257, 2441—2445.
Schafer, J. L. (1999 ). Multiple Imputation: a Primer. StatMethods ,8,3—15.
Schlesselman, J. J. (1982 ).Case-Control Studies. Oxford University Press, New York.
Schoenfeld, D. (1982 ). Partial Residuals for Proportional Hazards Regression Model.
Biometrica. 69, 239—241.
Schwarz, G. (1978 ). Estimating the Dimension of a Model. Annals of Statistics 6,
461—222.
Seaman, S. R., and Bird, S. M. (2001 ). Proportional Hazards Model for Interval-
Censored Failure Times and Time-Dependent Covariates: Application to Hazard
of HIV Infection of Injecting Drug Users in Prison. Statistics in Medicine ,20(12),
1855—1870.506
Segal, M. R., and Bloch, D. A. (1989 ). A Comparison of Estimated Proportional
Hazards Models and Regression Trees. Statistics in Medicine ,8, 539—550.
Sellke, T., and Siegmund, D. (1983 ). Sequential Analysis of the Proportional Hazards
Model. Biometrika ,70, 315—326.
Shapiro, S. S., and Wilk, M. B. (1965a ). An Analysis of Variance Test for Normality
(Complete Samples ).Biometrika ,52, 591.
Shapiro, S. S., and Wilk, M. B. (1965b ). Testing for Distributional Assumptions:
Exponential and Uniform Distributions. Unpublished manuscript.
Shibata, R. (1980 ). Asymptotically Efficient Selection of the Order of the Model for
Estimating Parameters of a Linear Process. Annals of Statistics ,8, 147—165.
Shipley, W. U., Thames, H. D., Sandler, H. M., Hanks, G. E., Zietman, Perez, C. A.,
Kuban, D. A., Hancock, S. L., and Smith, C. D. (1999 ). Radiation Therapy for
ClinicallyLocalizedProstate Cancer. Journal of the American Medical Association ,
281(17), 1598—1604.
Shryock, H. S., Sigel, J. S., and Associates (1971 ).The Methods and Materials of
Demography ,Vols. I and II. U.S. Departmentof Commerce, Bureau of the Census,
U.S. Government Printing Office, Washington, DC.
Sichieri, R., Everhart, J. E., and Roth, H. P. (1990 ). Low Incidence of Hospitalization
with Gallbladder Disease among Blacks in the United States. American Journal of
Epidemiology ,131, 826—835.
Siegmund, K. D., Todorov, A. A., and Province, M. A. (1999 ). A Frailty Approach for
Modelling Diseases with Variable Age of Onset in Families: The NHLBI Family
Heart Study. Statistics in Medicine, 18(12), 1517—1528
Sillitto, G. P. (1949 ). Note on Approximations to the Power Function of the ‘‘2 /p592
Comparative Trial.’’ Biometrika ,36, 347—352.
Sirott, M. N., Bajorin, D. F., Wong, G. Y., Tao, Y., Chapman, P. B., Templeton, M. A.,
and Houghton, A. N. (1993 ). Prognostic Factors in Patients with Metastatic
Malignant Melanoma: A Multivariate Analysis. Cancer,72(10), 3091—3098.
Slud, E. V., and Wei, L. J. (1982 ). Two-Sample Repeated Significance Tests Based on
the Modified Wilcoxon Statistic. Journal of the American Statistical Society ,77,
862—868.
Snedecor,G. W., and Cochran,W.G. (1967 ).Statistical Methods. Iowa State University
Press, Ames, IA.
SPSS (2000 ).SPSS-S User ’s Guide, Version 10.1. SPSS, Chicago.
Stacy, E. W. (1962 ). A Generalizationof the Gamma Distribution. Annals of Mathemat-
ical Statistics ,33, 1187—1192.
Stacy, E. W., and Mihram, G. A. (1965 ). Parameter Estimation for a Generalized
Gamma Distribution. Technometrics ,7, 349—358.
Statistics and Epidemiology Research Corporation (SERC )(1988 ).EGRET Statistical
Software. SERC, Seattle, WA.
Steering Committee on the Physicians Health Study Research Group (1989 ). Final
Report on the Aspirin Component of the Ongoing Physicians’ Health Study. New
England Journal of Medicine ,321, 129—135.
Tai, B. C., Peregoudov, A., and Machin, D. (2001 ). A Competing Risk Approach to the
Analysis of Trials of Alternative Intra-uterine Devices (IUDs )for Fertility Regu-
lation. Statistics in Medicine ,20(23), 3589—3600. 507
Tarone, R. E. (1982 ). The Use of Historical Control Information in Testing a Trend in
Poisson Means. Biometrics ,38(2), 457—462.
Tarone, R. E., and Ware, J. (1977 ). On Distribution-Free Tests for Equality of Survival
Distribution. Biometrics ,64, 156—160.
Teitelman, A. M., Welch, L. S., Hellenbrand, K. G., and Bracken, M. B. (1990 ). Effect
of Maternal Work Activity on Preterm Birth and Low Birth Weight. American
Journal of Epidemiology ,131, 104—113.
Therneau, T. M., Grambsch, P. M., and Fleming, T. R. (1990 ). Martingale-Based
Residuals and Survival Models. Biometrica ,77, 147—160.
Thoman, D. R., and Bain, L. J. (1969 ). Two Sample Tests in the Weibull Distribution.
Technometrics ,11, 805—815.
Thoman, D. R., Bain, L. J., and Antle, C. E. (1969 ). Inferences on the Parameters of the
Weibull Distribution. Technometrics ,11, 445—460.
Thoman, D. R., Bain, L. J., and Antle, C. E. (1970 ). Maximum Likelihood Estimation,
Exact Confidence Intervals for Reliability and Tolerance Limits in the Weibull
Distribution. Technometrics ,12, 363—373.
Truett, J., Cornfield, J., and Kannel, W. B. (1967 ). A Multivariate Analysis of the Risk
of Coronary Heart Disease in Framingham. Journal of Chronic Diseases ,20,
511—524.
Tsiatis, A. A. (1980 ). A Note of a Goodness-of-Fit Test for the Logistic Regression
Model. Biometrika ,67, 250—251.
Tsiatis, A. A. (1981 ). A Large Sample Study of Cox’s Regression Model. Annals of
Statistics ,9,93—108.
Tsiatis, A. A. (1982 ). Repeated Significance Testing for a General Class of Statistics
Used in Censored Survival Analysis. Journal of the American Statistical Associ-
ation,77, 855—861.
Tsumagari, K., Yamamoto, H., Suganuma, N., Kato, M., Ikeda, S., Imai, K., Kira, S.,
and Taketa, K. (2000 ). Epidemiological Studies of Coincidental Outbreaks of
Enterohemorrhagic Escherichia Coli O157:H7 Infection and Infectious Gastroen-teritis in Niimi City, ActaMedicaOkayama ,54, 265—273.
Upton, G. J. G. (1978 ).The Analysis of Cross-Tablated Data . Wiley New York.
Vaida, F., and Xu, R. (2000 ). Proportional Hazards Model with Random Effects.
StatisticsinMedicine ,19(22), 339—3324.
Vasan, R. S., Larson, M. G., Leip, E. P., Evans, J. C., O’Donnell, C. J., Kannel, W. B.,
and Levy, D. (2001 ). Impact of High-Normal Blood Pressure on the Risk of
Cardiovascular Disease. New England Journal of Medicine .345(18), 1337—1340.
Vaupel, J. W., Manton, K. G., and Stallard, E. (1979 ). The Impact of Heterogenity in
Individual Frailty on the Dynamics of Mortality. Demography ,16, 439—454.
Vasan, R. S., Larson, M. G., Levy, D., Evans, J. C., and Benjamin, E. J. (1997 ).
Distribution and Categorization of Echocardiographic Measurements in Relationto Reference Limits: the Framingham Heart Study: Formulation of a Height-
and Sex-specific Classification and Its Prospective Validation. Circulation ,96,
1863—1873.
Vega, G. L., and Grundy, S. M. (1989 ). Comparison of Lovastatin and Gemfibrozil in
Normolipidemic Patients with Hypoalphalipoproteinemia. Journal of the Ameri-
can Medical Association ,262, 3148—3153.508
Wald, A. (1947 ).Sequential Analysis. Wiley, New York.
Wang, Wenyu (1984 ). The Bayesian Estimater of the Orders of AR (k)and ARMA (p,q)
Models of Time Series. Acta Mathematicae Applicatae Sinica ,7(2), 185—195.
Wang, Wenyu (1989 ). Statistical Inference on Aggregated Markov Processes. Ph.D.
dissertation. Department of Mathematics, University of Maryland.
Waters,M.A.,Selvin,S.,and Rappaport,S. M. (1991 ).AMeasureofGoodness-of-fitfor
the Lognormal Model Applied to Occupational Exposures, AmericanIndustrial
HygieneAssociationJournal ,52,4 93—502.
Watson, G. S., and Wells, W. T. (1961 ). On the Possibility of Improving the Mean
Useful Life of Items by Eliminating Those with Short Lives. Technometrics ,3,
281—298.
Wei, L. J. (1984 ). Testing Goodness of Fit for Proportional Hazards Model with
Censored Observations. Journal of the American Statistical Association ,79, 649—
652.
Wei, L. J. (1992 ). On Predictive Least Squares Principle. Annals of Statistics ,20,1—42.
Wei, L. J., Lin, D. Y., and Weissfeld, L. (1989 ). Regression Analysis of Multivariate
Incomplete Failure Time Data by Modeling Marginal Distribution. Journal of the
American Statistical Association ,84, 1065—1073.
Weibull, W. (1939 ). A Statistical Theory of the Strength of Materials. Ingenioers
vetenskaps akakemien Handlingar ,151,2 93—297.
Weibull, W. (1951 ). A Statistical Distribution of Wide Applicability. Journal of Applied
Mathematics, 18,2 93—297.
Weiss,H. (1963 ).A Survey of Some MathematicalMethods in the Theory of Reliability,
inStatistical Theory of Reliability , edited by M. Zelen. University of Wisconsin
Press, Madison, WI.
Well, M. D., Lamborn, K., Edwards, M. S. B., and Wara, W. M. (1998 ). Influence of a
Child’s Sex on Medulloblastoma. Journal of the American Medical Association ,
279(18).
Whayne, T. F., Alaupovic, P., Curry, M. D., Lee, E. T., Anderson, P. S., and Schechter,
E.(1981 ). Plasma Apolipoprotein B and VLDL-, LDL-, and HDL-Cholesterol as
Risk Factors in the Development of Coronary Heart Disease in Male PatientsExamined by Angiography. Atherosclerosis ,39, 411—424.
Wienke, A., Holm, N. V., Skytthe, A., and Yashin, A. L. (2001 ). The Heritability of
Mortality Due to Heart Diseases: A Correlated Frailty Model Applied to Danish
Twins. Twin Research ,4(4), 266—274.
Wilcoxon,F. (1945 ).IndividualComparisonby RankingMethods. Biometrics ,1,80—83.
Wilk, M. B., Gnanadesikan, R., and Huyett, M. J. (1962a ). Estimation of Parameters of
the Gamma Distribution Using Order Statistics. Biometrika ,49, 525—545.
Wilk, M. B., Gnanadesikan, R., and Huyett, M. J. (1962b ). Probability Plots for the
Gamma Distribution. Technometrics ,4,1—20.
Wilkinson, L. (1987 ).SYSTAT: The System for Statistics. Systat, Inc., Evanston, IL.
Wilks, S. S. (1948 ). Order Statistics. Bulletin of the American Mathematical Society ,54,
6—50.
Wilks, S. S. (1950 ). Mathematical Statistics. Princeton University Press, Princeton, NJ. 509
Williams, C. A., Jr. (1950 ). On the Choice of the Number and Width of Classes for the
Chi-Square Test of Goodness of Fit. Journal of the American Statistical Associ-
ation,45,7 7—86.
Williams, J. E., Nieto, F. J., Sanford, C. P., and Tyroler, H. A. (2002 ). The Association
between Trait Anger and Incident Stroke Risk: The Atherosclerosis Risk in
Communities (ARIC )Study. Stroke,33(1),1 3—20.
Winkleby, M. A., Ragland, D. R., and Syme, L. (1988 ). Self-Reported Stressors and
Hypertension:Evidence of an Inverse Association. American Journal of Epidemiol-
ogy,127, 124—134.
Winter, F. D., Snell, P. G., and Stray-Gundersen, J. (1989 ). Effects of 100%Oxygen on
Performance of Professional Soccer Players. Journal of the American Medical
Association ,262, 227—229.
Woolf, B. (1955 ). On Estimating the Relation between Blood Group and Disease.
Annals of Human Genetics ,19, 251—253.
Wurpel, J. N., Dundore, R. L., Barbella, Y. R., Balaban, C. D., Keil, L. C., and Severs,
W. B. (1986 ). Barrel Rotation Evoked by Intracerebroventricular Vasopressin
Injections in Conscious Rats. I. Description and General Pharmacology. Brain
Research ,365,2 1—29.
Xue, X. (2001 ). Analysis of Childhood Brain Tumour Data in New York City Using
Frailty Models. Statistics in Medicine ,20(22), 3459—3473.
Yakovlev, A. Y., Tsodikov, A. D., Boucher, K., and Kerber, R. (1999 ). The Shape of the
Hazard Function in Breast Carcinoma: Curability of the Disease Revisited.Cancer,85, 1789—1798.
Yan, Y., Moore, R. D., and Hoover, D. R. (2000 ). Competing Risk Adjustment Reduces
Overestimation of Opportunistic Infection Rates in AIDS. Journal of Clinical
Epidemiology. 53(8), 817—822.
Yashin,A. I.,and Iachine,I. A. (1997 ).How FrailtyModelsCan BeUsed forEvaluating
Longevity Limits: Taking Advantage of an Interdisciplinary Approach. Demogra-
phy,34(1),3 1—48.
Young, E. M., and Fors, S. W. (2001 ). Factors Related to the Eating Habits of Students
in Grades 9—12.JournalofSchoolHealth ,71, 483—488.
Zelen, M. (1966 ). Applications of Exponential Models to Problems in Cancer Research.
Journal of the Royal Statistical Society, Series A ,129, 368—398.
Zhang, M. J., and Klein, J. P. (2001 ). Confidence Bands for the Difference of Two
Survival Curves under Proportional Hazards Model. LifetimeDataAnalysis ,7,
243—254.
Zippin, C., and Armitage, P. (1966 ). Use of Concomitant Variables and Incomplete
Survival Information in the Estimation of an Exponential Survival Parameter.Biometrics ,22, 665—672.510
Index
Accelerated failure time (AFT )model, 259
Age-specific failure rate, 11AIC, 230, 241, 288, 289
Anderson —Gill model, 368
Annual survival ratio, 94
BIC, 230, 241, 288, 289
BMDP, 7, 94, 115, 173, 180, 188, 196, 235, 269,
273, 277, 283, 306, 319, 324, 347,
351, 356, 367, 368, 389, 394, 398, 403,405, 413, 419, 424
Case-control study, 399
Censored observations, 2
progressively censored data, 4
singly censored data, 4
Censoring
interval, 4, 260
left, 4, 260
random, 4right, 4, 260
type I, 2
type II, 2type III, 3
Chi-square test, 379, 381
Competing risk, 352Conditional mortality rate, 11
Conditional probability, 399, 401
Corrected survival rate, 94
Cox’s F-test, 116, 249
Cox —Mantel test, 109
Cox —Snell residual, 199, 215, 290, 331
Cross-product ratio, 382
Cumulative hazard function, 13
Cumulative survival rate, 9Density function:
definition of, 10types of:
exponential, 135
extended generalized gamma, 153gamma, 150
generalized gamma, 152
log-logistic, 154lognormal, 145
Weibull, 139
Deviance residual, 331Dichotomous outcomes, 377
Exponential distribution, 134, 263
unit exponential distribution, 135two-parameter, 136
goodness-of-fit test, 226, 233
testfor equalityof two distributions,246,249
Five-year survival rate, 94
Force of mortality, 11
Frailty model, 375
Gamma and generalized gamma distribution,
148, 277
goodness-of-fit test, 227, 234, 235test for equality of two distributions, 252
Gap time, 364
Gehan’s generalized Wilcoxon test, 107
Gompertz distribution, 157, 242
Guarantee time, 136, 173, 175
Hazard function:
definition of, 11
exponential, 135
511
Hazard function (Continued )
gamma, 152
log-logistic, 155lognormal, 145
Weibull, 140
Hazard plotting, 29, 209
exponential, 210
log-logistic, 215
lognormal, 213
Weibull, 212
Hollander —Proschan’s test, 236
Hosmer —Lemeshow test of goodness-of-fit, 388
Incomplete gamma function, 151
Instantaneous failure rate, 11
Kaplan —Meier method, 20, 68, 216, 237
Kruskal —Wallis test, 125
multiple comparison, 128
K-sample test for censored data, 130
Life tables:
abridged, 86
clinical, 87
cohort, 77current, 77
population, 77
Likelihood ratio test, 243, 246Link function, 410
logit, 410
probit, 410complementary log-log, 411
Linear exponential distribution, 155
Log-logistic distribution, 154, 280
goodness-of-fit test, 235
Log odds, 386
Logistic regression, 45, 385
conditional, 398
dichotomous outcomes, 377
polychotomous outcomes, 377, 413
nominal, 414
ordinal, 419
Logistic transform, 386Log-likelihood ratio statistic, 223, 224, 388
Lognormal distribution, 143, 274
three-parameter, 146goodness-of-fit test, 227, 234
Logrank test, 111
Mantel —Haenszel method, 28, 121
Maximum likelihood estimate:
exponential, 166
Gompertz, 196interval-censored data, 165, 260
left-censored data, 165, 260
log-logistic, 195lognormal, 180
right-censored data, 162, 260
standard and generalized gamma, 188two-parameter exponential, 174
Weibull, 178
Martingale residual, 331
Matched design, 401
1:R, 401
n/p16:n/p15,40 3
Median remaining lifetime, 91
Model selection method, 230, 233, 286, 288,
289, 310
Normit function, 410Odds ratio, 379, 382
Partial likelihood function, 301, 304, 340, 348,
354, 357, 364, 369, 371
Peto and Peto’s generalized Wilcoxon test, 116Polychotomous outcomes, 377, 413
nominal, 414
ordinal, 419
Printice, Williams, and Peterson (PWP )model,
357, 363
Probability density function, 10Probability plotting, 29, 200
exponential, 204
log-logistic, 208lognormal, 206
normal, 203
Weibull, 205
Probit function, 410, 420
Product-limit estimate, 65
estimate of mean survival time, 74variance, 70
variance of estimated mean survival time, 75
Prognostic factors, 32, 256, 339, 377Prognostic homogeneity, 21
Proportional hazard model, 264, 270, 298
assessment of, 326
Proportional odds, 280
Rayleigh distribution, 158
Recurrent events, 356
Related observations, 374Relative mortality, 32
Relative survival rate, 94
Retrospective study, 398512
SAS, 7, 94, 115, 172, 173, 80, 188, 194, 196, 235,
268, 273, 276, 279, 283, 289, 291,
306, 310, 317, 322, 333, 346, 350, 355,366, 389, 392, 397, 403, 408, 410,
412, 418, 424
Schoenfeld residual, 331
weighted, 332
Score statistic, 225
SPSS, 7, 94, 115, 318, 323, 336, 347, 351, 356,
367, 389, 392, 397, 405, 413, 419, 424
Standardized rate and ratio, 97
direct method, 98indirect method, 99
SIR, 97
SMR, 97
Standardized mortality ratio, 97
Stritification, 328, 348
Survival curve, 9Survivorship function:
definition of, 8
estimation of, 319types:
exponential, 135
gamma, 152log-logistic, 154
lognormal, 145
two-parameter-exponential, 136Weibull, 140
Test of goodness of fit, 221, 222, 226, 227, 233,
234, 235, 330
Tied survival times, 302
Time-dependent covariates, 326, 339
Unconditional failure rate, 11
Wald statistic, 223, 224, 262, 388
Wei, Lin, and Weissfeld (WLW )model, 370
Weibull distribution, 138, 269
three-parameter, 141
goodness-of-fit test, 226, 234
test for equality of two distributions, 251 513
WILEY SERIES IN PROBABILITY AND STATISTICS
Established by WALTER A. SHEWHART and SAMUEL S. WILKSEditors: David J. Balding, Peter Bloomfield, Noel A. C. Cressie,
Nicholas I. Fisher, Iain M. Johnstone, J. B. Kadane, Louise M. Ryan, David W. Scott, Adrian F. M. Smith, Jozef L. TeugelsEditors Emeriti: Vic Barnett, J. Stuart Hunter, David G. Kendall
A complete list of the titles in this series appears at the end of this volume.p&s-cp.qxd 3/25/03 9:47 AM Page 1
WILEY SERIES IN PROBABILITY AND STATISTICS
ESTABLISHED BY WALTER A. S HEWHART AND SAMUEL S. W ILKS
Editors: David J. Balding, Peter Bloomfield, Noel A. C. Cressie,
Nicholas I. Fisher, Iain M. Johnstone, J. B. Kadane, Louise M. Ryan,
David W. Scott, Adrian F. M. Smith, Jozef L. Teugels
Editors Emeriti: Vic Barnett, J. Stuart Hunter, David G. Kendall
The Wiley Series in Probability and Statistics is well established and authoritative. It covers
many topics of current research interest in both pure and applied statistics and probabilitytheory. Written by leading statisticians and institutions, the titles span both state-of-the-artdevelopments in the field and classical methods.
Reflecting the wide range of current research in statistics, the series encompasses applied,
methodological and theoretical statistics, ranging from applications and new techniquesmade possible by advances in computerized practice to rigorous treatment of theoreticalapproaches.
This series provides essential and invaluable reading for all statisticians, whether in aca-
demia, industry, government, or research.
ABRAHAM and LEDOLTER ·Statistical Methods for Forecasting
AGRESTI ·Analysis of Ordinal Categorical Data
AGRESTI ·An Introduction to Categorical Data Analysis
AGRESTI ·Categorical Data Analysis, Second Edition
ANDE L ·Mathematics of Chance
ANDERSON ·An Introduction to Multivariate Statistical Analysis, Second Edition
*ANDERSON ·The Statistical Analysis of Time Series
ANDERSON, AUQUIER, HAUCK, OAKES, VANDAELE, and WEISBERG ·
Statistical Methods for Comparative Studies
ANDERSON and LOYNES ·The Teaching of Practical Statistics
ARMITAGE and DAVID (editors) ·Advances in Biometry
ARNOLD, BALAKRISHNAN, and NAGARAJA ·Records
*ARTHANARI and DODGE ·Mathematical Programming in Statistics
*BAILEY ·The Elements of Stochastic Processes with Applications to the Natural
Sciences
BALAKRISHNAN and KOUTRAS ·Runs and Scans with Applications
BARNETT ·Comparative Statistical Inference, Third Edition
BARNETT and LEWIS ·Outliers in Statistical Data, Third Edition
BARTOSZYNSKI and NIEWIADOMSKA-BUGAJ ·Probability and Statistical Inference
BASILEVSKY ·Statistical Factor Analysis and Related Methods: Theory and
Applications
BASU and RIGDON ·Statistical Methods for the Reliability of Repairable Systems
BATES and WATTS ·Nonlinear Regression Analysis and Its Applications
BECHHOFER, SANTNER, and GOLDSMAN ·Design and Analysis of Experiments for
Statistical Selection, Screening, and Multiple Comparisons
BELSLEY ·Conditioning Diagnostics: Collinearity and Weak Data in Regression
BELSLEY, KUH, and WELSCH ·Regression Diagnostics: Identifying Influential
Data and Sources of Collinearity
BENDAT and PIERSOL ·Random Data: Analysis and Measurement Procedures,
Third Edition
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 2
BERRY, CHALONER, and GEWEKE ·Bayesian Analysis in Statistics and
Econometrics: Essays in Honor of Arnold Zellner
BERNARDO and SMITH ·Bayesian Theory
BHAT and MILLER ·Elements of Applied Stochastic Processes, Third Edition
BHATTACHARYA and JOHNSON ·Statistical Concepts and Methods
BHATTACHARYA and WAYMIRE ·Stochastic Processes with Applications
BILLINGSLEY ·Convergence of Probability Measures, Second Edition
BILLINGSLEY ·Probability and Measure, Third Edition
BIRKES and DODGE ·Alternative Methods of Regression
BLISCHKE AND MURTHY (editors) ·Case Studies in Reliability and Maintenance
BLISCHKE AND MURTHY ·Reliability: Modeling, Prediction, and Optimization
BLOOMFIELD ·Fourier Analysis of Time Series: An Introduction, Second Edition
BOLLEN ·Structural Equations with Latent Variables
BOROVKOV ·Ergodicity and Stability of Stochastic Processes
BOULEAU ·Numerical Methods for Stochastic Processes
BOX ·Bayesian Inference in Statistical Analysis
BOX ·R. A. Fisher, the Life of a Scientist
BOX and DRAPER ·Empirical Model-Building and Response Surfaces
*BOX and DRAPER ·Evolutionary Operation: A Statistical Method for Process
Improvement
BOX, HUNTER, and HUNTER ·Statistics for Experimenters: An Introduction to
Design, Data Analysis, and Model Building
BOX and LUCEÑO ·Statistical Control by Monitoring and Feedback Adjustment
BRANDIMARTE ·Numerical Methods in Finance: A MATLAB-Based Introduction
BROWN and HOLLANDER ·Statistics: A Biomedical Introduction
BRUNNER, DOMHOF, and LANGER ·Nonparametric Analysis of Longitudinal Data in
Factorial Experiments
BUCKLEW ·Large Deviation Techniques in Decision, Simulation, and Estimation
CAIROLI and DALANG ·Sequential Stochastic Optimization
CHAN ·Time Series: Applications to Finance
CHATTERJEE and HADI ·Sensitivity Analysis in Linear Regression
CHATTERJEE and PRICE ·Regression Analysis by Example, Third Edition
CHERNICK ·Bootstrap Methods: A Practitioner’s Guide
CHERNICK and FRIIS ·Introductory Biostatistics for the Health Sciences
CHILÈS and DELFINER ·Geostatistics: Modeling Spatial Uncertainty
CHOW and LIU ·Design and Analysis of Clinical Trials: Concepts and Methodologies
CLARKE and DISNEY ·Probability and Random Processes: A First Course with
Applications, Second Edition
*COCHRAN and COX ·Experimental Designs, Second Edition
CONGDON ·Bayesian Statistical Modelling
CONOVER ·Practical Nonparametric Statistics, Second Edition
COOK ·Regression Graphics
COOK and WEISBERG ·Applied Regression Including Computing and Graphics
COOK and WEISBERG ·An Introduction to Regression Graphics
CORNELL ·Experiments with Mixtures, Designs, Models, and the Analysis of Mixture
Data, Third Edition
COVER and THOMAS ·Elements of Information Theory
COX ·A Handbook of Introductory Statistical Methods
*COX ·Planning of Experiments
CRESSIE ·Statistics for Spatial Data, Revised Edition
CSÖRGO´´and HORVÁTH ·Limit Theorems in Change Point Analysis
DANIEL ·Applications of Statistics to Industrial Experimentation
DANIEL ·Biostatistics: A Foundation for Analysis in the Health Sciences, Sixth Edition
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 3
*DANIEL ·Fitting Equations to Data: Computer Analysis of Multifactor Data,
Second Edition
DASU and JOHNSON ·Exploratory Data Mining and Data Cleaning
DAVID ·Order Statistics, Second Edition
*DEGROOT, FIENBERG, and KADANE ·Statistics and the Law
DEL CASTILLO ·Statistical Process Adjustment for Quality Control
DETTE and STUDDEN ·The Theory of Canonical Moments with Applications in
Statistics, Probability, and Analysis
DEY and MUKERJEE ·Fractional Factorial Plans
DILLON and GOLDSTEIN ·Multivariate Analysis: Methods and Applications
DODGE ·Alternative Methods of Regression
*DODGE and ROMIG ·Sampling Inspection Tables, Second Edition
*DOOB ·Stochastic Processes
DOWDY and WEARDEN ·Statistics for Research, Second Edition
DRAPER and SMITH ·Applied Regression Analysis, Third Edition
DRYDEN and MARDIA ·Statistical Shape Analysis
DUDEWICZ and MISHRA ·Modern Mathematical Statistics
DUNN and CLARK ·Applied Statistics: Analysis of Variance and Regression, Second
Edition
DUNN and CLARK ·Basic Statistics: A Primer for the Biomedical Sciences,
Third Edition
DUPUIS and ELLIS ·A Weak Convergence Approach to the Theory of Large Deviations
*ELANDT-JOHNSON and JOHNSON ·Survival Models and Data Analysis
ENDERS ·Applied Econometric Time Series
ETHIER and KURTZ ·Markov Processes: Characterization and Convergence
EVANS, HASTINGS, and PEACOCK ·Statistical Distributions, Third Edition
FELLER ·An Introduction to Probability Theory and Its Applications, Volume I,
Third Edition, Revised; Volume II, Second Edition
FISHER and VAN BELLE ·Biostatistics: A Methodology for the Health Sciences
*FLEISS ·The Design and Analysis of Clinical Experiments
FLEISS ·Statistical Methods for Rates and Proportions, Second Edition
FLEMING and HARRINGTON ·Counting Processes and Survival Analysis
FULLER ·Introduction to Statistical Time Series, Second Edition
FULLER ·Measurement Error Models
GALLANT ·Nonlinear Statistical Models
GHOSH, MUKHOPADHYAY, and SEN ·Sequential Estimation
GIFI ·Nonlinear Multivariate Analysis
GLASSERMAN and YAO ·Monotone Structure in Discrete-Event Systems
GNANADESIKAN ·Methods for Statistical Data Analysis of Multivariate Observations,
Second Edition
GOLDSTEIN and LEWIS ·Assessment: Problems, Development, and Statistical Issues
GREENWOOD and NIKULIN ·A Guide to Chi-Squared Testing
GROSS and HARRIS ·Fundamentals of Queueing Theory, Third Edition
*HAHN and SHAPIRO ·Statistical Models in Engineering
HAHN and MEEKER ·Statistical Intervals: A Guide for Practitioners
HALD ·A History of Probability and Statistics and their Applications Before 1750
HALD ·A History of Mathematical Statistics from 1750 to 1930
HAMPEL ·Robust Statistics: The Approach Based on Influence Functions
HANNAN and DEISTLER ·The Statistical Theory of Linear Systems
HEIBERGER ·Computation for the Analysis of Designed Experiments
HEDAYAT and SINHA ·Design and Inference in Finite Population Sampling
HELLER ·MACSYMA for Statisticians
HINKELMAN and KEMPTHORNE: ·Design and Analysis of Experiments, Volume 1:
Introduction to Experimental Design
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 4
HOAGLIN, MOSTELLER, and TUKEY ·Exploratory Approach to Analysis
of Variance
HOAGLIN, MOSTELLER, and TUKEY ·Exploring Data Tables, Trends and Shapes
*HOAGLIN, MOSTELLER, and TUKEY ·Understanding Robust and Exploratory
Data Analysis
HOCHBERG and TAMHANE ·Multiple Comparison Procedures
HOCKING ·Methods and Applications of Linear Models: Regression and the Analysis
of Variance, Second Edition
HOEL ·Introduction to Mathematical Statistics, Fifth Edition
HOGG and KLUGMAN ·Loss Distributions
HOLLANDER and WOLFE ·Nonparametric Statistical Methods, Second Edition
HOSMER and LEMESHOW ·Applied Logistic Regression, Second Edition
HOSMER and LEMESHOW ·Applied Survival Analysis: Regression Modeling of
Time to Event Data
HØYLAND and RAUSAND ·System Reliability Theory: Models and Statistical Methods
HUBER ·Robust Statistics
HUBERTY ·Applied Discriminant Analysis
HUNT and KENNEDY ·Financial Derivatives in Theory and Practice
HUSKOVA, BERAN, and DUPAC ·Collected Works of Jaroslav Hajek—
with Commentary
IMAN and CONOVER ·A Modern Approach to Statistics
JACKSON ·A User’s Guide to Principle Components
JOHN ·Statistical Methods in Engineering and Quality Assurance
JOHNSON ·Multivariate Statistical Simulation
JOHNSON and BALAKRISHNAN ·Advances in the Theory and Practice of Statistics: A
Volume in Honor of Samuel Kotz
JUDGE, GRIFFITHS, HILL, LÜTKEPOHL, and LEE ·The Theory and Practice of
Econometrics, Second Edition
JOHNSON and KOTZ ·Distributions in Statistics
JOHNSON and KOTZ (editors) ·Leading Personalities in Statistical Sciences: From the
Seventeenth Century to the Present
JOHNSON, KOTZ, and BALAKRISHNAN ·Continuous Univariate Distributions,
Volume 1, Second Edition
JOHNSON, KOTZ, and BALAKRISHNAN ·Continuous Univariate Distributions,
Volume 2, Second Edition
JOHNSON, KOTZ, and BALAKRISHNAN ·Discrete Multivariate Distributions
JOHNSON, KOTZ, and KEMP ·Univariate Discrete Distributions, Second Edition
JUREC KOVÁ and SEN ·Robust Statistical Procedures: Aymptotics and Interrelations
JUREK and MASON ·Operator-Limit Distributions in Probability Theory
KADANE ·Bayesian Methods and Ethics in a Clinical Trial Design
KADANE AND SCHUM ·A Probabilistic Analysis of the Sacco and Vanzetti Evidence
KALBFLEISCH and PRENTICE ·The Statistical Analysis of Failure Time Data, Second
Edition
KASS and VOS ·Geometrical Foundations of Asymptotic Inference
KAUFMAN and ROUSSEEUW ·Finding Groups in Data: An Introduction to Cluster
Analysis
KEDEM and FOKIANOS ·Regression Models for Time Series Analysis
KENDALL, BARDEN, CARNE, and LE ·Shape and Shape Theory
KHURI ·Advanced Calculus with Applications in Statistics, Second Edition
KHURI, MATHEW, and SINHA ·Statistical Tests for Mixed Linear Models
KLUGMAN, PANJER, and WILLMOT ·Loss Models: From Data to Decisions
KLUGMAN, PANJER, and WILLMOT ·Solutions Manual to Accompany Loss Models:
From Data to Decisions
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 5
KOTZ, BALAKRISHNAN, and JOHNSON ·Continuous Multivariate Distributions,
Volume 1, Second Edition
KOTZ and JOHNSON (editors) ·Encyclopedia of Statistical Sciences: Volumes 1 to 9
with Index
KOTZ and JOHNSON (editors) ·Encyclopedia of Statistical Sciences: Supplement
Volume
KOTZ, READ, and BANKS (editors) ·Encyclopedia of Statistical Sciences: Update
Volume 1
KOTZ, READ, and BANKS (editors) ·Encyclopedia of Statistical Sciences: Update
Volume 2
KOVALENKO, KUZNETZOV, and PEGG ·Mathematical Theory of Reliability of
Time-Dependent Systems with Practical Applications
LACHIN ·Biostatistical Methods: The Assessment of Relative Risks
LAD ·Operational Subjective Statistical Methods: A Mathematical, Philosophical, and
Historical Introduction
LAMPERTI ·Probability: A Survey of the Mathematical Theory, Second Edition
LANGE, RYAN, BILLARD, BRILLINGER, CONQUEST, and GREENHOUSE ·
Case Studies in Biometry
LARSON ·Introduction to Probability Theory and Statistical Inference, Third Edition
LAWLESS ·Statistical Models and Methods for Lifetime Data, Second Edition
LAWSON ·Statistical Methods in Spatial Epidemiology
LE ·Applied Categorical Data Analysis
LE ·Applied Survival Analysis
LEE and WANG ·Statistical Methods for Survival Data Analysis, Third Edition
LEPAGE and BILLARD ·Exploring the Limits of Bootstrap
LEYLAND and GOLDSTEIN (editors) ·Multilevel Modelling of Health Statistics
LIAO ·Statistical Group Comparison
LINDVALL ·Lectures on the Coupling Method
LINHART and ZUCCHINI ·Model Selection
LITTLE and RUBIN ·Statistical Analysis with Missing Data, Second Edition
LLOYD ·The Statistical Analysis of Categorical Data
MAGNUS and NEUDECKER ·Matrix Differential Calculus with Applications in
Statistics and Econometrics, Revised Edition
MALLER and ZHOU ·Survival Analysis with Long Term Survivors
MALLOWS ·Design, Data, and Analysis by Some Friends of Cuthbert Daniel
MANN, SCHAFER, and SINGPURWALLA ·Methods for Statistical Analysis of
Reliability and Life Data
MANTON, WOODBURY, and TOLLEY ·Statistical Applications Using Fuzzy Sets
MARDIA and JUPP ·Directional Statistics
MASON, GUNST, and HESS ·Statistical Design and Analysis of Experiments with
Applications to Engineering and Science, Second Edition
McCULLOCH and SEARLE ·Generalized, Linear, and Mixed Models
McFADDEN ·Management of Data in Clinical Trials
McLACHLAN ·Discriminant Analysis and Statistical Pattern Recognition
McLACHLAN and KRISHNAN ·The EM Algorithm and Extensions
McLACHLAN and PEEL ·Finite Mixture Models
McNEIL ·Epidemiological Research Methods
MEEKER and ESCOBAR ·Statistical Methods for Reliability Data
MEERSCHAERT and SCHEFFLER ·Limit Distributions for Sums of Independent
Random Vectors: Heavy Tails in Theory and Practice
*MILLER ·Survival Analysis, Second Edition
MONTGOMERY, PECK, and VINING ·Introduction to Linear Regression Analysis,
Third Edition
MORGENTHALER and TUKEY ·Configural Polysampling: A Route to Practical
Robustness
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 6
MUIRHEAD ·Aspects of Multivariate Statistical Theory
MURRAY ·X-STAT 2.0 Statistical Experimentation, Design Data Analysis, and
Nonlinear Optimization
MYERS and MONTGOMERY ·Response Surface Methodology: Process and Product
Optimization Using Designed Experiments, Second Edition
MYERS, MONTGOMERY, and VINING ·Generalized Linear Models. With
Applications in Engineering and the Sciences
NELSON ·Accelerated Testing, Statistical Models, Test Plans, and Data Analyses
NELSON ·Applied Life Data Analysis
NEWMAN ·Biostatistical Methods in Epidemiology
OCHI ·Applied Probability and Stochastic Processes in Engineering and Physical
Sciences
OKABE, BOOTS, SUGIHARA, and CHIU ·Spatial Tesselations: Concepts and
Applications of Voronoi Diagrams, Second Edition
OLIVER and SMITH ·Influence Diagrams, Belief Nets and Decision Analysis
PANKRATZ ·Forecasting with Dynamic Regression Models
PANKRATZ ·Forecasting with Univariate Box-Jenkins Models: Concepts and Cases
*PARZEN ·Modern Probability Theory and Its Applications
PEÑA, TIAO, and TSAY ·A Course in Time Series Analysis
PIANTADOSI ·Clinical Trials: A Methodologic Perspective
PORT ·Theoretical Probability for Applications
POURAHMADI ·Foundations of Time Series Analysis and Prediction Theory
PRESS ·Bayesian Statistics: Principles, Models, and Applications
PRESS ·Subjective and Objective Bayesian Statistics, Second Edition
PRESS and TANUR ·The Subjectivity of Scientists and the Bayesian Approach
PUKELSHEIM ·Optimal Experimental Design
PURI, VILAPLANA, and WERTZ ·New Perspectives in Theoretical and Applied
Statistics
PUTERMAN ·Markov Decision Processes: Discrete Stochastic Dynamic Programming
*RAO ·Linear Statistical Inference and Its Applications, Second Edition
RENCHER ·Linear Models in Statistics
RENCHER ·Methods of Multivariate Analysis, Second Edition
RENCHER ·Multivariate Statistical Inference with Applications
RIPLEY ·Spatial Statistics
RIPLEY ·Stochastic Simulation
ROBINSON ·Practical Strategies for Experimenting
ROHATGI and SALEH ·An Introduction to Probability and Statistics, Second Edition
ROLSKI, SCHMIDLI, SCHMIDT, and TEUGELS ·Stochastic Processes for Insurance
and Finance
ROSENBERGER and LACHIN ·Randomization in Clinical Trials: Theory and Practice
ROSS ·Introduction to Probability and Statistics for Engineers and Scientists
ROUSSEEUW and LEROY ·Robust Regression and Outlier Detection
RUBIN ·Multiple Imputation for Nonresponse in Surveys
RUBINSTEIN ·Simulation and the Monte Carlo Method
RUBINSTEIN and MELAMED ·Modern Simulation and Modeling
RYAN ·Modern Regression Methods
RYAN ·Statistical Methods for Quality Improvement, Second Edition
SALTELLI, CHAN, and SCOTT (editors) ·Sensitivity Analysis
*SCHEFFE ·The Analysis of Variance
SCHIMEK ·Smoothing and Regression: Approaches, Computation, and Application
SCHOTT ·Matrix Analysis for Statistics
SCHUSS ·Theory and Applications of Stochastic Differential Equations
SCOTT ·Multivariate Density Estimation: Theory, Practice, and Visualization
*SEARLE ·Linear Models
SEARLE ·Linear Models for Unbalanced Data
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 7
SEARLE ·Matrix Algebra Useful for Statistics
SEARLE, CASELLA, and McCULLOCH ·Variance Components
SEARLE and WILLETT ·Matrix Algebra for Applied Economics
SEBER and LEE ·Linear Regression Analysis, Second Edition
SEBER ·Multivariate Observations
SEBER and WILD ·Nonlinear Regression
SENNOTT ·Stochastic Dynamic Programming and the Control of Queueing Systems
*SERFLING ·Approximation Theorems of Mathematical Statistics
SHAFER and VOVK ·Probability and Finance: It’s Only a Game!
SMALL and M CLEISH ·Hilbert Space Methods in Probability and Statistical Inference
SRIVASTAVA ·Methods of Multivariate Statistics
STAPLETON ·Linear Statistical Models
STAUDTE and SHEATHER ·Robust Estimation and Testing
STOYAN, KENDALL, and MECKE ·Stochastic Geometry and Its Applications, Second
Edition
STOYAN and STOYAN ·Fractals, Random Shapes and Point Fields: Methods of
Geometrical Statistics
STYAN ·The Collected Papers of T. W. Anderson: 1943–1985
SUTTON, ABRAMS, JONES, SHELDON, and SONG ·Methods for Meta-Analysis in
Medical Research
TANAKA ·Time Series Analysis: Nonstationary and Noninvertible Distribution Theory
THOMPSON ·Empirical Model Building
THOMPSON ·Sampling, Second Edition
THOMPSON ·Simulation: A Modeler’s Approach
THOMPSON and SEBER ·Adaptive Sampling
THOMPSON, WILLIAMS, and FINDLAY ·Models for Investors in Real World Markets
TIAO, BISGAARD, HILL, PEÑA, and STIGLER (editors) ·Box on Quality and
Discovery: with Design, Control, and Robustness
TIERNEY ·LISP-STAT: An Object-Oriented Environment for Statistical Computing
and Dynamic Graphics
TSAY ·Analysis of Financial Time Series
UPTON and FINGLETON ·Spatial Data Analysis by Example, Volume II:
Categorical and Directional Data
VAN BELLE ·Statistical Rules of Thumb
VIDAKOVIC ·Statistical Modeling by Wavelets
WEISBERG ·Applied Linear Regression, Second Edition
WELSH ·Aspects of Statistical Inference
WESTFALL and YOUNG ·Resampling-Based Multiple Testing: Examples and
Methods for p-Value Adjustment
WHITTAKER ·Graphical Models in Applied Multivariate Statistics
WINKER ·Optimization Heuristics in Economics: Applications of Threshold Accepting
WONNACOTT and WONNACOTT ·Econometrics, Second Edition
WOODING ·Planning Pharmaceutical Clinical Trials: Basic Statistical Principles
WOOLSON and CLARKE ·Statistical Methods for the Analysis of Biomedical Data,
Second Edition
WU and HAMADA ·Experiments: Planning, Analysis, and Parameter Design
Optimization
YANG ·The Construction Theory of Denumerable Markov Processes
*ZELLNER ·An Introduction to Bayesian Inference in Econometrics
ZHOU, OBUCHOWSKI, and M CCLISH ·Statistical Methods in Diagnostic Medicine
*Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 8
The Third Edition ofthe
leading reference onsurvival data analysis
‘hestudy ofsurvival dataattempts topredict theprobability ofresponse, survival, ormean lifetime;
compare; thesurvival distributions ofexperimental animals orofhuman patients; andidentify
riskand/or prognosis factors. Statistical Methods forSurvival Data Analysis, Third Edition examines the
statistical methods foranalyzing survival datafromlaboratory studiesofanimals, clinical andepidemi-
ological studies ofhumans, andother appropriate applications.
Emphasizing applicarions over rigorous mathematics, thisextremely useful reference provides thor-
‘ough discussions ofthemost commonly usedparametric andnonparametric methods insurvival analy-
sis,aswell asguidelines fortheplanning anddesign ofclinical trials. The authors give special
consideration tothestudy ofsurvival datainbiomedical sciences, though themethods aresuitable for
applications inindustrial reliability, thesocialsciences, andbusiness.
This Third Edison brings thisstandard intheficlduptodatewith newmaterial andrevised refer-
‘ences including:
‘Anew introduction toleftand interval censored dara
Thegeneralized gamma andlog-logistic distribution
Estimation procedures forleftandinterval censored dara
Parametric models with covariates
‘Cox'sproportional hazardsmodelincluding stratification andtime-dependent covariates,
andsome non-proportional hazards models
Goodness-of- Fitrestsandmodelselection methods
Multiple responses tothelogistic regression model
Numerous real-life examples which illustrate keyconcepts
‘Computer programming codes inSAS, BMDP, andSPSS formost examples
Related FTPsiteproviding lange datasets
These additions andrevisions make Statistical Methods forSurvival Data Analysis, Third Edition, more
valuable than everasanessential reference forbiomedical investigators, statisticians, epidemiologists,
andresearchers inother disciplines involved orinterested intheanalysis ofsurvival data.
isGeorge Lynn Cross Research Professor ofBiostatistics andEpidemiology and
Director oftheCenter forAmerican Indian Health Research attheUniversity ofOklahoma Health
Sciences Center. Shereceived amaster’s degree from theUniversity ofCalifornia atBerkeley andher
doctorate from New York University. Theauthor oftheprevious editions ofStatistical Methods for
Survival Data Analysis, Professor LeeisaFellowoftheAmerican Statistical Association andmember of
theSociety forEpidemiological Research andtheAmerican Diabetes Association.
I 1G,F isanAssociate Professor ofBiostatistics attheUniversity of
Oklahoma Health Sciences Center. Hereceived amaster’s degree from theAcademy ofSciences of
Chinaandhisdoctorare fromtheUniversity ofMaryland.
‘Subscribe tooutfreeStatisticseNewsletter at ISBN 0-47L-36997-7wwrewdley.com jenewsletters. r 90000
i ONES 1 INTERSCIENCEoWagereneel Ae