Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Probability and Statistics / Survival Analysis

Lee Elisa T., Wang John Wenyu - Statistical Methods for Survival Data Analysis (3rd Edition, Wiley)

PDF · 536 pages · 6.9 MB
Open PDF file

Published textbook by Elisa T. Lee and John Wenyu Wang, not Phil's own work, aimed at biomedical researchers and statisticians. It covers survival functions, censored data, product-limit and life-table estimates, and nonparametric and parametric comparisons. Further topics are exponential, Weibull, lognormal, gamma and log-logistic models, goodness of fit, Cox proportional hazards and logistic regression, with statistical tables.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
FinancialEbooks.NET ||ry Visit FinancialEbooks.NET for more financial information @)WILEY Statistical Methods for Survival Data Analysis Third Edition Elisa T.Lee ftp:/, JohnWenyuWang WILEY SERIES INPROBABILITY AND STATISTICS StatisticalMethodsfor SurvivalDataAnalysis StatisticalMethodsfor SurvivalDataAnalysis ThirdEdition ELISAT.LEE JOHNWENYUWANG DepartmentofBiostatisticsandEpidemiologyand CenterforAmericanIndianHealthResearchCollegeofPublicHealthUniversityofOklahomaHealthSciencesCenterOklahomaCity,Oklahoma A JOHN WILEY & SONS, INC., PUBLICATION Copyright /p72003byJohnWiley&Sons,Inc.Allrightsreserved. PublishedbyJohnWiley&Sons,Inc.,Hoboken,NewJersey. PublishedsimultaneouslyinCanada. Nopartofthispublicationmaybereproduced,storedinaretrievalsystem,ortransmittedinany formorbyanymeans,electronic,mechanical,photocopying,recording,scanning,orotherwise,exceptaspermittedunderSection107or108ofthe1976UnitedStatesCopyrightAct, withouteitherthepriorwrittenpermissionofthePublisher,orauthorizationthroughpaymentof theappropriateper-copyfeetotheCopyrightClearanceCenter,Inc.,222RosewoodDrive,Danvers,MA01923,978-750-8400,fax978-750-4470,oronthewebatwww.copyright.com. RequeststothePublisherforpermissionshouldbeaddressedtothePermissionsDepartment, JohnWiley&Sons,Inc.,111RiverStreet,Hoboken,NJ07030, (201)748-6011,fax (201)748-6008, e-mail:permreq /p28wiley.com. LimitofLiability/DisclaimerofWarranty:Whilethepublisherandauthorhaveusedtheirbest effortsinpreparingthisbook,theymakenorepresentationsorwarrantieswithrespecttothe accuracyorcompletenessofthecontentsofthisbookandspecificallydisclaimanyimplied warrantiesofmerchantabilityorfitnessforaparticularpurpose.Nowarrantymaybecreatedorextendedbysalesrepresentativesorwrittensalesmaterials.Theadviceandstrategiescontained hereinmaynotbesuitableforyoursituationYoushouldconsultwithaprofessionalwhere appropriate.Neitherthepublishernorauthorshallbeliableforanylossofprofitoranyothercommercialdamages,includingbutnotlimitedtospecial,incidental,consequential,orother damages. ForgeneralinformationonourotherproductsandservicespleasecontactourCustomerCare DepartmentwithintheU.S.at877-762-2974,outsidetheU.S.at317-572-3993orfax317-572-4002. Wileyalsopublishesitsbooksinavarietyofelectronicformats.Somecontentthatappearsin print,however,maynotbeavailableinelectronicformat. Library of Congress Cataloging-in-Publication Data: Lee,ElisaT. Statisticalmethodsforsurvivaldataanalysis.--3rded./ElisaT.LeeandJohnWenyuWang. p.cm.-- (Wileyseriesinprobabilityandstatistics ) Includesbibliographicalreferencesandindex.ISBN0-471-36997-7 (cloth:alk.paper ) 1.Medicine--Research--Statisticalmethods.2.Failuretimedataanalysis.3. Prognosis--Statisticalmethods.I.Wang,JohnWenyu.II.Title.III.Series. R853.S7L432003 610/p30.72--dc21 2002027025 PrintedintheUnitedStatesofAmerica. 10987654321 Tothememoryofourparents Mr.Chi-LanTanandMrs.Hwei-ChiLeeTan (E.T.L. ) Mr.BeijunZhangandMrs.XiangyiWang (J.W.W. ) Contents Preface xi 1 Introduction 1.1 Preliminaries,1 1.2 CensoredData,1 1.3 ScopeoftheBook,5 BibliographicalRemarks,7 2 Functions of Survival Time 8 2.1 Definitions,8 2.2 RelationshipsoftheSurvivalFunctions,15 BibliographicalRemarks,17Exercises,17 3 Examples of Survival Data Analysis 19 3.1 Example3.1:ComparisonofTwoTreatmentsandThree Diets,19 3.2 Example3.2:ComparisonofTwoSurvivalPatterns UsingLifeTables,26 3.3 Example3.3:FittingSurvivalDistributionstoRemission Data,29 3.4 Example3.4:RelativeMortalityandIdentificationof PrognosticFactors,32 3.5 Example3.5:IdentificationofRiskFactors,40 BibliographicalRemarks,47 Exercises,47 vii 4 Nonparametric Methods of Estimating Survival Functions 64 4.1 Product-LimitEstimatesofSurvivorshipFunction,65 4.2 Life-TableAnalysis,77 4.3 Relative,Five-Year,andCorrectedSurvivalRates,94 4.4 StandardizedRatesandRatios,97 BibliographicalRemarks,102 Exercises,102 5 Nonparametric Methods for Comparing Survival Distributions 106 5.1 ComparisonofTwoSurvivalDistributions,106 5.2 Mantel —HaenszelTest,121 5.3 Comparisonof K(K/p572)Samples,125 BibliographicalRemarks,131 Exercises,131 6 Some Well-Known Parametric Survival Distributions and Their Applications 134 6.1 ExponentialDistribution,134 6.2 WeibullDistribution,1386.3 LognormalDistribution,143 6.4 GammaandGeneralizedGammaDistributions,148 6.5 Log-LogisticDistribution,1546.6 OtherSurvivalDistributions,155 BibliographicalRemarks,160 Exercises,160 7 Estimation Procedures for Parametric Survival Distributions without Covariates 162 7.1 GeneralMaximumLikelihoodEstimationProcedure,162 7.2 ExponentialDistribution,166 7.3 WeibullDistribution,1787.4 LognormalDistribution,180 7.5 StandardandGeneralizedGammaDistributions,188 7.6 Log-LogisticDistribution,1957.7 OtherParametricSurvivalDistributions,196 BibliographicalRemarks,196 Exercises,197viii  8 Graphical Methods for Survival Distribution Fitting 198 8.1 Introduction,198 8.2 ProbabilityPlotting,200 8.3 HazardPlotting,209 8.4 Cox —SnellResidualMethod,215 BibliographicalRemarks,219 Exercises,219 9 Tests of Goodness of Fit and Distribution Selection 221 9.1 Goodness-of-FitTestStatisticsBasedonAsymptotic LikelihoodInferences,222 9.2 TestsforAppropriatenessofaFamilyofDistributions,225 9.3 SelectionofaDistributionUsingBIC orAICProcedures,230 9.4 TestsforaSpecificDistributionwith KnownParameters,233 9.5 HollanderandProschan’sTestforAppropriateness ofaGivenDistributionwithKnownParameters,236 BibliographicalRemarks,238 Exercises,240 10 Parametric Methods for Comparing Two Survival Distributions 243 10.1 LikelihoodRatioTestforComparingTwoSurvival Distributions,243 10.2 ComparisonofTwoExponentialDistributions,246 10.3 ComparisonofTwoWeibullDistributions,25110.4 ComparisonofTwoGammaDistributions,252 BibliographicalRemarks,254 Exercises,254 11 Parametric Methods for Regression Model Fitting and Identification of Prognostic Factors 256 11.1 PreliminaryExaminationofData,257 11.2 GeneralStructureofParametricRegressionModels andTheirAsymptoticLikelihoodInference,259 11.3 ExponentialRegressionModel,26311.4 WeibullRegressionModel,269 11.5 LognormalRegressionModel,274 11.6 ExtendedGeneralizedGammaRegressionModel,277 ix 11.7 Log-LogisticRegressionModel,280 11.8 OtherParametricRegressionModels,283 11.9 ModelSelectionMethods,286 BibliographicalRemarks,295Exercises,295 12 Identification of Prognostic Factors Related to Survival Time: Cox Proportional Hazards Model 298 12.1 PartialLikelihoodFunctionforSurvivalTimes,298 12.2 IdentificationofSignificantCovariates,31412.3 EstimationoftheSurvivorshipFunctionwithCovariates,319 12.4 AdequacyAssessmentoftheProportionalHazardsModel,326 BibliographicalRemarks,336Exercises,337 13 Identification of Prognostic Factors Related to Survival Time: Nonproportional Hazards Models 339 13.1 ModelswithTime-DependentCovariates,339 13.2 StratifiedProportionalHazardsModels,34813.3 CompetingRisksModel,352 13.4 RecurrentEventsModels,356 13.5 ModelsforRelatedObservations,374 BibliographicalRemarks,376 Exercises,376 14 Identification of Risk Factors Related to Dichotomous and Polychotomous Outcomes 377 14.1 UnivariateAnalysis,378 14.2 LogisticandConditionalLogisticRegressionModels forDichotomousResponses,385 14.3 ModelsforPolychotomousOutcomes,413 BibliographicalRemarks,425Exercises,425 Appendix A Newton--Raphson Method 428 Appendix B Statistical Tables 433References 488Index 511x  Preface Statisticalmethodsforsurvivaldataanalysishavecontinuedtoflourishinthe lasttwodecades.Applicationsofthemethodshavebeenwidenedfromtheirhistorical use in cancer and reliability research to business, criminology,epidemiology,andsocialandbehavioralsciences.Thethirdeditionof Statisti- cal Methods for Survival Data Analysis isintendedtoprovideacomprehensive introductionofthemostcommonlyusedmethodsforanalyzingsurvivaldata.Itbeginswithbasicdefinitionsandinterpretationsofsurvivalfunctions.Fromthere,the readerisguidedthroughmethods,parametricandnonparametric,forestimatingandcomparingthesefunctionsandthesearchforatheoreticaldistribution (or model )to fit the data. Parametric and nonparametric ap- proachestotheidentificationofprognosticfactorsthatarerelatedtosurvivalare then discussed. Finally, regression methods, primarily linear logistic re-gressionmodels,toidentifyriskfactorsfordichotomousandpolychotomousoutcomesareintroduced. The third edition continues to be application-oriented, with a minimum levelofmathematics.Inafewchapters,someknowledgeofcalculusandmatrixalgebrais needed.Thefewsections thatintroducethe generalmathematicalstructureforthemethodscanbeskippedwithoutlossofcontinuity.Alargenumberofpracticalexamplesaregiventoassistthereaderinunderstandingthemethodsandapplicationsandininterpretingtheresults.Readerswithonlycollegealgebrashouldfindthebookreadableandunderstandable. Therearemanyexcellentbooksonclinicaltrials.Wethereforehavedeleted thetwochaptersonthesubjectthatwereinthesecondedition.Instead,wehaveincludeddiscussionsofmorestatisticalmethodsforsurvivaldataanalysis.A brief summary of the improvements made for the third edition is givenbelow. 1. Twoadditionaldistributions,thelog-logisticdistributionandageneral- izedgammadistribution,havebeenaddedtotheapplicationofparamet-ric models that can be used in model fitting and prognostic factoridentification (Chapters6,7,and11 ). xi 2. Inseveralsections (Sections7.1,9.1,10.1,11.2,and12.1 ),discussionsof the asymptotic likelihood inference of the methods covered in thechaptersaregiven.Thesesectionsareintendedtoprovideamoregeneralmathematicalstructureforstatisticians. 3. The Cox —Snell residual method has been added to the chapter on graphicalmethodsforsurvivaldistributionfitting (Chapter8 ).Inaddi- tion,thesectionsonprobabilityandhazardplottinghavebeenrevised sothatnospecialgraphicalpapersarerequiredtomaketheplots. 4. More tests of goodness of fit are given, including the BIC and AIC procedures (Chapters9and11 ). 5. For Cox’s proportional hazards model (Chapter 12 ), we have now includedmethodstoassessitsadequencyandprocedurestoestimatethesurvivorshipfunctionwithcovariates. 6. Theconceptofnonproportionalhazardsmodelsisintroduced (Chapter 13), which includes models with time-dependent covariates, stratified models,competingrisksmodels,recurrenteventmodels,andmodelsforrelatedobservations. 7. Thechapteronlinearlogisticregression (Chapter14 )hasbeenexpanded to cover regression models for polychotomous outcomes. In addition,methods for a general m:nmatching design have been added to the sectiononconditionallogisticregressionforcase —controlstudies. 8. ComputerprogrammingcodesforsoftwarepackagesBMDP,SAS,and SPSSareprovidedformostexamplesinthetext. Wewouldliketothankthemanyresearchers,teachers,andstudentswho haveusedthesecondeditionofthebook.Thesuggestionsforimprovementthatmanyofthemhaveprovidedareinvaluable.SpecialthanksgotoXingWang, Linda Hutton, Tracy Mankin, and Imran Ahmed for typing themanuscript. Steve Quigley of John Wiley convinced us to work on a thirdedition.Wethankhimforhisenthusiasm. Finally,wearemostgratefultoourfamilies,Sam,Vivian,Benedict,Jennifer, andAnnelisa (E.T.L. ),andAliceandXing (J.W.W. ),fortheconstantjoy,love, andsupporttheyhavegivenus. ET.L JWW Oklahoma City, OK April 18, 2001xii  CHAPTER 1 Introduction 1.1 PRELIMINARIES This book is for biomedical researchers, epidemiologists, consulting statisti- cians, students taking a first course on survival data analysis, and othersinterested in survival time study. It deals with statistical methods for analyzingsurvival data derived from laboratory studies of animals, clinical and epi-demiologic studies of humans, and other appropriate applications. Survival time can be defined broadly as the time to the occurrence of a given event. This event can be the development of a disease, response to a treatment,relapse,ordeath.Therefore,survivaltime canbetumor-freetime,thetimefromthe start of treatment to response, length of remission, and time to death.Survival data can include survival time, response to a given treatment, andpatient characteristics related to response, survival, and the development of adisease. The study of survival data has focused on predicting the probability ofresponse, survival, or mean lifetime, comparing the survival distributions ofexperimentalanimalsor of humanpatientsand the identificationof risk and/orprognostic factors related to response, survival, and the development of adisease.In this book,specialconsiderationis givento thestudy ofsurvival datain biomedical sciences, although all the methods are suitable for applicationsin industrial reliability, social sciences, and business. Examples of survival datain these fields are the lifetime of electronic devices, components, or systems(reliability engineering ); felons’ time to parole (criminology ); duration of first marriage (sociology ); length of newspaper or magazine subscription (market- ing); and worker’s compensation claims (insurance )and their various influenc- ing risk or prognostic factors. 1.2 CENSORED DATA Many researchers consider survival data analysis to be merely the application of twoconventionalstatisticalmethodsto a specialtypeof problem: parametric if the distribution of survival times is known to be normal and nonparametric 1 if the distribution is unknown. This assumption would be true if the survival times of all the subjects were exact and known; however, some survival timesare not. Further, the survival distribution is often skewed, or far from beingnormal. Thus there is a need for new statistical techniques. One of the mostimportant developments is due to a special feature of survival data in the lifesciences that occurs when some subjects in the study have not experienced theevent of interest at the end of the study or time of analysis. For example, somepatients may still be alive or disease-free at the end of the study period. The exact survival times of these subjects are unknown. These are called censored observations orcensored times and can also occur when people are lost to follow-up after a period of study. When these are not censored observations,the set of survival times is complete. There are three types of censoring. Type I Censoring Animal studies usually start with a fixed number of animals, to which thetreatment or treatments is given. Because of time and/or cost limitations, theresearcher often cannot wait for the death of all the animals. One option is toobserve for a fixed period of time, say six months, after which the survivinganimals are sacrificed. Survival times recorded for the animals that died duringthe study period are the times from the start of the experiment to their death.These are called exactoruncensored observations . The survival times of the sacrificedanimals are not known exactly but are recorded as at least the lengthof the study period. Theseare called censored observations. Someanimalscould be lost or die accidentally. Their survival times, from the start of experimentto loss or death, are also censored observations. In type I censoring , if there are no accidental losses, all censored observations equal the length of the studyperiod. For example, suppose that six rats have been exposed to carcinogens by injecting tumor cells into their foot pads. The times to develop a tumor of agiven size are observed. The investigator decides to terminate the experimentafter 30 weeks. Figure 1.1 is a plot of the development times of the tumors.Rats A, B, and D developed tumors after 10, 15, and 25 weeks, respectively.Rats C and E did not develop tumors by the end of the study; their tumor-freetimes are thus 30-plus weeks. Rat F died accidentally without tumors after 19weeks of observation. The survival data (tumor-free times )are 10, 15, 30 /p59, 25, 30/p59, and 19/p59weeks. (The plus indicates a censored observation. ) Type II Censoring Another option in animal studies is to wait until a fixed portion of the animalshave died, say 80 of 100, after which the surviving animals are sacrificed. Inthis case, type II censoring , if there are no accidental losses, the censored observations equal the largest uncensored observation. For example, in anexperimentof six rats (Figure 1.2 ), the investigatormay decide to terminate the study after four of the six rats have developed tumors. The survival ortumor-free times are then 10, 15, 35 /p59, 25, 35, and 19 /p59weeks.2  Figure 1.1 Example of type I censored data. Figure 1.2 Example of type II censored data. Type III Censoring In most clinical and epidemiologic studies the period of study is fixed andpatients enter the study at different times during that period. Some may diebefore the end of the study; their exact survival times are known. Others maywithdraw before the end of the study and are lost to follow-up. Still others maybe alive at the end of the study. For ‘‘lost’’ patients, survival times are at leastfrom their entrance to the last contact. For patients still alive, survival timesare at least from entry to the end of the study. The latter two kinds ofobservations are censored observations. Since the entry times are not simulta-neous, the censored times are also different. This is type III censoring . For example, suppose that six patients with acute leukemia enter a clinical study  3 Figure 1.3 Example of type III censored data. during a total study period of one year. Suppose also that all six respond to treatment and achieve remission. The remission times are plotted in Figure 1.3.Patients A, C, and E achieve remission at the beginning of the second, fourth,and ninth months, and relapse after four, six, and three months, respectively.Patient B achieves remission at the beginning of the third month but is lost tofollow-up four months later; the remission duration is thus at least fourmonths. Patients D and F achieve remission at the beginning of the fifth andtenth months, respectively, and are still in remission at the end of the study;their remission times are thus at least eight and three months. The respectiveremission times of the six patients are 4, 4 /p59,6 ,8/p59, 3, and 3 /p59months. Type I and type II censored observations are also called singly censored data, and type III, progressively censored data , by Cohen (1965 ). Another commonly used name for type III censoring is random censoring . All of these types of censoring are right censoring orcensoring to the right . There are also left censoring and interval censoring cases. L eft censoring occurs when it is knownthattheevent ofinterestoccurredpriorto acertaintime t, but theexact timeof occurrenceis unknown.For example,anepidemiologistwishes toknowthe age at diagnosis in a follow-up study of diabetic retinopathy. At the time ofthe examination, a 50-year-old participant was found to have already develop-ed retinopathy,but there is no recordof the exacttime at whichinitial evidencewas found. Thus the age at examination (i.e., 50 )is a left-censored observation. It means that the age of diagnosis for this patient is at most50 years. Interval censoring occurs when the event of interest is known to have occurred between times aand b. For example, if medical records indicate that at age 45, the patient in the example above did not have retinopathy, his ageat diagnosis is between 45 and 50 years. We will study descriptive and analytic methods for complete, singly cen- sored, and progressively censored survival data using numerical and graphical4  techniques.Analytic methods discussed include parametricand nonparametric. Parametric approaches are used either when a suitable model or distributionis fitted to the data or when a distribution can be assumed for the populationfrom which the sampleis drawn. Commonly used survival distributions are theexponential,Weibull,lognormal,and gamma.If a survivaldistributionis foundto fit the data properly, the survival pattern can then be described by theparameters in a compact way. Statistical inference can be based on thedistribution chosen. If the search for an appropriate model or distribution is too time consuming or not economicalor no theoreticaldistribution adequate-ly fits the data, nonparametric methods, which are generally easy to apply,should be considered. 1.3 SCOPE OF THE BOOK This book is divided into four parts. Part I (Chapters 1, 2, and 3 )defines survival functions and gives examples of survival data analysis. Survival distribution is most commonly described bythree functions: the survivorship function (also called the cumulative survival rate or survival function ), the probability density function, and the hazard function (hazard rate or age-specific rate ). In Chapter 2 we define these three functions and their equivalence relationships. Chapter 3 illustrates survivaldata analysis with five examples taken from actual research situations. Clinicaland laboratory data are systematically analyzed in progressive steps and theresults are interpreted. Section and chapter numbers are given for quickreference. The actual calculations are given as examples or left as exercises inthe chapters where the methods are discussed. Four sets of data are providedin the exercise section for the reader to analyze. These data are referred to inthe various chapters. In Part II (Chapters 4 and 5 )we introduce some of the most widely used nonparametric methods for estimating and comparing survival distributions.Chapter 4 deals with the nonparametric methods for estimating the threesurvival functions: the Kaplan and Meier product-limit (PL)estimate and the life-table technique (population life tables and clinical life tables ). Also covered is standardization of rates by direct and indirect methods, including thestandardized mortality ratio. Chapter 5 is devoted to nonparametric tech-niques for comparing survival distributions. A common practice is to comparethe survival experiences of two or more groups differing in their treatment orin a given characteristic. Several nonparametric tests are described. Part III (Chapters 6 to 10 )introduces the parametric approach to survival data analysis. Although nonparametric methods play an important role insurvival studies, parametric techniques cannot be ignored. In Chapter 6 weintroduce and discuss the exponential, Weibull, lognormal, gamma, andlog-logistic survival distributions. Practical applications of these distributionstaken from the literature are included.    5 An important part of survival data analysis is model or distribution fitting. Once an appropriate statistical model for survival time has been constructedand its parameters estimated, its information can help predict survival, developoptimal treatment regimens, plan future clinical or laboratory studies, and soon. The graphical technique is a simple informal way to select a statisticalmodel and estimate its parameters. When a statistical distribution is found tofit the data well, the parameters can be estimated by analytical methods. InChapter 7 we discuss analytical estimation procedures for survival distribu- tions. Most of the estimationprocedures are based on the maximum likelihoodmethod. Mathematical derivations are omitted; only formulas for the estimatesand examples are given. In Chapter 8 we introduce three kinds of graphicalmethods: probability plotting, hazard plotting, and the Cox —Snell residual method for survival distribution fitting. In Chapter 9 we discuss several testsof goodness of fit and distribution selection. In Chapter 10 we describe severalparametric methods for comparing survival distributions. A topic that has received increasing attention is the identification of prognostic factors related to survival time. For example, who is likely tosurvivelongest after mastectomy,and what are the most important factorsthatinfluence that survival? Another subject important to both biomedical re-searchers and epidemiologists is identification of the risk factors related to thedevelopment of a given disease and the response to a given treatment. Whatare the factorsmost closely related to the developmentof a given disease? Whois more likely to develop lung cancer, diabetes, or coronary disease? In manydiseases, such as cancer, patients who respond to treatment have a betterprognosis than patients who do not. The question, then, relates to what thefactors are that influence response. Who is more likely to respond to treatmentand thus perhaps survive longer? Part IV (Chapters 11 to 14 )deals with prognostic/risk factors and survival times. In Chapter 11 we introduce parametric methods for identifying impor-tant prognostic factors. Chapters 12 and 13 cover, respectively, the Coxproportional hazards model and several nonproportional hazards models forthe identification of prognostic factors. In the final chapter, Chapter 14, weintroduce the linear logistic regression model for binary outcome variables andits extension to handle polychotomous outcomes. In Appendix A we describe a numerical procedure for solving nonlinear equations, the Newton —Raphson method. This method is suggested in Chap- ters 7, 11, 12, and 13. Appendix B comprises a number of statistical tables. Most nonparametric techniques discussed here are easy to understand and simple to apply. Parametric methods require an understanding of survivaldistributions. Unfortunately, most of survival distributions are not simple.Readers without calculus may find it difficult to apply them on their own.However, if the main purpose is not model fitting, most parametric techniquescan be substituted for by their nonparametric competitors. In fact, a largepercentage of survival studies in clinical or epidemiological journals areanalyzed by nonparametric methods. Researchers not interested in survival6  modelfitting shouldreadthe chaptersand sectionsonnonparametricmethods. Computerprograms for survivaldata analysis are available in several commer-cially available software packages: for example, BMDP, SAS, and SPSS. Thesecomputer programs are referred to in various chapters when applicable.Computer programming codes are given for many of the examples. Bibliographical Remarks Cross and Clark (1975 )was the first book to discuss parametric models and nonparametric and graphical techniques for both complete and censoredsurvival data. Since then, several other books have been published in additionto the first edition of this book (Lee, 1980, 1992 ). Elandt-Johnsonand Johnson (1980 )discuss extensively the construction of life tables, model fitting, compet- ing risk, and mathematical models of biological processes of disease pro-gression and aging. Kalbfleisch and Prentice (1980 )focus on regression problems with survival data, particularly Cox’s proportional hazards model.Miller (1981 )covers a number of parametric and nonparametric methods for survival analysis. Cox and Oakes (1984 )also cover the topic concisely with an emphasis on the examination of explanatory variables. Nelson (1982 )providesa gooddiscussionofparametric,nonparametric,and graphical methods. The book is more suited for industrial reliability engineersthan for biomedical researchers, as are Hahn and Shapiro (1967 )and Mann et al.(1974 ). In addition, Lawless (1982 )gives a broad coverage of the area with applications in engineering and biomedical sciences. More recent publications include Marubini and Valsecchi (1994 ), Klein- baum (1995 ), Klein and Moeschberger (1997 ), and Hosmer and Lemeshow (1999 ). Most of these books take a more rigorous mathematical approach and require knowledge of mathematical statistics.    7 CHAPTER 2 Functions of Survival Time Survival time data measure the time to a certain event, such as failure, death, response, relapse, the development of a given disease, parole, or divorce. Thesetimes are subject to random variations, and like any random variables, form adistribution. The distribution of survival times is usually described or charac-terized by three functions: (1)the survivorship function, (2)the probability density function, and (3)the hazard function. These three functions are mathematically equivalent — if one of them is given, the other two can bederived. In practice, the three functions can be used to illustrate different aspects of the data. A basic problem in survival data analysis is to estimate from thesampled data one or more of these three functions and to draw inferencesabout the survival pattern in the population. In Section 2.1 we define the threefunctions and in Section 2.2, discuss the equivalence relationship among thethree functions. 2.1 DEFINITIONS LetTdenote the survival time. The distribution of Tcan be characterized by three equivalent functions. Survivorship Function (or Survival Function) This function, denoted by S(t), is defined as the probability that an individual survives longer than t: S(t)/p58P(an individual survives longer than t) /p58P(T/p57t) (2.1.1 ) From the definition of the cumulative distribution function F(t)o fT, S(t)/p581-P(an individual fails before t) /p581/p57F(t)( 2 .1.2) 8 Figure 2.1 Two examples of survival curves.HereS(t) is a nonincreasing function of time twith the properties S(t)/p58/p71 for t/p580 0 for t/p58/p45 That is, the probability of surviving at least at the time zero is 1 and that of surviving an infinite time is zero. The function S(t) is also known as the cumulativesurvivalrate. To depict the course of survival, Berkson (1942 )recommended a graphic presentation of S(t). The graph of S(t) is called the survival curve. A steep survival curve, such as the one shown in Figure 2.1 a, represents low survival rate or short survival time. A gradual or flat survival curve such as in Figure 2.1 brepresents high survival rate or longer survival. The survivorship function or the survival curve is used to find the 50th percentile (the median )and other percentiles (e.g., 25th and 75th )of survival time and to compare survival distributions of two or more groups. The mediansurvival times in Figure 2.1 aandbare approximately 5 and 36 units of time, respectively. The mean is generally used to describe the central tendency of adistribution, but in survival distributions the median is often better because asmall number of individuals with exceptionally long or short lifetimes willcause the mean survival time to be disproportionately large or small. In practice, if there are no censored observations, the survivorship function is estimated as the proportion of patients surviving longer than t: S/p19(t)/p58number of patients surviving longer than t total number of patients (2.1.3 ) where the circumflex denotes an estimate of the function. When censored observations are present, the numerator of (2.1.3 )cannot always be determined. For example, consider the following set of survival data: 4, 6, 6 /p59,1 0/p59, 15, 20. 9 Figure 2.2 Two examples of density curves.Using (2.1.3 ), we can compute S/p19(5)/p585/6/p580.833. However, we cannot obtain S/p19(11) since the exact number of patients surviving longer than 11 is unknown. Either the third or the fourth patient (6/p59and 10 /p59)could survive longer than or less than 11. Thus, when censored observations are present, (2.1.3 )is no longer appropriate for estimating S(t). Nonparametric methods of estimating S(t) for censored data are discussed in Chapter 4. Probability Density Function (or Density Function) Like any other continuous random variable, the survival time Thas a probability density function defined as the limit of the probability that anindividual fails in the short interval ttot/p59/afii9773tper unit width /afii9773t, or simply the probability of failure in a small interval per unit time. It can be expressed as f(t)/p58lim/p9/p82/p29/p15P[an individual dying in the interval (t,t/p59/afii9773t)] /afii9773t(2.1.4) The graph of f(t) is called the density curve. Figure 2.2 aandbgive two examples of the density curve. The density function has the following twoproperties: 1.f(t) is a nonnegative function: f(t)/p460 for all t/p460 /p580 fort/p580 2. The area between the density curve and the taxis is equal to 1. In practice, if there are no censored observations, the probability density functionf(t) is estimated as the proportion of patients dying in an interval per10     unit width: f/p19(t)/p58number of patients dying in the interval beginning at time t (total number of patients )/p59(interval width )(2.1.5 ) Similar to the estimation of S(t), when censored observations are present, (2.1.5 )is not applicable. We discuss an appropriate method in Chapter 4. The proportion of individuals that fail in any time interval and the peaks of high frequency of failure can be found from the density function. The densitycurve in Figure 2.2 agives a pattern of high failure rate at the beginning of the study and decreasing failure rate as time increases. In Figure 2.2 b, the peak of high failure frequency occurs at approximately 1.7 units of time. The propor-tion of individuals that fail between 1 and 2 units of time is equal to the shadedarea between the density curve and the axis. The density function is also knownas theunconditional failure rate. Hazard Function The hazard function h(t) of survival time Tgives the conditional failure rate. This is defined as the probability of failure during a very small time interval,assuming that the individual has survived to the beginning of the interval, oras the limit of the probability that an individual fails in a very short interval,t/p59/afii9773t, given that the individual has survived to time t: h(t)/p58lim/p9/p82/p29/p15P /p3an individual fails in the time interval (t,t/p59/afii9773t) given the individual has survived to t /p4 /afii9773t(2.1.6) The hazard function can also be defined in terms of the cumulative distribution function F(t) and the probability density function f(t): h(t)/p58f(t) 1/p57F(t)(2.1.7) The hazard function is also known as the instantaneous failure rate ,force of mortality ,conditional mortality rate , andage-specific failure rate. Iftin(2.1.6 ) is age, it is a measure of the proneness to failure as a function of the age of theindividual in the sense that the quantity /afii9773th(t) is the expected proportion of agetindividuals who will fail in the short time interval t/p59/afii9773t. The hazard function thus gives the risk of failure per unit time during the aging process. Itplays an important role in survival data analysis. In practice, when there are no censored observations the hazard function is estimated as the proportion of patients dying in an interval per unit time, given 11 Figure 2.3 Examples of the hazard function.that they have survived to the beginning of the interval: h/p19(t)/p58number of patients dying in the interval beginning at time t (number of patients surviving at t)/p59(interval width ) /p58number of patients dying per unit time in the interval number of patients surviving at t(2.1.8 ) Actuaries usually use the average hazard rate of the interval in which the number of patients dying per unit time in the interval is divided by the averagenumber of survivors at the midpoint of the interval: h/p19(t)/p58 number of patients dying per unit time in the interval (number of patients surviving at t)/p57(number of deaths in the interval )/2 (2.1.9 ) The actuarial estimate in (2.1.9 )gives a higher hazard rate than (2.1.8 )and thus a more conservative estimate. The hazard function may increase, decrease, remain constant, or indicate a more complicated process. Figure 2.3 is a plot of several kinds of hazardfunction. For example, patients with acute leukemia who do not respond totreatment have an increasing hazard rate, h/p16(t),h/p17(t) is a decreasing hazard function that, for example, indicates the risk of soldiers wounded by bulletswho undergo surgery. The main danger is the operation itself and this dangerdecreases if the surgery is successful. An example of a constant hazard function,h/p18(t), is the risk of healthy persons between 18 and 40 years of age whose main risks of death are accidents. The bathtub curve ,h/p19(t), describes the process of12     Table 2.1 Survival Data and Estimated Survival Functions of40 Myeloma Patients Number of Patients Surviving at Number of Patients Survival Time Beginning of Dying int(months ) Interval Interval S/p19(t)f/p19(t)h/p19(t) 0—5 40 5 1.000 0.025 0.027 5—10 35 7 0.875 0.035 0.044 10—15 28 6 0.700 0.030 0.048 15—20 22 4 0.550 0.020 0.040 20—25 18 5 0.450 0.025 0.065 25—30 13 4 0.325 0.020 0.072 30—35 9 4 0.225 0.020 0.114 35—40 5 0 0.125 0.000 0.000 40—45 5 2 0.125 0.010 0.100 45—50 3 1 0.075 0.005 0.080 /p4650 2 2 0.050 — —human life. During an initial period, the risk is high (high infant mortality ). Subsequently, h(t) stays approximately constant until a certain time, after which it increases because of wear-out failures. Finally, patients with tubercu-losis have risks that increase initially, then decrease after treatment. Such anincreasing, then decreasing hazard function is described by h/p20(t). Thecumulative hazard function is defined as H(t)/p58 /p16/p82 /p15h(x)dx (2.1.10) It will be shown in Section 2.2 that H(t)/p58/p57 logS(t) (2.1.11 ) Thus, at t/p580,S(t)/p581,H(t)/p580, and at t/p58/p45,S(t)/p580,H(t)/p58/p45. The cumulative hazard function can be any value between zero and infinity. All logfunctions in this book are natural logs (basee)unless otherwise indicated. The following example illustrates how these functions can be estimated from a complete sample of grouped survival times without censored observations. Example 2.1 The first three columns of Table 2.1 give the survival data of 40 patients with myeloma. The survival times are grouped into intervals of fivemonths. The estimated survivorship function, density function, and hazardfunction are also given, with the corresponding graphs plotted in Figure2.4a—c. 13 Figure 2.4 Estimated survival functions of myeloma patients.14     Figure 2.4 (Continued). The estimated survivorship function ,S/p19(t), is calculated following (2.1.3 )at the beginning or the end of each interval. For example, at the beginning of the firstinterval, all 40 patients are alive, S/p19(0)/p581, and at the beginning of the second interval, 35 of the 40 patients are still alive, S/p19(5)/p5835/40/p580.875. Similarly, S/p19(10)/p5828/40/p580.700. The estimated density function f/p19(t) is computed follow- ing(2.1.5 ). For example, the density function of the first interval (0—5)is 5/(40/p595)/p580.025, and that of the second interval (5—10)is 7/(40/p595)/p580.035. The estimated density function is plotted at the midpoint of each interval(Figure 2.4 b). The estimated hazard function, h/p19(t), is computed following the actuarial method given in (2.1.9 ). For example, the hazard function of the first interval 5/[5 (40/p575/2)]/p580.027 and that of the second interval is 7/[5 (35/p577/ 2)]/p580.044. The estimated hazard function is also plotted at the midpoint of each interval (Figure 2.4 c). From Table 2.1 or Figure 2.4 a, the median survival time of myeloma patients is approximately 17.5 months, and the peak of high frequency of deathoccurs in 5 to 10 months. In addition, the hazard function shows an increasingtrend and reaches its peak at approximately 32.5 months and then fluctuates. 2.2 RELATIONSHIPS OF THE SURVIVAL FUNCTIONS The three functions defined in Section 2.1 are mathematically equivalent. Given any one of them, the other two can be derived. Readers not interested in themathematical relationship among the three survival functions can skip this     15 section without loss of continuity. 1. From (2.1.2 )and (2.1.7 ), h(t)/p58f(t) S(t)(2.2.1) This relationship can also be derived from (2.1.6 )using basic definitions of conditional probabilities. 2. Since the probability density function is the derivative of the cumulative distribution function, f(t)/p58d dt[1/p57S(t)]/p58/p57S/p30(t)( 2 .2.2) 3. Substituting (2.2.2 )into (2.2.1 )yields h(t)/p58/p57S/p30(t) S(t)/p58/p57d dtlogS(t)( 2 .2.3) 4.Integrating (2.2.3 )from zero to tand using S(0)/p581, we have /p57/p16/p82 /p15h(x)dx/p58logS(t) or H(t)/p58/p57 logS(t) or S(t)/p58exp[/p57H(t)]/p58exp/p3/p57/p16/p82 /p15h(x)dx/p4(2.2.4 ) 5. From (2.2.1 )and (2.2.4 )we obtain f(t)/p58h(t) exp[ /p57H(t)] (2 .2.5) Hence, iff(t) is known, the survivorship function can be obtained from the basic relationship between f(t),F(t), and (2.1.2 ). The hazard function can then be determined from (2.2.1 ).I fS(t) is known, f(t) andh(t) can be determined from (2.2.2 )and(2.2.1 ), respectively, or h(t) can be derived first from (2.2.3 )and thenf(t) from (2.2.1 ).I fh(t) is given, S(t) andf(t) can be obtained, respectively, from (2.2.4 )and (2.2.5 ). Thus, given any one of the three survival functions, the other two can easily be derived. The following example illustrates theseequivalence relationships.16     Example 2.2 Suppose that the survival time of a population has the following density function: f(t)/p58e/p92/p82t/p460 Using the definition of the cumulative distribution function, F(t)/p58/p16/p82 /p15f(x)dx/p58/p16/p82 /p15e/p92/p86dx/p58/p57e/p92/p86/p11/p82 /p15/p581/p57e/p92/p82 From (2.1.2 )we obtain the survivorship function S(t)/p58e/p92/p82 The hazard function can then be obtained from (2.2.1 ): h(t)/p58e/p92/p82 e/p92/p82/p581 A complete treatment of this distribution is given in Section 6.1. Bibliographical Remarks The three survival functions and their equivalents are discussed in every text cited in the Bibliographical Remarks in Chapter 1. EXERCISES 2.1 Consider the survival data given in Exercise Table 2.1. Compute and plot the estimated survivorship function, the probability density function, andthe hazard function. Exercise Table 2.1 Year of Number Alive at Number Dying in Follow-up Beginning of Interval Interval 0—1 1100 240 1—2 860 180 2—3 680 184 3—4 496 138 4—5 358 118 5—6 240 60 6—7 180 52 7—8 128 44 8—98 4 3 2 /p4695 2 2 8 17 2.2 Exercise Table 2.2 is a life table for the total population (of 100,000 live births )in the United States, 1959 —1961. Compute and plot the estimated survivorship function, the probability density function, and the hazardfunction. Exercise Table 2.2 Age Number Living at Number Dying in Interval Beginning of Age Interval Age Interval 0—1 100,000 2,593 1—5 97,407 409 5—10 96,998 233 10—15 96,765 214 15—20 96,551 440 20—25 96,111 594 25—30 95,517 612 30—35 94,905 761 35—40 94,144 1,080 40—45 93,064 1.686 45—50 91,378 2,622 50—55 88,756 4,045 55—60 84,711 5,644 60—65 79,067 7,920 65—70 71,147 10,290 70—75 60,857 12,687 75—80 48,170 14,594 80—85 33,576 15,034 85 and over 18,542 18,542 Source: U.S. National Center for Health Statistics, Life Tables 1959 —1961, Vol. 1, No. 1, ‘‘United States Life Tables 1959 —61,’’ December 1964, pp. 8 —9. 2.3 Derive (2.2.1 )using (2.1.6 )and basic definitions of conditional probabil- ity. 2.4 Given the hazard function h(t)/p58c derive the survivorship function and the probability density function. 2.5 Given the survivorship function S(t)/p58exp(/p57t/p65) derive the probability density function and the hazard function.18     CHAPTER 3 Examples of Survival Data Analysis The investigator who has assembled a large amount of data must decide what to do with it and what it indicates. In this chapter we take several sets ofsurvival data from actual research situations and analyze them. In Example 3.1we analyze two sets of data obtained, respectively, from two and threetreatment groups to compare the treatment’s abilities to prolong life. Example3.2 is an example of the life-table technique for large samples. Example 3.3 givesremission data from two treatments; the investigator seeks a well-knowndistribution for the remission patterns to compare the two groups. In Example3.4 we study survival data and several other patient characteristics to identifyimportant prognostic factors; the patient characteristics are analyzed individ-ually and simultaneously for their prognosticvalues. In Example 3.5 weintroduce a case in which the interest is to identify risk factors in thedevelopment of a given disease. Four sets of real data are presented in theexercises so that the reader can plan analysis. 3.1 EXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 3.1.1 Comparison of Two Treatments Thirty melanoma patients (stages 2 to 4 )were studied to compare the immunotherapies BCG (Bacillus Calmette-Guerin )andCorynebacterium par- vumfor their abilities to prolong remission duration and survival time. The age, gender, disease stage, treatment received, remission duration, and survival timeare given in Table 3.1. All the patients were resected before treatment beganand thus had no evidence of melanoma at the time of first treatment. The usual objective with this type of data is to determine the length of remission and survival and to compare the distributions of remission andsurvival time in each group. Before comparing the remission and survival 19 Table 3.1 Data for 30 Resected Melanoma Patients Initial Treatment Remission Survival Patient Age Gender /p63Stage Received /p64Duration /p65 Time /p65 1 59 2 3B 1 33.7 /p59 33.7/p59 2 50 2 3B 1 3.8 3.9 3 76 1 3B 1 6.3 10.54 66 2 3B 1 2.3 5.4 5 33 1 3B 1 6.4 19.5 6 23 2 3B 1 23.8 /p59 23.8/p59 7 40 2 3B 1 1.8 7.9 8 34 1 3B 1 5.5 16.9 /p59 9 34 1 3B 1 16.6 /p59 16.6/p59 10 38 2 2 1 33.7 /p59 33.7/p59 11 54 2 2 1 17.1 /p59 17.1/p59 12 49 1 3B 2 4.3 8.0 13 35 1 3B 2 26.9 /p59 26.9/p59 14 22 1 3B 2 21.4 /p59 21.4/p59 f15 30 1 3B 2 18.1 /p59 18.1/p59 16 26 2 3B 2 5.8 16.0 /p59 17 27 1 3B 2 3.0 6.9 18 45 2 3B 2 11.0 /p59 11.0/p59 19 76 2 3A 2 22.1 24.8 /p59 20 48 1 3A 2 23.0 /p59 23.0/p59 21 91 1 4A 2 6.8 8.3 22 82 2 4A 2 10.8 /p59 10.8/p59 23 50 2 4A 2 2.8 12.2 /p59 24 40 1 4A 2 9.2 12.5 /p59 25 34 1 3A 2 15.9 24.4 26 38 1 4A 2 4.5 7.7 27 50 1 2 2 9.2 14.8 /p59 28 53 2 2 2 8.2 /p59 8.2/p59 29 48 2 2 2 8.2 /p59 8.2/p59 30 40 2 2 2 7.8 /p59 7.8/p59 Source: Data courtesy of Richard Ishmael. /p631, male; 2, female. /p641, BCG; 2, C. parvum. /p65Remission and survival times are in months. distributions, we attempt to determine if the two treatment groups are comparable with respect to prognostic factors. Let us use the survival time toillustrate the steps. (The remission time could be analyzed similarly. ) 1.Estimate and plot the survival function of the two treatment groups . The resulting curves are called survival curves. Points on the curve estimate the proportion of patients who will survive at least a given period of time. For suchsmall samples with progressively censored observations, the Kaplan —Meier product-limit (PL)method is appropriate for estimating the survival function.20 EXAMPLES OF SURVIVAL DATA ANALYSIS Table 3.2 Kaplan--Meier Product-Limit Estimate of Survival Function S(t) BCG Patients Death time ( t) 3.9 5.4 7.9 10.5 19.5 S/p19(t) 0.909 0.818 0.727 0.636 0.477 C. parvum Patients Death time ( t) 6.9 7.7 8.0 S/p19(t) 0.947 0.895 0.839 It does not require any assumptions about the form of the function that is being estimated. We discuss this method in detail in Section 4.1. Computerprograms for the method can be found in BMDP (Dixon et al. 1990 ), SPSS Version 10.1 (2000 ), and SAS Version 8.1 (2000 ). Examples for computer codes will be given in Section 4.1. Table 3.2 gives the PL estimate of the survival function S/p19(t) for the two treatment groups. Note that S/p19(t) is estimated only at death times; however, the censored observations were used to estimate S(t). Themedian survival time can be estimated by linear interpolation. For BCG patients the median survivaltime was about 18.2 months. The median survival time for the C.parvum group cannot be calculated since 15 of the 19 patients were still alive. Most computerprograms give not only S/p19(t) but also the standard error of S/p19(t), and the 75-, 50-, and 25-percentile points. Figure 3.1 plots the estimated survival function S/p19(t) for patients receiving the two treatments: The median survival time (50-percentile point )for the BCG group can also be determined graphically. The survival curves clearly showthatC. parvum patients had slightly better survival experience than BCG patients. For example, 50%of the BCG patients survived at least 18.2 months,whereas about 61% of the C. parvum patients survived that long. 2.Examine the prognostic homogeneity of the two groups . The next question to ask is whether the difference in survival between the two treatment groupsis statistically significant. Is the difference shown by the data significant orsimply random variation in the sample? A statistical test of significance isneeded. However, a statistical test without considering patient characteristicsmakes sense only if the two groups of patients are homogeneous with respectto prognosticfac tors. It has been assumed thus far that the patients in the twogroups are comparable and that the only difference between them is treatment.Thus, before performing a statistical test it is necessary to examine thehomogeneity between the two groups. Although prognosticfac tors for melanoma patients are not well established, it has been reported that women and the young have a better survivalEXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 21 Figure3.1 Survival curves of patients receiving BCG and C. parvum. experience than men and the elderly. Also, the disease stage plays an important role in survival. Let us check the homogeneity of the two treatment groupswith respect to age, gender, and disease stage. The age distributions are estimated and plotted in Figure 3.2. The median age is 39 for the BCG group and 43 for the C. parvum patients. To test the significance of the difference between the two age distributions, the two-samplet-test (Armitage, 1971; Daniel, 1987 )or nonparametrictests suc h as the Mann—Whitney U-test or the Kolmogorov —Smirnov test (Marascuilo and McSweeney, 1977 )are appropriate. However, the generalized Wilcoxon tests given in Section 5.1 can also be used, since they reduce to the Mann —Whitney U-test. Using Geham’s generalized Wilcoxon test, the difference between the two age distributions is not found to be statistically significant. More about thetest will be given in Section 5.1. The number of male and female patients in the two treatment groups is given in Table 3.3. Sixty-four percent of the BCG patients and 42% of the C. parvum patients are women. A chi-square test can be used to compare the two proportions (see Section 14.1 ). It can be used only for r/p59ctables in which the entries are frequencies, not for tables in which the entries are mean values ormedians of a certain variable. For a 2 /p592 table, the chi-square value can be computed by hand. Computer programs for the test can be found in manycomputer program packages, such as BMDP (Dixon et al., 1990 ), SPSS Version 10.1 (2000 ), and SAS Version 8.1 (SAS Institute, 2000 ). The chi-square value for treatment by gender in Table 3.3 is 1.29 with 1 degree of freedom, which is not significant at the 0.05 or 0.10 level. Therefore,the difference between the two proportions is not statistically significant. Thenumber of stage 2 patients and the number of patients with more advanced22 EXAMPLES OF SURVIVAL DATA ANALYSIS Figure3.2 Age distribution of two treatment groups. Table 3.3 Treatment by Gender and Disease Stage BCG C. parvum BCG C. Parvum Disease Gender Number % Number % Total Stage Number % Number % Total Male 4 36 11 58 15 2 2 18 4 21 6 Female 7 64 8 42 15 3 and 4 9 82 15 79 24—— — —— —11 19 30 11 19 30disease in the two treatment groups are also given in Table 3.3. Eighteen percent of the BCG patients are at stage 2 against 21% of the C. parvum patients. However, a chi-square test result shows that the difference is notsignificant. Thus we can say that the data do not show heterogeneity between the two treatment groups. If heterogeneity is found, the groups can be divided intosubgroups of members who are similar in their prognoses.EXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 23 3.Compare the two survival distributions . There are several parametricand nonparametric tests to compare two survival distributions. They are describedin Chapters 5 and 10. Since we have no information of the survival distributionthat the data follow, we would continue to use nonparametric methods tocompare the two survival distributions. The four tests described in Sections5.1.1 to 5.1.4 are suitable. The performance of these tests is discussed at the endof Section 5.1. We chose Gehan’s generalized Wilcoxon test here to demon-strate the analysis procedure only because of its simplicity of calculation. In testing the significance of the difference between two survival distribu- tions, the hypothesis is that the survival distribution of the BCG patients is thesame as that of the C. parvum patients. Let S/p16(t) andS/p17(t) be the survival function of the BCG and C.parvum groups, respectively. The null hypothesis is H/p15:S/p16(t)/p58S/p17(t) The alternative hypothesis chosen is two-sided: H/p16:S/p16(t)/p34S/p17(t) since we have no prior information concerning the superiority of either of the two treatments. The slight difference between the two estimated survival curvescould be due to random variation. The one-sided alternative H/p16:S/p16(t)/p58S/p17(t) should be considered inappropriate. Using Gehan’s generalized Wilcoxon test, the difference in survival distribu- tion of the two treatment groups is found to be insignificant (p/p580.33). Therefore, we do not reject the null hypothesis that the two survival distribu-tions are equal. Although our conclusion is that the data do not provideenough evidence to reject the hypothesis, ‘‘not to reject the null hypothesis’’does not automatically mean ‘‘to accept the null hypothesis.’’ The differencebetween the two statements is that the error probability of the latter statementis usually much larger than that of the former. 3.1.2 Comparison of Three Diets A laboratory investigator interested in the relationship between diet and the development of tumors divided 90 rats into three groups and fed them low-fat,saturated fat, and unsaturated fat diets, respectively (King et al., 1979 ). The rats were of the same age and species and were in similar physical condition. Anidentical amount of tumor cells were injected into a foot pad of each rat. Therats were observed for 200 days. Many developed a recognizable tumor earlyin the study period. Some were tumor-free at the end of the 200 days. Rat 16in the low-fat group and rat 24 in the saturated group died accidentally after140 days and 170 days, respectively, with no evidence of tumor. Table 3.4 givesthe tumor-free time, the time from injection to the time that a tumor developsor to the end of the study. Fifteen of the 30 rats on the low-fat diet developeda tumor before the experiment was terminated. The rat that died had atumor-free time of at least 140 days. The other 14 rats did not develop any24 EXAMPLES OF SURVIVAL DATA ANALYSIS Table 3.4 Tumor-Free Time (Days) of 90 Rats on Three Different Diets Rat Low-Fat Rat Saturated Fat Rat Unsaturated Fat 1 140 1 124 1 112 2 177 2 58 2 683 50 3 56 3 844 65 4 68 4 109 5 86 5 79 5 153 6 153 6 89 6 143 7 181 7 107 7 608 191 8 86 8 709 77 9 142 9 98 10 84 10 110 10 16411 87 11 96 11 63 12 56 12 142 12 63 13 66 13 86 13 7714 73 14 75 14 9115 119 15 117 15 9116 140 /p59 16 98 16 66 17 200 /p59 17 105 17 70 18 200 /p59 18 126 18 77 19 200 /p59 19 43 19 63 20 200 /p59 20 46 20 66 21 200 /p59 21 81 21 66 22 200 /p59 22 133 22 94 23 200 /p59 23 165 23 101 24 200 /p59 24 170 /p59 24 105 25 200 /p59 25 200 /p59 25 108 26 200 /p59 26 200 /p59 26 112 27 200 /p59 27 200 /p59 27 115 28 200 /p59 28 200 /p59 28 126 29 200 /p59 29 200 /p59 29 161 30 200 /p59 30 200 /p59 30 178 Source: King et al. (1979 ). Data are used by permission of the author. tumor by the end of the experiment; their tumor-free times were at least 200 days. Among the 30 rats in the saturated fat diet group, 23 developed a tumor,one died tumor-free after 170 days, and six were tumor-free at the end of theexperiment. All 30 rats in the unsaturated fat diet group developed tumorswithin 200 days. The two early deaths can be considered losses to follow-up.The data are singly censored if the two early deaths are excluded. The investigator’s main interest here is to compare the three diets’ abilities to keep the rats tumor-free. To obtain information about the distribution ofthe tumor-free time, we can first estimate the survival (tumor-free )function of the three diet groups. The three survival functions were estimated using theEXAMPLE 3.1: COMPARISON OF TWO TREATMENTS AND THREE DIETS 25 Figure3.3 Survival curves of rats in three diet groups. Kaplan—Meier PL method and plotted in Figure 3.3. The median tumor-free times for the low-fat, saturated fat, and unsaturated fat groups were 188, 107,and 91 days, respectively. Since the three groups are homogeneous, we can skipthe step that checks for homogeneity and compare the three distributions oftumor-free time. TheK-sample test described in Section 5.3.3 can be used to test the significance of the differences among the three diet groups. Using this test, the investigator finds that the differences among the three groups are highlysignificant (p/p580.002 ). Note that the K-sample test can tell the investigator only that the differences among the groups are statistically significant. It cannottell which two groups contribute the most to the differences—whether thelow-fat diet produces a significantly different tumor-free time than thesaturated fat diet or whether the saturated fat diet is significantly different fromthe unsaturated fat diet. All one can conclude is that the data show a significantdifference among the tumor-free times produced by the three diets. 3.2 EXAMPLE 3.2: COMPARISON OF TWO SURVIVAL PATTERNS USING LIFE TABLES When the sample of patients is so large that their groupings are meaning- ful, the life-table technique can be used to estimate the survival distribution.A method developed by Mantel and Haenszel (1959 )and applied to life26 EXAMPLES OF SURVIVAL DATA ANALYSIS Table 3.5 Life Table for Male Patients with Localized Cancer of Rectum Diagnosed in Connecticut, 1935--1944 and 1945--1954 /p63 1935—1944 1945 —1954 Interval (t/p71) n/p30/p71d/p71w/p71/p59l/p71n/p71S/p19(t/p71)n/p30/p71d/p71w/p71/p59l/p71n/p71S/p19(t/p71) 1 388 167 2 387.0 0.5685 749 185 10 744.0 0.7513 2 219 45 1 218.5 0.4514 554 88 10 549.0 0.6309 3 173 45 1 172.5 0.3336 456 55 10 451.0 0.55394 127 19 0 127.0 0.2837 391 43 10 386.0 0.4922 5 108 17 0 108.0 0.2390 338 32 14 331.0 0.4446 6 91 11 1 90.5 0.2100 292 31 52 266.0 0.39287 79 8 0 79.0 0.1887 209 20 38 190.0 0.3514 8 71 5 0 71.0 0.1754 151 7 24 139.0 0.3337 9 66 6 1 65.5 0.1593 120 6 25 107.5 0.3151 10 59 7 0 59.0 0.1404 89 6 24 77.0 0.2905 Source: Myers (1969 ). /p63Symbols: n/p30/p71, number of patients alive at beginning of interval t/p71;d/p71, number of patients dying during interval t/p71;w/p71/p59l/p71, number of patients withdrawn alive or lost to follow-up during interval t/p71;n/p71/p58n/p30/p71/p57/p16/p17(w/p71/p59l/p71);S/p19(t/p71), cumulative proportion surviving from beginning of study to end of intervalt/p71.tables by Mantel (1966 )can be used to compare two survival patterns in the life-table analysis. Consider the data of male patients with localized cancer of the rectum diagnosed in Connecticut from 1935 to 1954 (Myers, 1969 ). A total of 388 patients were diagnosed between 1935 and 1944, and 749 patients werediagnosed between 1945 and 1954. For such large sample sizes the data can begrouped and tabulated as shown in Table 3.5. The 10 intervals indicate thenumber of years after diagnosis. For the tabulated life tables the survival functionS(t/p71) can be estimated for each interval t/p71. In Section 4.2 we discuss the estimation procedures of S(t/p71) and density and hazard functions. The survival, density, and hazard functions are the three most important functionsthat characterize a survival distribution. TheS/p19(t/p71) column in Table 3.5 gives the estimated survival function for the two time periods; these are plotted in Figure 3.4. Patients diagnosed in the1945—1954 period had considerably longer survival times (median 3.87 years ) than did patients diagnosed in the 1935 —1944 period (median 1.58 years ). The five-year survival rate is frequently used by cancer researchers and can easilybe determined from a life table. Patients diagnosed in 1935 —1944 had a five-year survival rate of 0.2390, or 23.9%. The patients diagnosed in 1945 — 1954 had a rate of 0.4446, or 44.5%. In comparing two sets of survival data,one can compare the proportions of patients surviving some stated period,such as five years, or the five-year survival rates. However, one cannotanticipate that two survival patterns will always stand in a superior —inferiorEXAMPLE 3.2: COMPARISON OF TWO SURVIVAL PATTERNS 27 Figure3.4 Survival curves for male patients with localized cancer of the rectum, diagnosed in Connecticut, 1935 —1944 versus 1945 —1954. relationship. It is more desirable to make a whole-pattern comparison (see Sections 4.3 and 5.2 ). The Mantel —Haenszel method described in Section 5.2 is a whole-pattern comparison and can be used to compare two survival patterns in life tables.Application of this method to the data in Table 3.5 results in a chi-square valueof 51.996 with 1 degree of freedom. We can conclude that the difference between the two survival patterns is highly significant (p/p580.001 ). Estimates of the survival function or survival rate depend on the life-table interval used. If each interval is very short, resulting in a large number ofintervals, the computation becomes very tedious and the life-table advantageis not fully taken. One assumption underlying the life table is that thepopulation has the same survival probability in each interval. If the intervallength is long, this assumption may be violated and the estimates inaccurate;this should be avoided except for rough calculations. Although the length ofeach interval and the total number of intervals are important, they will notcause trouble in most clinical studies since the study periods normally cover ashort period of time, such as one, two, or three years. Life tables with about10 to 20 intervals of several months to one year each are reasonable. Theinvestigator should also consider the disease under study. If the variation insurvival is large in a short period of time, the interval length should be short.However, in some demographicor other studies it is often of interest to c overa life span from birth to age 85 or more. The number of intervals would be28 EXAMPLES OF SURVIVAL DATA ANALYSIS very large if short intervals were used. In this case five-year intervals are sufficient to take into account the important variations in survival rateestimates (Shryock et al., 1971 ). 3.3 EXAMPLE 3.3: FITTING SURVIVAL DISTRIBUTIONS TO REMISSION DATA The remission times of 42 patients with acute leukemia were reported by Freireich et al. (1963 )in a clinical trial undertaken to assess the ability of 6-mercaptopurine (6-MP )to maintain remission. /p16Each patient was ran- domized to receive 6-MP or a placebo. The study was terminated after oneyear. The following remission times, in weeks, were recorded: 6-MP (21 patients ): 6, 6, 6, 7, 10, 13, 16, 22, 23, 6 /p59,9/p59,1 0/p59,1 1/p59,1 7/p59, 19/p59,2 0/p59,2 5/p59,3 2/p59,3 2/p59,3 4/p59,3 5/p59 Placebo (21 patients ): 1, 1, 2, 2, 3, 4, 4, 5, 5, 8, 8, 8, 8, 11, 11, 12, 12, 15, 17, 22, 23 Suppose that we are interested in a distribution to describe the remission times of these patients but that no information is available as to which distributionwill fit. We need to find a distribution that fits the data well. If we can find one,the remission experience can then be described by the properties of thedistribution, and the remission time of new patients can be predicted. Paramet-ric tests can be used to compare the effectiveness of the two treatments, butsince there are a large number of well-known functions and distributions tochoose from, the search becomes an art as much as a scientific task. The simplest and most efficient tool is the graph. Probability plotting can be done for complete data; for data that include censored observations, hazardplotting and the Cox —Snell method are more appropriate. It is not difficult to use the computer to generate these plots. Detailed discussions of probabilityplotting and hazard plotting are presented in Chapter 8. In both probabilityand hazard plotting, a linear configuration indicates that the distribution fitswell and its parameters can be estimated from the graph. Let us begin by trying to fit a distribution to the remission duration of 6-MP patients. Since the data consist of both censored and uncensored observations,we use the technique of hazard plotting. In this example we limit ourselves tothree distributions: the exponential, Weibull, and lognormal. In practice, moredistributions may need to be considered. Figures 3.5, 3.6, and 3.7 give thehazard plots for the exponential, Weibull, and lognormal distributions, respec-tively. A straight line is fitted to the points by eye in each of the plots. Amongthese graphs, the Weibull distribution appears to provide the best fit to theremission data. The straight line fits the points fairly closely. The estimates of /p16Data are used by permission of the publisher.EXAMPLE 3.3: FITTING SURVIVAL DISTRIBUTIONS 29 Figure3.5 Exponential hazard plot of the remission times of 21 leukemia patients who received 6-MP. Figure3.6 Weibull hazard plot of the remission times of 21 leukemia patients who received 6-MP.30 EXAMPLES OF SURVIVAL DATA ANALYSIS /afii9818/p92/p16/p431/p57exp[H(t)]/p44 Figure3.7 Log-normal hazard plot of the remission times of 21 leukemia patients who received 6-MP. the parameters /afii9838and/afii9828of the Weibull distribution obtained from the line are equal to 0.033 and 1.143, respectively (methods discussed in Chapter 8 ). After knowing that the Weibull distribution provides a good fit, we can use ananalytical method, the maximum likelihood method, to obtain a more accurateestimate of the parameters. Following the procedures discussed in Chapter 7, the maximum likelihood estimates of /afii9838and/afii9828are/afii9838/p19/p580.03 and /afii9828/p24/p581.354, which are quite close to the graphical estimates. After an appropriate distribution has been identified and parameters es- timated, we can estimate the probability of having a given duration ofremission and other probabilities. For example, the probability of having aremission time longer than 10 weeks can be predicted as P(T/p5710)/p58e/p92/p7/p16/p15 /afii9838/p19/p8/p65/p19/p58e/p57(10*0.03)/p16/p13/p18/p20/p19/p580.822 For the placebo group, we can use the probability plotting technique since the data are complete. Figures 3.8, 3.9, and 3.10 give the exponential, Weibull, andlognormal probability plots. Comparing the three graphs, again, the straightline in the Weibull plot appears to give the best fit. From the Weibull plot,estimates of /afii9838and/afii9828are found to be 0.111 and 1.250, respectively. The maximum likelihood estimates of /afii9838and/afii9828are, respectively, 0.105 and 1.371. Again, the graphical estimates are very close to the maximum likelihoodestimates. Based on the maximum likelihood estimates of the parameters, weEXAMPLE 3.3: FITTING SURVIVAL DISTRIBUTIONS 31 log/p431/[1/p57F(t)]/p44 Figure3.8 Exponential probability plot of the remission times of 21 leukemia patients who received placebo. can estimate the probability of having a remission time longer than 10 weeks. Using the same formula as given above, the probability for a patient receiving placebo to have a remission duration longer than 10 weeks is found to be 0.34,which is smaller than that of a patient receiving 6-MP. These graphical methods are subjective. The judgment as to whether the assumed distribution fits the data is based on a visual examination rather thanon an objective statistical test. However, the methods are very simple and doprovide a great deal of information. Even in a case where none of thedistributions discussed in this book fit well, graphs can help find the reasonsand thus help modify the model. Therefore, graphical methods are usuallyrecommended as the first thing to try. 3.4 EXAMPLE 3.4: RELATIVE MORTALITY AND IDENTIFICATION OF PROGNOSTIC FACTORS One thousand and twelve Oklahoma Indians (379 men and 633 women ) with non-insulin-dependent diabetes mellitus (NIDDM )were examined in32 EXAMPLES OF SURVIVAL DATA ANALYSIS log log /p431/[1/p57F(t)]/p44 Figure3.9 Weibull probability plot of the remission times of 21 leukemia patients who received placebo. /afii9818/p92/p16(F(t)) Figure3.10 Log-normal probability plot of the remission times of 21 leukemia patients who received placebo.EXAMPLE 3.4: RELATIVE MORTALITY 33 1972—1980 and a mortality follow-up study was conducted in 1986 —1989 (Lee et al., 1993 ). The mean [standard deviation (SD)] age and duration of diabetes at baseline examination were 52 (11)and 7 (6)years. The average duration of follow-up was 10 (SD 4 )years. As of December 31, 1989, 548 patients were alive, 452 (187 men and 265 women )were dead, and 12 could not be traced. Table 3.6 gives the survival time in years (T)of the first 40 male patients along with 12 potential prognosticfac tors: age, duration of diabetes (DUR )in years, family history of diabetes (FAM ), use of insulin within one year of diagnosis (INS), use of diuretics (DIU ), hypertension (HBP ), retinopathy (EVD ), pro- teinuria (PRO ), fasting plasma glucose (GLU )in milligrams per deciliter, cholesterol (TC)in milligrams per deciliter, triglyceride (TG)in milligrams per deciliter, and body mass index (BMI ), which is defined as weight in kilograms divided by height in meters squared. Among other things, the authors compared the mortality experience of the diabeticpatients with that of the general population in Oklahoma over thefollow-up period. Taking changes in age distribution into consideration, thepatients were divided into five groups according to their age at baselineexamination: /p5835, 35—44, 45—54, 55—64, and /p4665. The expected survival rates were calculated on a yearly basis following the methods described in Section4.3 and using the death rates given in the 1970 and 1980 Oklahoma populationlife tables. Death rates for the years between 1970 and 1980 and between 1980and 1989 were estimated based on the 1970 and 1980 statistics and theassumption that changes in death rates between 1970 and 1980 and after 1980follow a linear trend. The observed and expected survival curves for the groupswere plotted (Figure 3.11 ), and ratios of the observed and expected number of deaths (O/E ratios )by age were tabulated (Table 3.7 ). Figure 3.11 shows that the diabeticpatients had a muc h lower survivorship than the general Oklahoma population for this age —gender distribution. At the beginning of the fifteenth year after baseline examination, the relative survivalfor the diabeticOklahoma Indians was only 60%. The overall O/E ratios inTable 3.7 are 2.92 [or standardized mortality ratio (SMR )292] for men and 4.09 (or SMR 409 )for women, which indicates a significantly higher mortality rate in the diabeticOklahoma Indians than in the general population.Although patients in every group experienced excessive mortality, the youngerpatients had the highest rate. The relationship between the 12 potential prognosticvariables and the survival time of men was examined using univariate and multivariate methods.The procedures are summarized below. 1.Examine the individual relationship of each variable to survival. One way to analyze the data is first to determine which of the 12 variables could beconsidered of significant prognostic importance. In addition to correlationanalysis of these variables, the survival times in subcategories are compared(Table 3.8 ). Patients are grouped into subgroups in a meaningful way or in a way that maximizes the observed difference in survival time between the34 EXAMPLES OF SURVIVAL DATA ANALYSIS Figure3.11 Observed and expected survivorship from baseline examination for dabeticOklahoma Indians. subgroups (subject to the constraint that each subgroup contains at least 10% of the total number of patients ). The survivorship function for every subgroup of each variable was estimated using the Kaplan —Meier method (discussed in Chapter 4 )and plotted. Figure 3.12 gives an example. Survival functions among the subgroups were comparedby the logrank test (one of the available tests discussed in Chapter 5 ). Table 3.8 shows that except cholesterol and triglyceride, every one of the 12 variablesis significant. The median survival time decreases as age and duration ofdiabetes increase. Patients with a family history of diabetes, elevated fastingplasma glucose, hypertension, or retinopathy have significantly shorter survivaldurations than those without these characteristics. Patients with baseline BMIvalues greater than or equal to 30 had much better survivorship than didpatients with a lower BMI value.EXAMPLE 3.4: RELATIVE MORTALITY 35 Table3.6 Data of First 40 MalePatie nts Enrolle d in Mortality Study of Oklahoma Diabe tic Indians Patient Status /p63T Age DUR FAM /p64INS /p64DIU /p64HBP /p64EVD /p64PRO /p64GLU TC TG BMI 1 1 1 2 . 4 4 4 . 0 3 100001 2 4 2 3 9 2 5 3 8 3 4 . 2 2 1 1 4 . 1 4 3 . 5 1 0000009 4 1 5 8 9 4 4 2 . 2 3 0 1 4 . 4 4 7 . 6 4 101100 1 0 0 1 9 5 4 0 5 3 3 . 1 4 0 1 4 . 2 3 6 . 3 3 100000 1 7 1 2 1 2 2 1 8 3 8 . 55 0 1 4 . 4 5 4 . 4 7 000000 1 1 2 2 0 4 7 7 3 1 . 76 0 1 2 . 4 5 0 . 8 4 1010008 3 2 0 6 1 7 8 4 1 . 57 0 1 2 . 4 5 0 . 0 2 100000 1 0 4 1 7 8 1 0 0 3 9 . 58 17 . 0 6 6 . 9 8 000001 1 6 1 1 8 9 9 9 2 9 . 7 9 1 1 3 . 6 4 0 . 2 1 4 100110 2 6 2 2 4 1 3 0 1 2 7 . 5 1 0 0 1 4 . 4 5 4 . 1 4 101010 1 1 5 1 8 3 3 9 2 2 4 . 41 1 0 1 2 . 4 3 8 . 9 3 100101 1 0 8 2 3 7 4 9 3 2 . 41 2 19 . 8 5 1 . 2 4 100001 1 8 4 1 1 4 1 1 8 2 6 . 51 3 10 . 0 5 3 . 0 6 001001 1 2 6 2 0 6 4 8 0 3 4 . 51 4 1 1 2 . 1 4 5 . 0 5 001100 1 1 5 2 3 8 1 7 7 1 8 . 9 1 5 0 1 4 . 4 3 8 . 3 2 100000 1 1 0 2 0 4 1 8 0 3 1 . 2 1 6 0 1 2 . 4 4 0 . 0 5 101111 2 2 7 1 5 9 3 3 7 3 9 . 21 7 0 1 4 . 4 4 4 . 4 9 100000 1 8 2 1 9 3 3 3 2 3 2 . 71 8 0 1 4 . 2 4 8 . 2 5 101111 1 8 4 1 7 1 2 3 8 3 3 . 51 9 0 1 2 . 4 3 6 . 3 6 100000 2 3 8 1 9 6 1 6 2 2 4 . 22 0 0 1 3 . 7 4 1 . 9 3 100010 2 8 4 1 7 0 1 2 5 3 0 . 7 36 2 1 1 1 3 . 4 4 9 . 7 1 5 000000 2 4 8 2 1 7 2 3 4 2 8 . 0 2 2 1 1 2 . 6 5 1 . 5 9 1011109 6 1 7 4 1 4 4 2 4 . 22 3 0 1 4 . 0 4 4 . 1 1 100000 1 1 6 2 5 1 1 5 3 3 3 . 3 2 4 0 1 2 . 4 3 5 . 9 3 100000 1 2 0 1 2 9 7 6 3 0 . 1 2 5 0 1 2 . 9 5 0 . 4 1 000100 1 2 8 1 9 0 1 2 3 2 7 . 72 6 0 1 2 . 4 4 8 . 0 1 0011009 5 2 0 7 5 9 2 8 . 12 7 0 1 4 . 5 4 0 . 1 3 0000008 9 2 5 2 1 1 2 3 1 . 72 8 0 1 3 . 4 5 4 . 5 0 010100 1 0 4 4 9 0 5 4 0 3 0 . 82 9 1 1 3 . 9 6 9 . 3 6 101100 1 2 2 1 6 2 2 0 9 2 4 . 2 3 0 1 1 2 . 0 3 8 . 5 0 101101 2 2 5 1 8 3 3 4 3 4 3 . 1 3 1 1 3 . 6 6 3 . 7 1 8 000110 2 1 1 2 1 7 1 2 4 2 5 . 13 2 1 1 5 . 4 7 1 . 8 1 2 100000 1 5 0 2 2 7 1 3 7 2 6 . 03 3 0 1 0 . 3 5 9 . 5 3 101100 1 8 0 1 8 8 3 0 8 2 8 . 13 4 1 5 . 8 5 0 . 1 1 100000 4 0 0 2 0 0 1 6 6 2 6 . 13 5 1 2 . 5 7 5 . 4 1 8 100101 2 8 1 6 9 2 3 6 4 4 9 . 7 3 6 0 1 5 . 0 2 9 . 6 1 100001 1 6 7 1 5 4 1 5 7 6 0 . 2 3 7 1 5 . 5 6 0 . 2 1 8 000010 3 5 1 2 0 6 1 4 1 2 6 . 03 8 1 4 . 5 6 3 . 9 4 100101 1 2 7 1 8 0 1 3 1 2 1 . 8 3 9 1 6 . 8 5 7 . 4 1 7 100110 1 5 3 7 8 6 4 6 6 3 4 . 1 4 0 1 3 . 6 7 1 . 3 1 8 111111 1 7 9 4 8 8 3 6 4 2 5 . 6 /p631, dead; 0, alive. /p641, yes; 0, no. 37 Figure3.12 Survival curves of diabetic patients by hypertension status at baseline.Table 3.7 Observed and Expected Number of Deaths and O /E Ratios During Follow-up Period by Gender and Age at Baseline Age at Males FemalesBaseline (yr) Observed Expected O/E Observed Expected O/E /p4544 34 6.48 5.25 31 5.51 5.63 45—54 79 25.64 3.08 85 18.40 4.62 55—64 32 11.67 2.74 69 15.03 4.59 65/p59 42 20.32 2.06 80 25.86 3.09 Total 187 64.11 2.92 265 64.80 4.09 2.Examine the simultaneous relationship of the variables to survival. Exam- ination of each variable can give only a preliminary idea of which variablesmight be of prognostic importance. The simultaneous effect of the variablesmust be analyzed by an appropriate multivariate statistical method to deter-mine the relative importance of each. Cox’s (1972 )proportional hazards model38 EXAMPLES OF SURVIVAL DATA ANALYSIS Table 3.8 Survival Time by Potential Prognostic Variable for Male Diabetic Patients Number of Number of Median Variable Patients Deaths Survival Time (yr)pValue Age (yr) /p5845 102 34 15.2 45—54 173 79 14.2/p580.00155—64 53 31 9.1 /p4665 46 43 5.1 Family history of diabetes No 104 62 11.4/p580.01Yes 238 104 14.8 Duration of diabetes (yr) /p587 207 84 15.2 7—13 91 49 12.2 /p580.001 /p4614 58 48 7.9 Use of diuretics No 254 117 14.5/p580.001Yes 102 64 10.0 Use of insulin /p581 year of diagnosis No 317 157 13.9/p580.05Yes 39 24 11.9 Hypertension No 211 84 15.3/p580.001Yes 151 99 9.8 Retinopathy No 332 163 13.9/p580.001Yes 24 18 6.5 Proteinuria Negative 250 112 14.5Slight 54 27 12.4 /p580.001 Heavy 57 44 8.0 Fasting plasma glucose /p58200 235 106 11.9/p580.05/p46200 139 81 8.4 Cholesterol /p58240 300 144 14.80.12/p46240 64 38 12.2 Triglyceride /p58220 223 105 14.10.44/p46220 141 77 13.2 BMI /p5830 189 114 11.8/p580.001/p4630 184 73 15.4EXAMPLE 3.4: RELATIVE MORTALITY 39 can be applied. This model, presented in Chapter 12, is a regression model that relates patient characteristics directly to the risk of failure and thusindirectly to survival. The assumption of this model is that the hazards fordifferent strata of each independent (or prognostic )variable are proportional over time. This assumption was verified by a graphical method (discussed in Chapter 12 )using each of the variables. Figure 3.13 gives an example of the graph of log[ /p57S(t)] versustfor the two hypertension groups. The two almost parallel curves indicate that the hazards of dying are proportional. Therefore, the assumption of the proportional hazards model is satisfied and the modelappropriate. This model can be fitted by a stepwise procedure that results in a ranking of the prognosticvariables. The first variable selec ted to enter the model is themost important single variable in predicting the risk of dying. The secondvariable is the second most important, and so on. A significance level can beobtained from a likelihood ratio test at each step, which indicates the level ofcontribution given by the additional variable. Using the proportional hazards model and a stepwise procedure, seven of the 12 variables were identified as significant at the 0.05 level based on thelikelihood ratio test at each step. These variables, the regression coefficients,and the significance levels based on the Ward test, which uses the regressioncoefficient and its standard error (S.E.)are given in Table 3.9. The sign of the coefficient indicates whether the variable is positively or negatively related tothe hazard of dying. For example, age and duration of diabetes are bothpositively related to the risk of dying and therefore negatively related to thesurvival time. Table 3.9 also gives the ratio of risk (or hazard )for values of each variable unfavorable to survival to values of that variable favorable tosurvival. For example, patients who were 60 years of age at baseline had a 3.05times higher risk of dying during the follow-up period (10—16 years, average 13 years )than did patients who were only 40 years old at baseline. For dichotomous variables, the ratio of risk is equal to exp (coefficient ), which is also interpreted as the relative risk of the variable adjusting for the othervariables. Consequently, the confidence interval for the relative risk can becalculated (not shown in Table 3.9 ). Based on this set of data, the authors conclude that age, hypertension, duration of diabetes, fasting plasma glucose,BMI, proteinuria, and use of diuretics are significantly related to survival. Themultivariate method also showed that high values of BMI might be protective. 3.5 EXAMPLE 3.5: IDENTIFICATION OF RISK FACTORS A study of the incidence of retinopathy in Oklahoma Indians with NIDDM was conducted in 1987 —1990 as part of a prospective study of diabetic complications (Lee et al., 1992 ). Among the 312 patients who were free of retinopathy at initial examination in the 1970s, 228 were found to have40 EXAMPLES OF SURVIVAL DATA ANALYSIS Figure3.13 Curves of log[ /p57logS(t)] for the two hypertension groups. Table 3.9 Significant Variables (at 0.05 Level) Identified by Proportional Hazards Model Relative Risk /p64 Ratio Regression pValue of Variable /p63 Coefficient (Ward Test )Favorable Unfavorable Risk Age 0.0558 /p580.001 9.32 28.45 3.05 Hypertension: 0.6360 /p580.001 1.00 1.89 1.89 1, yes, 0, no Duration of 0.0559 /p580.001 1.32 2.19 1.66 diabetes Fasting plasma 0.0023 /p580.010 1.35 1.58 1.17 glucose BMI /p570.0330 0.035 0.32 0.44 0.72 Proteinuria: 0.3744 0.025 1.00 1.45 1.45 1, yes, 0, no Use of diuretics: 0.4191 0.030 1.00 1.52 1.52 1, yes; 0, no /p63Variables are listed in order of entry into model with a p-value limit for entry of 0.05. /p64Favorable categories are 40 years of age, no hypertension, duration of diabetes 5 years, fasting plasma glucose 130mg/dL, BMI 35, no proteinuria, and no diuretics use. Unfavorable categoriesare 60 years of age, hypertensive, duration of diabetes 14 years, fasting plasma glucose 200mg/dL, BMI 25, having proteinuria, and diuretics use.EXAMPLE 3.5: IDENTIFICATION OF RISK FACTORS 41 developed the eye disease during the 10 to 16-year follow-up period (average follow-up time 12.7 years ). Twelve potential factors (assessed at time of baseline examination )were examined by univariate and multivariate methods for their relationship to retinopathy (RET ): age, gender, duration of diabetes (DUR ), fasting plasma glucose (GLU ), initial treatment (TRT ), systolic (SBP )and diastolicblood pressure (DBP ), body mass index (BMI ), plasma cholesterol (TC), plasma triglyceride (TG), and presence of macrovascular disease (LVD ) or renal disease (RD). Table 3.10 gives the data for the first 40 patients. Among other things, the authors related these variables to the development ofretinopathy. 1.Examine the individual relationship of each variable to the development of diabetic retinopathy. Table 3.11 gives some summary statistics of the eight continuous variables for patients who have developed retinopathy and forthose who have not. Notice that patients who have developed the disease wereyounger at baseline and had much higher fasting plasma glucose, systolic anddiastolicblood pressure, and plasma triglyc eride than did patients who havenot. Table 3.12 summarizes the contingency table analysis of retinopathyincidence rates. The number of patients at risk of developing retinopathy andthe number of patients who developed the disease (and rate )are given by subcategory of each potential risk factor. Using the chi-square test, it is foundthat there was a significant difference in the retinopathy rate among thesubcategories of several variables using a significance level of 0.05: duration ofdiabetes, fasting plasma glucose, systolic and diastolic blood pressure, andtreatment. It appears that patients with poor glucose control or high bloodpressure or treated with oral agents or insulin have a higher incidence ofretinopathy. In addition, patients with high triglyceride levels tend to havehigher incidence of retinopathy (p/p580.064). However, patients who had developed macrovascular disease at the time of baseline examination had alower retinopathy incidence. The authors state that this may be due to the factthat 68% of the patients who had macrovascular disease either died (54% ) during the follow-up period or were lost to follow-up (14% ). Many of these patients may have developed retinopathy, particularly the patients who havedied, but were not included. Therefore, the lower incidence of retinopathy inpatients who had macrovascular disease at baseline is probably the result of aselection bias. Similarly, the large number of death plus the losses to follow-upmay also contribute to the drop in retinopathy rate in patients who had haddiabetes for more than 12 years at baseline. Among the 80 patients in thisduration of diabetes category, 56% have died and 10% did not participate inthe follow-up examination. The large number of deaths may also be responsiblefor the finding that patients who survived long enough to develop retinopathywere younger at baseline. The deceased patients were significantly older (mean 57 years )than the survivors who participated in the follow-up examination (mean 48 years ).42 EXAMPLES OF SURVIVAL DATA ANALYSIS Table 3.10 First 40 Patients Involved in Study of Risk Factors in Development of Diabetic Retinopathy Patient RET /p63Age Gender DUR TRT /p64GLU SBP DBP TC TG BMI LVD /p63RD /p63 1 1 47.6 M 4 2 100 156 98 195 405 33.1 0 0 2 0 54.4 M 7 1 112 112 78 204 77 31.7 0 03 0 50.8 M 4 2 83 134 80 206 178 41.5 1 04 1 49.4 F 5 2 276 102 70 190 222 24.1 0 05 1 50.0 M 2 2 104 142 86 178 100 39.5 0 06 1 50.7 F 7 2 242 142 78 217 268 31.6 0 1 7 1 35.3 F 2 2 130 134 80 390 564 47.0 0 0 8 0 50.2 F 5 2 130 100 70 174 128 29.8 0 09 1 45.0 M 5 2 115 134 100 238 177 18.9 1 0 10 0 38.3 M 2 2 110 132 80 204 180 31.2 1 011 0 45.8 F 5 3 130 118 68 185 316 26.6 1 012 1 51.6 F 3 2 141 112 78 152 77 31.2 0 1 13 1 36.3 M 6 2 238 142 94 194 162 24.2 0 0 14 1 44.7 F 16 2 190 152 90 132 161 32.0 0 015 1 37.2 F 1 2 126 136 90 133 211 32.7 0 016 1 52.6 F 9 2 159 140 76 151 132 26.7 0 017 1 44.1 M 1 1 116 126 76 251 153 33.3 0 018 1 35.9 M 3 1 120 132 80 129 76 30.1 0 0 19 1 50.4 M 1 2 128 144 90 190 123 27.7 0 0 20 1 48.0 M 1 2 95 128 74 207 59 28.1 1 021 0 47.5 F 1 2 85 124 82 161 190 31.6 1 022 1 50.1 F 6 2 138 106 72 181 135 30.5 0 1 (Continued overleaf ) 43 Table3.10 Continued Patient RET /p63Age Gender DUR TRT /p64GLU SBP DBP TC TG BMI LVD /p63RD /p63 23 0 43.3 F 1 1 104 128 86 204 198 26.1 0 0 24 1 54.5 M 0 3 104 142 84 490 540 30.8 1 025 1 52.2 F 3 3 304 132 84 192 119 36.9 1 026 1 53.3 F 3 2 249 128 72 120 85 35.5 0 027 1 64.3 F 6 2 297 138 80 145 64 30.1 1 028 1 44.6 F 4 1 139 112 80 156 111 36.1 0 1 29 1 47.1 F 2 2 169 130 84 198 99 39.0 0 1 30 1 46.5 F 7 1 159 128 78 238 157 34.5 0 031 0 51.5 F 3 1 147 128 78 185 182 32.0 0 032 1 59.5 M 3 2 180 132 78 188 308 28.1 1 033 1 52.0 F 6 2 183 142 84 175 68 37.6 0 034 1 45.7 F 6 2 180 138 80 179 189 44.1 1 0 35 1 48.2 F 18 2 267 158 100 195 112 25.2 1 0 36 1 57.4 M 0 1 159 172 108 219 294 33.7 0 137 1 42.0 F 4 2 158 106 68 224 157 33.6 0 038 1 50.7 F 1 2 211 142 84 390 645 37.1 0 039 1 53.8 F 0 1 177 154 80 175 208 35.3 0 140 0 56.9 M 1 1 98 116 70 146 97 26.5 0 1 /p631, yes; 0, no. /p641, diet only; 2, oral agent; 3, insulin. 44 Table 3.11 Summary Statistics for Eight Variables by Retinopathy Status at Follow-up Retinopathy Status No Yes Variable Mean S.D. Mean S.D. pValue Age 50.0 9.0 47.2 7.4 0.01 Duration of diabetes 4.2 4.5 4.8 4.4 0.34 Fasting plasma glucose 141.8 65.6 196.3 76.6 /p580.0001 Systolicblood pressure 128.0 15.7 132.6 17.3 0.04Diastolicblood pressure 80.3 10.8 84.9 10.1 /p580.001 Body mass index 32.3 6.3 32.5 5.9 0.76Cholesterol 204.4 66.0 206.8 58.7 0.76 Triglyceride 180.5 111.1 234.4 273.3 0.01 2.Examine the simultaneous relationship of the variables to the development of retinopathy. Univariate analysis of each variable using the contingency table or the chi-square test gives a preliminary idea of which individual variablemight be of prognostic importance. The simultaneous effect of all the variablescan be analyzed by the linear logistic regression model (discussed in Section 14.2)to determine the relative importance of each. The 12 variables were fitted to the linear logisticregression model using a stepwise selection procedure. The variables most significantly related to thedevelopment of retinopathy were found to be initial treatment, fasting plasmaglucose, age, and diastolic blood pressure (p/p450.001). Table 3.13 gives the regression coefficients of the four most significant variables (p/p450.05), the standard errors, and adjusted odds ratios [exp (coefficient )]. Thepvalues used here are the significance levels based on the likelihood ratio test or theimprovement in the maximum likelihood due to the addition of the variable inthe stepwise procedure. This method is more powerful than the Wald test,which is based on the standardized regression coefficients (Chapter 14 ). The results are consistent with those in the univariate analysis. On the basis of the regression coefficients, the probability of developing retinopathy during a 10 to 16-year follow-up can be estimated by substitutingvalues of the risk factors into the regression equation, logP 1/p57P /p58/p572.373/p591.495(oral agent )/p590.882 (insulin ) /p590.014 (GLU )/p570.074 (age)/p590.048 (DBP ) For example, for a 50-year-old patient who is on oral agents and whose fasting plasma glucose and diastolic blood pressure are 170 mg/dl and 95 mmHg,EXAMPLE 3.5: IDENTIFICATION OF RISK FACTORS 45 Table 3.12 Cumulative Incidence Rates of Retinopathy by Baseline Variables Developed Retinopathy Number of Variable Persons at Risk Number Percent pvalue Gender Female 211 151 71.60.384Male 101 77 76.2 Age (yr) /p5835 13 10 76.9 35—44 101 77 76.20.24245—54 155 115 74.2 /p4655 43 26 60.5 Duration of diabetes (yr) /p584 153 105 68.6 4—7 113 86 76.10.0338—11 23 22 95.7 /p4612 23 15 65.2 Fasting plama glucose (mg/dl ) /p58140 117 62 53.0 140—199 90 74 82.2 /p580.001 /p46200 105 92 87.6 Systolicblood pressure (mmHg ) /p58130 145 95 65.5 130—159 149 115 78.8 0.016 /p46160 20 18 85.7 Diastolicblood pressure (mmHg ) /p5885 179 118 65.9 85—94 87 73 83.9 0.004 /p4695 46 37 80.4 Plasma cholesterol (mg/dl ) /p58240 267 193 72.30.442/p46240 45 35 77.8 Plasma triglyceride (mg/dl ) /p58250 237 167 70.50.064/p46250 75 61 81.3 Body mass index (kg/m /p17) /p5828 73 49 67.1 28—33 121 94 77.7 0.261 /p4634 118 85 72.0 Renal disease No 251 179 71.30.155Yes 61 49 80.3 Macrovascular disease No 205 157 76.60.053Yes 107 71 66.4 Treatment (initial ) Diet alone 115 62 53.9 Oral agent 158 136 86.1 /p580.001 Insulin 37 29 78.446 EXAMPLES OF SURVIVAL DATA ANALYSIS Table 3.13 Results of Logistic Regression Analysis Standard Variable Coefficient Error exp (coefficient )Coefficient/S.E. Constant /p572.373 1.557 Initial treatment Oral agent 1.495 0.330 4.459 4.53 Insulin 0.882 0.488 2.416 1.81 Fasting plasma 0.014 0003 1.014 4.67 glucose Age /p570.074 0019 0.929 /p573.89 Diastolicblood 0.048 0.015 1.049 3.20 pressure respectively, the chance of developing retinopathy in the next 10 to 16 years is 91%. The linear logisticregression model is useful in identifying important risk factors. However, complete measurements of all the variables are needed;missing data are a problem. In this example, complete data are available onmost of the patients. This may not always be the case. Although there aremethods of coping with missing data (discussed in Section 11.1 ), none is perfect. Thus it is extremely important for investigators to make every effort toobtain complete data on every subject. Bibliographical Remarks It is impossible to cite all the published examples of survival data analysis similar to those in this chapter. Other similar studies can be found in theliterature: for example, Biometrics ,Biometrika ,Cancer,Journal of Chronic Disease,Journal of the National Cancer Institute ,American Journal of Epi- demiology ,Journal of the American Medical Association , andNew England Journal of Medicine . An easy way to find examples is to use the National Library of Medicine’s Web site and search the file PubMed with appropriatekeywords. EXERCISES The four sets of data below are taken from actual research situations. Although the data can be used for various analyses throughout the book, the reader isasked here only to describe in detail how the data can be analyzed. The dataappear in examples and other exercises in subsequent chapters. 3.1Thirty-three patients with hypernephroma were treated with combined chemotherapy,immunotherapy,andhormonaltherapy.ExerciseTable3.1EXERCISES 47 Exercise Table 3.1 Data for 33 Patients with Hypernephroma Date of Date Death or Skin Test Results /p65 Treatment Last Patient Age Gender Started Response /p63Follow-up Status /p64Monilia Mumps PPD PHA SK-SD 1 53 F 3/31/77 1 10/1/77 0 7 /p5972 3 /p5923 0 /p5902 5 /p5925 0 /p590 2 61 M 6/18/76 0 8/21/76 1 10 /p5910 15 /p5920 0 /p5901 3 /p5913 9 /p599 3 53 F 2/1/77 3 10/1/77 0 0 /p5907 /p5970 /p5902 5 /p5925 0 /p590 4 48 M 12/19/74 2 1/15/76 1 0 /p5900 /p5900 /p5900 /p5900 /p590 5 55 M 11/10/75 0 1/15/76 1 12 /p5912 ND 10 /p5910 8 /p5985 /p595 6 62 F 10/7/74 2 4/5/75 1 10 /p5910 5 /p5950 /p5907 /p5975 /p595 7 57 M 10/28/74 0 1/6/75 1 15 /p5915 15 /p5915 0 /p5900 /p5901 0 /p5910 8 53 M 10/6/75 2 6/18/77 1 0 /p590N D 0 /p5901 2 /p5912 0 /p590 9 45 M 4/11/77 0 10/1/77 0 6 /p5944 /p5940 /p5900 /p5900 /p590 10 58 M 8/4/76 3 2/11/77 1 13 /p5913 13 /p5913 22 /p5922 23 /p5923 0 /p590 11 61 F 1/1/77 3 10/1/77 0 0 /p5908 /p5981 7 /p5917 11 /p5911 0 /p590 12 61 M 7/25/76 1 10/1/77 0 9 /p5991 2 /p5912 0 /p5902 0 /p5920 0 /p590 13 77 M 5/8/75 0 9/26/75 1 0 /p5900 /p5900 /p5900 /p5900 /p590 14 55 M 4/27/77 2 10/1/77 0 0 /p5900 /p5901 5 /p5915 10 /p5910 0 /p590 15 50 M 4/20/77 3 10/1/77 0 0 /p5901 4 /p5914 5 /p5953 2 /p5932 21 /p5921 16 42 M 8/24/76 0 10/1/77 0 11 /p5911 7 /p5970 /p5901 2 /p5912 0 /p590 48 17 50 F 1/8/75 0 6/30/75 1 0 /p5900 /p5900 /p5900 /p5900 /p590 18 66 F 9/8/76 3 10/1/77 0 9 /p5991 0 /p5910 6 /p5961 5 /p5915 11 /p5911 19 58 M 2/18/75 0 10/1/77 0 0 /p5900 /p5900 /p5900 /p590N D 20 62 M 5/12/76 0 10/17/76 1 2 /p592N D N D 3 /p5932 /p592 21 71 F 10/22/76 3 12/12/76 1 10 /p5910 6 /p5960 /p5901 2 /p5912 0 /p590 22 44 M 6/6/77 3 10/1/77 0 10 /p5910 10 /p5910 0 /p5902 0 /p5920 0 /p590 23 69 M 6/21/76 0 10/13/76 1 0 /p5901 5 /p5915 25 /p5925 25 /p5925 0 /p590 24 56 M 6/7/77 2 10/1/77 0 0 /p5907 /p5970 /p5900 /p5900 /p590 25 57 M 11/16/76 0 12/10/76 1 11 /p5911 5 /p5950 /p5902 0 /p5920 0 /p590 26 69 M 5/10/77 0 7/25/77 1 0 /p5900 /p5900 /p5901 5 /p5915 0 /p590 27 60 M 6/29/77 0 7/7/77 1 0 /p5900 /p5900 /p5902 6 /p5926 0 /p590 28 60 M 7/21/75 3 10/1/77 0 11 /p5911 20 /p5920 10 /p5910 18 /p5918 0 /p590 29 72 M 7/19/75 0 10/18/75 1 10 /p5910 0 /p5907 /p5971 0 /p5910 0 /p590 30 42 F 3/3/75 0 4/23/75 1 0 /p590N D 0 /p5900 /p5900 /p590 31 57 M 2/24/77 2 10/1/77 0 5 /p5958 /p5980 /p5902 5 /p5915 0 /p590 32 66 M 6/15/77 3 10/1/77 0 0 /p5901 5 /p5915 0 /p5901 0 /p5910 0 /p590 33 59 M 3/4/77 0 4/2/77 1 0 /p5900 /p5900 /p5901 6 /p5916 0 /p590 Source: Data courtesy of Richard Ishmael. /p630, no response; 1, complete response; 2, partial response; 3, stable. /p640, alive; 1, dead. /p65ND, not done. 49 gives the age, gender, date treatment began, response status, date of death or last follow-up, survival status, and results of five pretreatment skintests. The investigator is interested in the response and survival of thepatients and in identifying prognosticfac tors. How would you analyze thedata? 3.2In a study undertaken to compare the treatments given to hyperneph- roma patients and to relate response and survival to surgery, metastasis, and treatment time, data from 58 patients were collected (Exercise Table 3.2). How would you analyze the data to answer these questions? (a)Do patients who had nephrectomy have a higher response rate? (b)Is the time of nephrectomy related to response and survival? (c)Are there significant differences between the treatments? (d)What are the most important variables related to response and survival? 3.3Exercise Table 3.3 gives the age, gender, family history of melanoma, remission duration, survival time, stage, and results of six pretreatmentskin tests (the larger diameter is given )of 102 stage 3 and 4 melanoma patients (Lee et al., 1982 ). (a)Study the immunocompetence of melanoma patients by investigating skin test results. (b)Determine if age, gender, or pretreatment skin test results are predic- tive to remission and survival time. (c)Find theoretical distributions that describe the survival and remission patterns. 3.4One hundred and forty-nine diabeticpatients were followed for 17 years (a subset of data from Lee et al., 1988 ). Exercise Table 3.4 gives the survival time from baseline examination, survival status, and severalpotential prognosticfac tors at baseline: age, body mass index (BMI ), age at diagnosis of diabetes, smoking status, systolicblood pressure (SBP ), diastolicblood pressure (DBP ), electrocardiogram reading (ECG ), and whether the patient had any coronary heart disease (CHD ). Identify the important prognostic factors that are associated with survival.50 EXAMPLES OF SURVIVAL DATA ANALYSIS Exercise Table 3.2 Data of 58 Patients with Hypernephroma Time of Survival Lung Bone Patient Gender Age Nephrectomy /p64Nephrectomy /p65Treatment /p66Response /p65Time Status /p68Metastasis /p64Metastasis /p64 12 5 3 1 0 . 0 1 17 7 01 0 21 6 9 1 4 . 0 1 21 8 10 1 31 6 1 0 /p579.0 1 0 8 1 1 0 42 5 2 1 2 . 0 1 26 8 11 0 51 4 6 1 2 . 0 1 23 5 10 161 5 5 1 0 . 0 1 0 8 11 072 6 2 1 0 . 0 1 22 6 11 081 5 3 1 0 . 0 1 28 4 10 191 7 0 0 /p579.0 1 0 17 1 1 0 10 1 48 1 0.0 1 3 52 1 1 0 11 1 58 1 1.5 1 0 26 1 1 112 1 61 1 5.0 1 1 108 0 1 013 1 77 1 4.0 1 0 18 1 1 014 1 56 1 0.0 1 3 72 1 0 115 1 55 1 0.0 1 2 38 1 1 1 16 1 50 1 4.0 1 3 /p5799 9 1 0 17 1 75 1 0.0 1 0 9 1 1 018 1 43 1 2.0 1 3 56 1 0 019 1 69 1 1.0 1 2 36 1 1 120 2 59 1 1.5 1 2 108 1 0 121 2 71 1 0.0 1 2 10 1 1 0 22 1 56 1 0.0 1 2 36 1 1 0 23 1 57 0 /p579.0 1 0 6 1 1 0 24 1 69 1 8.0 1 0 9 1 1 0 (Continued overleaf ) 51 Exercise Table 3.2 Continued Time of Survival Lung Bone Patient Gender Age Nephrectomy /p64Nephrectomy /p65Treatment /p66Response /p65Time Status /p68Metastasis /p64Metastasis /p64 25 1 72 0 /p579.0 1 0 12 1 1 0 26 1 67 1 0.0 1 9 5 0 1 027 1 41 1 2.0 1 2 104 0 1 0 28 1 77 1 10.0 1 0 6 1 1 0 29 2 63 1 2.0 1 3 115 1 0 130 2 42 1 12.0 1 0 9 1 0 0 31 1 59 0 /p579.0 1 0 21 1 0 0 32 1 62 1 5.0 1 0 14 1 1 033 1 65 1 0.0 1 0 52 1 1 0 34 2 53 0 /p579.0 1 0 9 1 0 1 35 1 57 1 0.0 1 2 48 1 1 036 2 60 0 /p579.0 1 0 15 1 1 1 37 1 59 1 0.0 1 0 5 1 1 038 1 75 1 0.0 2 3 28 0 1 039 2 53 1 0.0 2 2 25 0 1 0 40 2 67 1 5.0 2 3 25 0 0 1 41 1 58 1 8.0 2 3 40 1 1 142 1 62 1 8.0 2 0 16 1 1 143 1 69 0 /p579.0 2 0 8 1 1 1 52 44 1 44 1 0.0 2 2 70 0 1 0 45 1 60 1 1.0 2 0 6 1 0 146 1 57 0 /p579.0 2 0 8 1 1 1 47 1 45 1 2.0 2 4 12 1 1 0 48 2 50 1 1.0 2 4 20 1 1 049 1 58 0 /p579.0 2 4 8 1 1 0 50 1 51 1 0.0 2 3 /p5799 0 0 1 51 1 59 1 3.0 2 3 12 1 1 052 1 53 1 0.0 2 1 181 0 1 0 53 1 70 1 0.5 2 0 20 1 1 1 54 1 69 1 3.0 2 0 14 1 1 055 1 62 1 0.0 2 3 26 1 0 056 1 52 1 2.0 2 0 16 1 1 057 1 77 1 2.0 2 2 30 1 1 058 1 61 1 8.0 2 0 20 1 1 0 Source: Data courtesy of Richard Ishmael. /p631, male; 2, female. /p641, yes; 0, no. /p65Number of years prior to treatment; negative value—no nephrectomy. /p661, combined chemotherapy and immunotherapy, 2, others. /p670, no response; 1, complete response; 2, partial response; 3, stable; 4, increasing disease; 9, unknown. /p681, dead; 0, alive; 9, unknown. 53 Exercise Table 3.3 Data of 102 Patients with Stages 3 and 4 Melanoma Family Remission Survival Skin Tests /p68 History of Time Remission Time Survival Patient Age Gender Melanoma /p64 (months )/p65Status /p66 (months )Status /p67Stage Monilia Mumps PPD PHA SK-SD Tricophyton 1 58 2 9 42.0 0 42.0 0 3B 18 16 0 20 14 99 25 0 2 0 3 . 3 13 . 9 1 3 B 7 8 0 0 70 37 6 1 0 6 . 1 1 1 0 . 5 1 3 B 0 1 5 0 0 9 90 46 6 2 9 2 . 3 16 . 0 0 3 B 8 0 0 1 0 005 33 1 9 5.1 1 20.6 1 3B 17 10 0 5 18 99 6 55 2 0 11.1 0 21.8 0 4B 99 99 99 99 99 99 7 25 2 0 36.5 0 36.5 0 3B 7 5 0 8 90 998 23 1 0 24.3 0 24.3 0 3B 10 20 0 7 30 0 9 30 1 0 28.7 0 28.7 0 3B 8 99 0 6 10 17 10 34 1 9 7.7 0 7.7 0 3B 15 99 0 0 0 011 34 1 0 29.3 0 29.3 0 3B 15 99 0 7 10 0 12 26 2 0 5.9 1 19.3 1 3AB 0 99 0 42 25 0 13 27 1 9 2.6 1 6.9 1 3AB 15 99 0 10 10 0 14 72 2 9 16.7 0 18.0 0 4B 10 99 99 15 0 7 15 70 2 0 14.6 0 14.6 0 3B 0 99 0 16 0 20 16 82 2 0 /p5799.0 9 23.6 0 4B 9 99 99 17 10 0 17 43 1 9 /p5799.0 9 3.9 1 4B 25 0 0 7 20 5 18 52 1 9 /p5799.0 9 7.3 1 4B 12 99 0 5 15 0 19 34 1 9 /p5799.0 9 9.8 1 4B 13 20 0 15 30 30 20 48 1 0 26.5 0 26.5 0 4A 10 99 0 88 5 0 21 62 1 9 18.0 1 25.4 0 3AB 5 99 0 10 0 24 22 49 1 9 4.3 1 8.0 1 3B 0 99 8 3 0 023 46 1 0 0.3 1 13.8 1 3B 0 5 0 7 0 0 24 53 2 1 21.5 0 21.5 0 3A 16 20 14 14 18 12 25 21 2 9 /p5799.0 9 9.3 1 4B 10 5 0 10 15 0 26 25 1 9 /p5799.0 9 1.2 1 4B 15 10 5 18 10 5 27 35 2 0 /p5799.0 9 20.0 1 4B 11 7 0 10 0 2 54 28 66 2 9 /p5799.0 9 12.5 0 4B 99 99 99 99 99 99 29 54 2 0 /p5799.0 9 7.4 1 4B 99 99 99 99 99 99 30 43 2 0 /p5799.0 9 4.7 0 4B 13 19 9 11 30 0 31 40 1 0 13.3 0 13.3 0 3B 0 5 25 12 10 032 16 1 0 0.0 0 0.0 0 3B 7 10 14 10 0 0 33 59 1 0 /p5799.0 9 25.8 0 4B 0 99 0 18 0 35 34 64 1 9 16.5 0 16.5 0 3B 8 7 0 0 0 035 52 1 9 /p5799.0 9 2.5 1 4B 9 14 0 12 5 0 36 /p5799 2 8 /p5799.0 9 13.8 1 4B 30 75 0 35 0 12 37 27 1 0 /p5799.0 9 4.2 1 4B 10 12 10 10 0 3 38 60 2 0 5.4 1 11.4 1 4B 20 10 0 10 0 0 39 73 1 9 /p5799.0 9 5.8 1 4B 0 9 0 40 0 30 40 50 2 0 13.5 0 13.5 0 3B 0 8 0 10 5 041 63 2 0 /p5799.0 9 2.7 1 4B 0 10 0 15 0 10 42 56 1 0 /p5799.0 9 0.9 1 4B 6 6 0 15 0 0 43 62 2 0 2.1 1 8.0 1 3A 0 8 0 32 0 2044 57 1 9 /p5799.0 9 0.0 0 3AB 0 24 11 20 0 0 45 56 2 0 12.1 1 16.1 0 3B 0 5 0 25 6 30 46 41 2 1 /p5799.0 9 13.3 1 4B 99 99 99 99 99 99 47 40 2 0 10.1 0 10.1 0 4A 0 15 0 15 17 20 48 81 2 0 /p5799.0 9 0.0 0 4B 99 99 99 99 99 99 49 61 1 0 8.4 0 8.4 0 3B 0 4 0 10 0 850 62 2 0 7.7 0 7.7 0 3AB 0 11 0 16 23 0 51 34 1 0 15.1 1 24.4 1 3B 0 9 15 5 20 0 52 62 2 9 1.1 1 10.5 1 4A 8 99 0 99 0 053 63 2 9 /p5799.0 9 22.2 1 4B 0 0 0 8 3 0 54 56 1 9 /p5799.0 9 7.4 1 4B 0 99 0 6 0 0 55 66 2 9 /p5799.0 9 1.3 1 4B 0 0· 0 22 0 0 56 62 1 0 11.1 0 20.5 1 4B 25 20 0 10 15 18 57 68 2 0 /p5799.0 9 13.8 1 4B 20 99 0 17 15 15 58 45 1 0 /p5799.0 9 6.3 1 4B 28 15 0 10 17 50 59 58 1 9 /p5799.0 9 8.5 0 3B 10 17 0 25 30 25 60 55 1 9 /p5799.0 9 5.8 0 4B 99 99 99 99 99 99 (Continued Overleaf ) 55 Exercise Table 3.3 Continued Family Remission Survival Skin Tests /p68 History of Time Remission Time Survival Patient Age Gender Melanoma /p64 (months )/p65Status /p66 (months )Status /p67Stage Monilia Mumps PPD PHA SK-SD Tricophyton 61 63 2 1 7.3 1 8.7 0 3B 0 9 0 23 12 15 62 53 1 9 36.4 0 36.4 0 3B 5 35 20 17 6 063 45 1 0 /p5799.0 9 5.9 1 4B 0 0 0 0 0 0 64 41 1 0 /p5799.0 9 1.7 1 4B 0 99 0 5 0 6 65 43 1 9 /p5799.0 9 3.9 1 4B 25 0 0 7 20 5 66 80 1 0 5.8 0 11.0 1 4B 99 99 99 99 99 99 67 75 2 9 /p5799.0 9 3.8 1 4B 0 99 0 6 0 0 68 47 2 9 /p5799.0 1 15.9 0 3B 0 0 0 20 15 0 69 64 2 9 6.7 0 6.7 0 3AB 0 5 0 18 0 0 70 38 1 9 /p5799.0 9 1.6 1 4B 99 99 99 99 99 99 71 27 1 0 6.0 0 6.0 0 3B 8 15 20 27 20 1072 56 1 9 /p5799.0 9 4.1 0 4B 0 0 0 0 0 0 73 60 2 9 /p5799.0 9 2.8 0 3A 99 99 99 99 99 99 74 80 2 9 /p5799.0 9 0.2 0 4B 0 20 99 40 0 20 75 38 1 9 /p5799.0 9 7.0 0 4B 0 0 0 15 12 12 76 71 1 9 6.2 0 6.2 0 4A 99 99 99 99 99 99 77 57 2 0 6.1 0 6.1 0 4B 28 20 0 19 20 2078 69 1 0 /p5799.0 9 2.1 0 4B 15 15 15 10 0 0 79 17 2 9 4.9 0 4.9 0 3B 99 99 99 99 99 99 80 64 2 0 /p5799.0 9 1.6 0 4B 99 99 99 99 99 99 81 91 1 0 6.5 0 8.3 1 4A 99 99 99 99 99 99 82 40 2 0 1.7 1 4.6 0 3B 99 99 99 99 99 99 56 83 63 1 0 7.3 0 28.0 1 3A 99 99 99 99 99 99 84 40 1 9 /p5799.0 9 16.1 1 4B 99 99 99 99 99 99 85 53 1 9 /p5799.0 9 4.5 1 4B 99 99 99 99 99 99 86 41 1 0 21.2 0 21.2 0 3A 99 99 99 99 99 9987 27 1 9 /p5799.0 9 4.0 1 4B 99 99 99 99 99 99 88 /p5799 /p5790 /p5799.0 9 7.8 1 4B 99 99 99 99 99 99 89 45 2 9 /p5799.0 9 4.4 0 4B 99 99 99 99 99 99 90 50 2 9 /p5799.0 9 4.2 1 4B 99 99 99 99 99 99 91 47 1 9 /p5799.0 9 1.5 1 4B 99 99 99 99 99 99 92 63 1 9 /p5799.0 9 3.5 0 4B 99 99 99 99 99 99 93 52 1 9 /p5799.0 9 0.4 1 4B 99 99 99 99 99 99 94 53 1 9 /p5799.0 9 2.5 1 4B 99 99 99 99 99 99 95 60 2 9 /p5799.0 9 1.1 0 4B 99 99 99 99 99 99 96 35 1 9 /p5799.0 9 11.1 1 4B 99 99 99 99 99 99 97 24 2 9 1.2 0 1.2 0 3B 99 99 99 99 99 99 98 80 2 0 /p5799.0 9 1.9 0 3A 99 99 99 99 99 99 99 /p5799 /p579 0 4.6 1 6.7 0 3B 99 99 99 99 99 99 100 60 2 9 0.9 0 0.9 0 3AB 99 99 99 99 99 99 101 60 2 9 /p5799.0 9 4.3 0 4B 99 99 99 99 99 99 102 35 1 0 5.2 0 5.2 0 3B 99 99 99 99 99 99 Source: Lee et al. (1979 ) /p631 ,m a l e ;2 ,f e m a l e ; /p579, unknown. /p641, yes; 0, No; 9, unknown. /p65/p5799; never in remission during study period. /p661, relapsed; 0, still in remission; 9, never in remission during study period. /p671, dead; 0, still alive. /p68In millimeters; 99, unknown. 57 Exe rciseTable3.4 Data of 149 Diabe tic Patie nts Variable at Baseline Survival Age at Time Age Diagnosis Smoking SBP DBP Patient Status /p63 (yr)( yr)BMI (yr) Status /p64 (mmHg )(mmHg )ECG /p65CHD /p66 1 1 12.4 44 34.2 41 0 132 96 1 0 2 1 12.4 49 32.6 48 2 130 72 1 03 1 9.6 49 22.0 35 2 108 58 1 14 1 7.2 47 37.9 45 0 128 76 2 15 1 14.1 43 42.2 42 2 142 80 1 06 1 14.1 47 33.1 44 0 156 94 1 0 7 1 12.4 50 36.5 48 0 140 86 2 1 8 1 14.2 36 38.5 33 2 144 88 1 09 1 12.4 50 41.5 47 1 134 78 1 1 10 1 14.5 49 34.1 45 0 102 68 1 011 1 12.4 50 39.5 48 2 142 84 1 012 1 10.8 54 42.9 43 0 128 74 1 0 13 0 10.9 42 29.8 36 2 156 86 1 0 14 1 10.3 44 33.2 43 2 102 58 1 015 0 13.6 40 27.5 26 2 146 98 1 016 1 11.9 48 25.3 48 0 120 68 2 117 1 12.5 50 31.6 44 1 142 76 1 018 1 5.9 47 26.3 38 1 144 82 1 0 19 1 12.4 38 32.4 36 2 150 98 2 1 20 1 14.1 35 47.0 33 1 134 78 1 021 0 9.8 51 26.5 47 2 130 76 1 022 1 7.2 40 43.9 34 0 122 92 1 0 58 23 1 3.5 54 32.3 52 1 132 80 1 0 24 1 0.0 53 34.5 47 2 150 88 3 125 0 12.1 45 18.9 40 1 134 98 1 0 26 1 1.9 41 32.0 31 1 142 90 2 1 27 1 8.6 34 33.9 30 2 124 66 1 028 1 14.0 38 23.7 28 0 102 60 1 029 1 14.3 43 24.8 43 0 134 80 1 030 1 12.4 45 26.6 41 2 118 66 2 131 1 12.4 40 39.2 35 2 192 108 1 0 32 1 14.4 44 32.7 36 2 122 78 1 0 33 1 14.2 48 33.5 43 1 122 92 1 034 1 14.5 51 32.2 49 2 112 74 1 035 1 12.4 36 24.2 30 2 142 90 1 036 1 14.3 52 31.6 48 1 152 96 1 037 0 13.7 41 30.7 39 2 112 74 1 0 38 1 13.4 49 28.0 35 2 118 84 1 0 39 1 12.5 44 32.0 29 0 152 88 1 040 1 14.4 37 32.7 36 2 136 88 1 0 41 1 12.6 51 24.2 42 2 134 90 1 0 42 1 13.8 47 18.7 42 0 130 78 2 143 1 14.0 45 25.6 36 0 108 72 1 0 44 1 6.8 38 22.8 27 2 126 66 2 1 45 1 12.4 35 30.1 33 0 132 78 1 046 1 12.9 50 27.7 49 1 144 88 1 047 1 8.9 53 27.6 49 2 126 68 1 048 1 12.4 48 28.1 47 1 128 70 1 049 1 14.5 40 31.7 37 2 132 82 1 0 50 1 13.0 43 26.1 42 2 128 80 1 0 51 1 13.4 54 30.8 54 1 142 80 2 152 1 10.6 52 36.9 50 1 132 80 2 153 1 13.9 69 24.2 63 1 148 78 1 0 (Continued overleaf ) 59 Exe rciseTable3.4 Continued Variable at Baseline Survival Age at Time Age Diagnosis Smoking SBP DBP Patient Status /p63 (yr)( yr)BMI (yr) Status /p64 (mmHg )(mmHg )ECG /p65CHD /p66 54 1 16.9 38 27.5 26 2 170 100 1 0 55 1 3.6 50 27.3 44 1 140 90 1 056 1 10.2 64 30.1 58 0 138 76 2 157 1 15.7 44 36.1 41 0 112 78 1 058 1 12.0 38 43.1 39 2 140 78 1 059 0 6.7 62 34.6 58 0 138 78 3 1 60 1 11.6 47 39.0 45 0 130 82 1 0 61 0 2.0 78 28.7 77 0 178 86 2 162 1 10.2 49 28.2 43 2 158 80 1 063 1 3.6 63 25.1 46 1 168 88 3 164 1 15.4 71 26.0 59 0 146 88 1 065 1 11.3 51 32.0 49 2 128 76 1 0 66 1 10.3 59 28.1 57 1 132 76 1 1 67 1 5.8 50 26.1 49 1 154 80 1 068 0 8.0 66 45.3 49 0 154 92 1 069 1 14.6 42 30.0 41 1 122 80 1 070 1 11.4 40 35.7 36 2 144 76 2 171 1 7.2 67 28.1 61 0 178 96 1 0 72 1 5.5 86 32.9 61 0 162 60 1 0 73 1 11.1 52 37.6 46 1 142 80 1 074 1 16.5 42 43.4 37 0 120 76 1 075 1 10.9 60 25.4 60 0 124 64 1 076 1 2.5 75 49.7 57 1 174 82 2 1 60 77 0 10.8 81 35.2 81 0 142 88 1 0 78 1 4.7 60 37.3 39 0 160 78 1 079 0 5.5 60 26.0 42 0 122 68 3 1 80 1 4.5 63 21.8 60 2 162 98 1 1 81 1 9.0 62 18.2 43 0 132 72 2 182 1 6.8 57 34.1 41 2 116 60 3 183 0 3.6 71 25.6 54 1 152 84 3 184 1 12.1 58 35.1 45 0 144 68 2 185 1 8.1 42 32.5 28 1 98 68 3 1 86 1 11.1 45 44.1 40 0 138 76 1 1 87 0 7.0 66 29.7 59 1 138 78 1 088 1 1.5 61 29.2 54 0 184 80 2 189 1 11.7 48 25.2 30 2 158 98 1 090 1 0.3 82 25.3 50 0 176 96 1 191 1 13.6 35 25.8 34 1 118 72 1 0 92 1 15.0 57 33.7 57 2 172 98 1 0 93 1 11.2 56 39.5 55 1 182 100 1 194 1 3.0 49 32.9 48 0 144 90 2 1 95 1 13.7 50 37.1 50 0 142 80 1 0 96 1 10.2 53 35.3 53 2 154 76 1 097 1 12.4 71 29.3 70 0 122 60 1 0 98 1 1.1 55 22.1 33 2 222 102 2 1 99 1 16.3 69 23.6 43 0 150 80 1 1 100 1 6.7 59 26.1 55 2 142 66 1 0101 1 15.4 47 32.5 45 2 128 82 1 0102 0 7.6 75 29.8 67 0 122 76 3 1103 0 3.6 80 24.4 80 1 162 88 2 1 104 1 11.5 57 26.3 54 0 172 82 2 1 105 1 13.5 52 30.8 46 2 132 70 1 1106 1 10.6 48 29.4 46 0 112 68 1 0 (Continued overleaf ) 61 Exe rciseTable3.4 Continued Variable at Baseline Survival Age at Time Age Diagnosis Smoking SBP DBP Patient Status /p63 (yr)( yr)BMI (yr) Status /p64 (mmHg )(mmHg )ECG /p65CHD /p66 107 0 6.5 57 29.1 47 1 138 92 2 1 108 0 14.3 58 30.1 56 0 128 74 1 0109 1 11.6 51 31.0 37 2 132 78 1 1110 1 15.4 33 34.0 33 2 120 78 1 0111 1 11.0 36 38.1 33 1 122 70 1 0112 0 11.0 52 37.0 46 0 140 98 1 0 113 0 4.8 64 31.2 57 2 172 88 3 1 114 1 14.8 31 38.8 29 1 136 76 1 0 115 1 1.8 69 22.3 56 0 152 74 3 1 116 1 15.8 59 25.0 58 0 126 80 1 0117 1 14.1 38 31.3 38 2 104 58 1 0118 1 4.6 49 59.7 49 1 142 82 1 0 119 1 15.5 49 34.0 41 0 128 76 1 0 120 0 7.2 68 29.4 66 1 122 58 3 1121 1 14.5 40 43.2 41 1 122 70 1 0122 1 10.5 36 35.1 32 2 122 68 1 0123 1 14.3 60 37.0 54 0 122 70 1 0124 0 2.2 74 27.1 54 1 168 84 2 1 125 1 5.0 61 27.6 51 0 162 82 1 0 126 1 12.4 54 25.2 51 0 116 76 1 0 62 127 1 1.1 35 25.8 34 2 126 82 1 0 128 1 15.4 46 32.2 42 2 180 98 1 0129 1 14.3 40 41.6 41 2 132 98 1 0 130 1 15.6 53 39.8 52 0 150 88 1 0 131 0 12.5 66 26.6 54 1 106 70 1 1132 1 12.3 61 33.3 55 0 154 88 1 0133 1 14.8 41 27.7 38 1 122 76 1 0134 1 10.2 64 26.6 51 2 130 68 1 0135 1 12.3 41 25.0 38 2 120 58 1 0 136 1 10.3 46 54.3 45 1 144 86 1 0 137 1 8.5 80 29.4 79 1 134 60 1 1138 1 10.2 63 33.1 60 1 148 80 2 1139 0 10.0 72 27.3 68 1 170 78 3 1140 1 7.3 41 36.9 33 0 160 92 2 1141 0 15.3 52 40.2 36 0 154 96 1 0 142 1 14.0 53 32.7 48 2 124 76 2 1 143 1 15.8 61 33.2 57 1 130 70 1 0144 1 11.4 53 41.4 47 1 156 78 1 0 145 0 5.5 75 35.8 66 0 162 78 1 0 146 1 11.0 40 34.0 38 2 132 76 1 0147 1 7.3 61 19.9 37 0 120 60 2 1 148 0 10.6 62 30.6 49 0 160 86 2 1 149 1 10.5 49 30.8 47 1 146 86 1 0 /p63Status: 0, dead; 1, alive. /p640, no; ex-smoker; 2, current. /p651, normal; 2, borderline; 3, abnormal. /p660, no; 1, yes. 63 CHAPTER4 NonparametricMethodsof EstimatingSurvivalFunctions Inthischapterwediscussmethodsofestimatingthethreesurvival (survivor- ship, density, and hazard )functions for censored data. Unfortunately, the simple method of Example 2.1 cannot be applied if some of the patients arealive at the time of analysis and therefore their exact survival times areunknown. Nonparametric or distribution-free methods are quite easy tounderstand and apply. They are less efficient than parametric methods whensurvival times followa theoretical distribution and more efficient w hen nosuitable theoretical distributions are known. Therefore, we suggest usingnonparametric methods to analyze survival data before attempting to fit atheoreticaldistribution. If the main objective is to find a model for the data,estimates obtained by nonparametric methods and graphs can be helpful inchoosingadistribution. Of the three survival functions, survivorship or its graphical presentation, the survival curve, is the most widely used. Section 4.1 introduces theproduct-limit (PL)method of estimating the survivorshipfunction developed byKaplanandMeier (1958 ).Withtheincreasedavailabilityofcomputers,this method is applicable to small, moderate, and large samples. However, if thedatahavealreadybeengroupedintointervals,orthesamplesizeisverylarge,sayinthe thousands,orthe interest isin alarge population,it maybe moreconvenient to perform a life-table analysis. Section 4.2 is devoted to thediscussionofpopulationandclinicallifetables.ThePLestimatesandlife-tableestimatesofthesurvivorshipfunctionareessentiallythesame.Manyauthorsusetheterm life-table estimates forthePLestimates.Theonlydifferenceisthat thePLestimateisbasedonindividualsurvivaltimes,whereasinthelife-tablemethod, survival times are grouped into intervals. The PL estimate can beconsidered as a special case of the life-table estimate where each intervalcontainsonlyoneobservation. 64 In Section 4.3 we discuss three other measures that describe the survival experience: the relative survival rate, the five-year survival rate, and thecorrected survival rate. In Section 4.4 we describe two methods, direct andindirectstandardization,toadjustratestoeliminatetheeffectofdifferencesinpopulationcompositionwithrespecttoageandothervariables.Inaddition,itintroducesthestandardizedmortalityrateandstandardizedincidencerate. 4.1 PRODUCT-LIMIT ESTIMATES OF SURVIVORSHIP FUNCTION Letusfirstconsiderthesimplecasewhereallthepatientsareobservedtodeath sothat the survivaltimes areexact and known.Let t/p16,t/p17,...,t/p76be the exact survivaltimesofthe nindividualsunderstudy.Conceptually,weconsiderthis groupofpatientsasarandomsamplefromamuchlargerpopulationofsimilarpatients.We relabelthe nsurvival times t/p16,t/p17,...,t/p76in ascendingordersuch thatt/p7/p16/p8/p45t/p7/p17/p8/p45···/p45t/p7/p76/p8.Following (2.1.2 )and (2.1.3 ),thesurvivorshipfunc- tionatt/p7/p71/p8canbeestimatedas S/p19(t/p7/p71/p8)/p58n/p57i n /p581/p57i n(4.1.1) wheren/p57iisthenumberofpeopleinthesamplesurvivinglongerthan t/p7/p71/p8.If two or more t/p7/p71/p8are equal (tied observations ),the largest ivalue is used. For example,if t/p7/p17/p8/p58t/p7/p18/p8/p58t/p7/p19/p8,then S/p19(t/p7/p17/p8)/p58S/p19(t/p7/p18/p8)/p58S/p19(t/p7/p19/p8)/p58n/p574 n Thisgivesaconservativeestimateforthetiedobservations. Sinceeverypersonisaliveatthebeginningofthestudyandnoonesurvives longerthan t/p7/p76/p8, S/p19(t/p7/p15/p8)/p581 and S/p19(t/p7/p76/p8)/p580( 4 .1.2) Inpractice, S/p19(t) iscomputedatevery distinctsurvivaltime.Wedonothaveto worryabouttheintervalsbetweenthedistinctsurvivaltimesinwhichnoonediesand S/p19(t) remainsconstant.Equations (4.1.1 )and (4.1.2 )showthat S/p19(t)i s a step function starting at 1.0 and decreasing in steps of 1/ n(if there are no ties)to zero.When S/p19(t) isplottedversus t,the variouspercentilesofsurvival timecanbereadfromthegraphorcalculatedfrom S/p19(t). Thefollowingexample illustratesthemethod. Example 4.1 Consideraclinicaltrialinwhich10lungcancerpatientsare followedtodeath.Table4.1liststhesurvivaltimes tinmonths.Thefunction-     65 Table 4.1 Computation of S/p19(t) for 10 Lung Cancer Patients ti S /p19(t) 41 /p24/p16/p15/p580.9 52 /p23/p16/p15/p580.8 63 /p22/p16/p15/p580.7 84 /p19/p16/p15/p580.4 85 /p19/p16/p15/p580.4 86 /p19/p16/p15/p580.4 10 7 /p17/p16/p15/p580.2 10 8 /p17/p16/p15/p580.2 11 9 /p16/p16/p15/p580.1 12 10 /p15/p16/p15/p580.0 S/p19(t) iscomputedfollowing (4.1.1 )andplottedasastepfunctioninFigure4.1 a andasasmoothcurveinFigure4.1 b.Theestimatedmediansurvivaltimeis8 months from Figure 4.1 aor 7.6 months from Figure 4.1 b. A more accurate estimatecanbeobtainedusinglinearinterpolation: tS /p19(t) 60 .7 m0.5 80 .4 8/p576 0.4/p570.7/p588/p57m 0.4/p570.5 m/p588/p572(0.1) 0.3/p587.3(months ) Theoretically, S/p19(t) should be plotted as a step function since it remains constant between two observed exact survival times. However, when themediansurvivaltimemustbeestimatedfromasurvivalcurve,asmoothcurve(suchasFigure4.1 b)maygiveamuchbetterestimatethanastepfunction,as indicatedintheexample. Thismethodcanbeappliedonlyifallthepatientsarefollowedtodeath.If someofthepatientsarestill aliveat the endofthestudy,adifferentmethodofestimating S/p19(t), suchasthePLestimategivenbyKaplanandMeier (1958 ), isrequired.Therationalecanbeillustratedbythefollowingsimpleexample. Suppose that 10 patients join a clinical study at the beginning of 2000;66       Figure 4.1Function S/p19(t) oflungcancerpatientsinExample4.1. during that year 6 patients die and 4 survive. At the end of the year, 20 additional patients join the study. In 2001, 3 patients who entered in thebeginningof2000and15patientswhoenteredlaterdie,leavingoneandfivesurvivors, respectively. Suppose that the study terminates at the end of 2001and you want to estimate the proportion of patients in the populationsurvivingfortwoyearsormore,thatis, S(2). The first group of patients in this example is followed for two years; the second group is followed for only one year. One possible estimate, thereduced-sample estimate ,i sS/p19(2)/p581/10/p580.1, which ignores the 20 patients whoarefollowedonlyforoneyear.KaplanandMeierbelievethatthesecondsample,underobservationforonlyoneyear,cancontributetotheestimateofS(2). Patients who survived two years may be considered as surviving the first yearandthensurvivingonemoreyear.Thus,theprobabilityofsurvivingfortwo years or more is equal to the probability of surviving the first year andthensurvivingonemoreyear.Thatis, S(2)/p58P(survivingfirstyearandthensurvivingonemoreyear ) whichcanbewrittenas S(2)/p58P(survivingtwoyearsgivenpatienthassurvivedfirstyear ) /p59P(survivingfirstyear )( 4.1.3 ) TheKaplan —Meierestimateof S(2)following (4.1.3 )is S/p19(2)/p58 /p1proportion of patients surviving two years given they survive for one year /p2 /p59(proportionofpatientssurvivingoneyear )(4.1.4 )-     67 For the data given above, one of the four patients who survived the first yearsurvivedtwo years,so the first proportionin (4.1.4 )is/p16/p19. Fourof the 10 patients who entered at the beginning of 2000 and 5 of the 20 patients whoenteredattheendof2000survivedoneyear.Therefore,thesecondproportionin(4.1.4 )is(4/p595)/(10/p5920).ThePLestimateof S(2)is S/p19(2)/p581 4 /p594/p595 10/p5920/p580.25/p590.3/p580.075 Thissimplerulemaybegeneralizedasfollows:Theprobabilityofsurviving k(/p462) or more years from the beginning of the study is a product of k observedsurvivalrates: S/p19(k)/p58p/p16/p59p/p17/p59p/p18/p59···/p59p/p73(4.1.5 ) wherep/p16denotestheproportionofpatientssurvivingatleastoneyear, p/p17the proportionofpatientssurvivingthesecondyearaftertheyhavesurvivedoneyear,p/p18the proportion of patients surviving the third year after they have survived two years, and p/p73the proportion of patients surviving the kth year aftertheyhavesurvived k/p571years. Therefore, the PL estimate of the probability of surviving any particular number of years from the beginning of study is the product of the sameestimate up to the preceding year, and the observed survival rate for theparticularyear,thatis, S/p19(t)/p58S/p19(t/p571)p/p82(4.1.6) ThePLestimatesaremaximumlikelihoodestimates. Inpractice,thePLestimatescanbecalculatedbyconstructingatablewith fivecolumnsfollowingtheoutlinebelow. 1. Column1containsallthesurvivaltimes,bothcensoredanduncensored, in order from smallest to largest. Affix a plus sign to the censoredobservation.Ifacensoredobservationhasthesame valueasanuncen-soredobservations,thelattershouldappearfirst. 2. Thesecondcolumn,labeled i,consistsofthecorrespondingrankofeach observationincolumn1. 3. The third column, labeled r, pertains to uncensored observations only. Letr/p58i. 4. Compute( n/p57r)/(n/p57r/p591), orp/p71,foreveryuncensoredobservation t/p7/p71/p8incolumn4togivetheproportionofpatientssurvivinguptoandthen throught/p7/p71/p8.68       Table 4.2 Calculation of the PL Estimate of S/p19(t) for Data in Example 4.2 RemissionTime Rank ti r (n/p57r)/(n/p57r/p591) S/p19(t) 3.01 1 /p24/p16/p15/p24/p16/p15/p580.900 4.0/p59 2— — — 5.7/p59 3— — — 6.5 4 4 /p21/p22/p24/p16/p15/p59/p21/p22/p580.771 /p63 6.5 5 5 /p20/p21/p24/p16/p15/p59/p21/p22/p59/p20/p21/p580.643 /p63 8.4/p59 6— — — 10.0 7 7 /p18/p19/p24/p16/p15/p59/p21/p22/p59/p20/p21/p59/p18/p19/p580.482 10.0/p59 8— — — 12.0 9 9 /p16/p17/p24/p16/p15/p59/p21/p22/p59/p20/p21/p59/p18/p19/p59/p16/p17/p580.241 15.0 10 10 0 0 /p630.643isusedas S/p19(6.5).Itisaconservativeestimate.5. Incolumn5, S/p19(t) istheproductofallvaluesof( n/p57r)/(n/p57r/p591) upto and including t. If some uncensored observations are ties, the smallest S/p19(t) shouldbeused. To summarizethis procedure,let nbe the total number of patientswhose survival times, censored or not, are available. Relabel the nsurvival times in orderofincreasingmagnitudesuchthat t/p7/p16/p8/p45t/p7/p17/p8/p45···/p45t/p7/p76/p8.Then S/p19(t)/p58/p147 t/p7/p80/p8/p45tn/p57r n/p57r/p591(4.1.7) whererruns through those positive integers for which t/p7/p80/p8/p45tandt/p7/p80/p8is uncensored.Thevaluesof rareconsecutiveintegers1,2,..., nifthereareno censoredobservations;iftherearecensoredobservations,theyarenot. Theestimatedmediansurvivaltimeisthe50thpercentile,whichisthevalue oftatS/p19(t)/p580.50.Thefollowingexampleillustratesthecalculationprocedures. Example 4.2 Supposethatthefollowingremissiondurationsareobserved from10patients (n/p5810)withsolidtumors.Sixpatientsrelapseat3.0,6.5,6.5, 10, 12, and 15 months; 1 patient is lost to follow-up at 8.4 months; and 3patients are still in remission at the end of the study after 4.0, 5.7, and 10months.Thecalculationof S/p19(t) isshowninTable4.2. Thesurvivorshipfunction S/p19(t) isplottedinFigure4.2;theestimatedmedian remissiontimeis m/p589.8months.Fromthecalculationwenoticethat S/p19(t)a t-     69 Figure 4.2Function S/p19(t) ofExample4.2. t/p58t/p7/p71/p8isrelatedto S/p19(t)a tt/p58t/p7/p71/p92/p16/p8and (4.1.6 )canberewrittenas S/p19(t/p7/p71/p8)/p58S/p19(t/p7/p71/p92/p16/p8)n/p57i n/p57i/p591(4.1.8 ) wheret/p7/p71/p8andt/p7/p71/p92/p16/p8areuncensoredobservations.Forexample, S/p19(12) /p58S/p19(10)/p59/p16/p17/p580.482/p59/p16/p17/p580.241 Iftherearenocensoredobservationsorlossesbefore t,(4.1.7 )isequivalentto (4.1.1 ). ThevarianceofthePLestimateof S/p19(t) isapproximatedby Var[S/p19(t)]/p60[S/p19(t)]/p17/p26 /p801 (n/p57r)(n/p57r/p591)(4.1.9) whererincludes thosepositiveintegers for which t/p7/p80/p8/p45tandt/p7/p80/p8corresponds toadeath.ForthedatainExample4.2,forexample, Var[S/p19(10)] /p58(0.482)/p17/p11 9/p5910/p591 6/p597/p591 5/p596/p591 3/p594/p2 /p580.0352 andtheestimatedstandarderroris0.1876.InExample4.1, Var[S/p19(6)]/p58(0.7)/p17/p11 9/p5910/p591 8/p599/p591 7/p598/p2/p580.021070       andtheestimatedstandarderroris0.145.Thevariancemaybeusedtoobtain confidenceintervalsfor S(t). CalculationofthePLestimateof S(t) inExample4.2canalsobeobtained byusingstatisticalsoftware.Let tdenotetheobservedremissiontime (uncen- sored or censored )in Table 4.2 and CENS denote an index (or dummy ) variablewithCENS /p580iftiscensoredand1otherwise.Assumethatthedata havebeensavedin‘‘C: /p33D4d2.DAT’’asatextfile,whichcontainstwocolumns, tandCENS,separatedbyaspace. The following SAS code can be used to obtain the PL estimate of S(t)i n Table4.2.Onecanadopt thiscode to obtainthePLestimateof S(t) forany observeduncensoredorcensoredsurvivaltimedata. dataw1; infile‘c: /p33d4d2.dat’missover; inputtcens; run;proclifetestdata /p58w1outsurv /p58wa; timet*cens (0); run; title’PLestimateofsurvivalfunction’;procprintdata /p58wa; run; IfBMDP1Lisused,thefollowingcodecanbeused. /input file /p58‘c:/p33d4d2.dat’. variables /p582. format /p58free. /variable names /p58t,cens. /form time /p58t. status /p58cens. response /p581. /estimate method /p58product. Print. /end IftheSPSSKMprocedureisused,thefollowingcodecanbeused. datalistfile /p58‘c:/p33d4d2.dat’free /tcens. kmt /status /p58censevent (1) /print. Example 4.3 Consider the tumor-free time in days of the 30 rats on a low-fatdietinTable3.4.Table4.3givesthecalculationsofthePLestimatesof S(t) andthestandarderrorof S/p19(t). Theestimated S(t) isplottedinFigure3.3. Themediantumor-freetimeisapproximately189days.-     71 Table 4.3 Calculation of S/p19(t) and Standard Error of S/p19(t) for 30 Rat s on a Low-FatDietin Table 3.4 Rat Tumor-free Number Time ti rn/p57r n/p57r/p591S/p19(t) StandardErrorof S/p19(t) 35 0 1 1 /p17/p24/p18/p15/p17/p24/p18/p15/p580.967 /p3(0.967 )/p17/p11 29/p5930/p2/p4/p16/p30/p17/p580.033 12 56 2 2 /p17/p23/p17/p240.967/p59/p17/p23/p17/p24/p580.933 /p3(0.933 )/p17/p11 29/p5930/p591 28/p5929/p2/p4/p16/p30/p17/p580.046 46 5 3 3 /p17/p22/p17/p230.933/p59/p17/p22/p17/p23/p580.900 /p3(0.900 )/p17/p11 29/p5930/p591 28/p5929/p591 27/p5928/p2/p4/p16/p30/p17/p580.055 13 66 4 4 /p17/p21/p17/p220.900/p59/p17/p21/p17/p22/p580.867 0.062 14 73 5 5 /p17/p20/p17/p210.833 0.068 97 7 6 6 /p17/p19/p17/p200.800 0.073 10 84 7 7 /p17/p18/p17/p190.767 0.077 58 6 8 8 /p17/p17/p17/p180.733 0.081 11 87 9 9 /p17/p16/p17/p170.700 0.084 15 119 10 10 /p17/p15/p17/p160.667 0.086 1 140 11 11 /p16/p24/p17/p150.633 0.088 16 140 /p5912 — — — 72 6 153 13 13 /p16/p22/p16/p230.633/p59/p16/p22/p16/p23/p580.598 0.090 2 177 14 14 /p16/p21/p16/p220.598/p59/p16/p21/p16/p22/p580.563 0.091 7 181 15 15 /p16/p20/p16/p210.528 0.092 8 191 16 16 /p16/p19/p16/p200.493 0.092 17 200 /p5917 — — — — 18 200 /p5918 — — — — 19 200 /p5919 — — — — 20 200 /p5920 — — — — 21 200 /p5921 — — — — 22 200 /p5922 — — — — 23 200 /p5923 — — — — 24 200 /p5924 — — — — 25 200 /p5925 — — — — 26 200 /p5926 — — — — 27 200 /p5927 — — — — 28 200 /p5928 — — — — 29 200 /p5929 — — — — 30 200 /p5930 — — — — 73 The mean survival time /afii9839can be shown to equal the area under the estimatedsurvivorshipfunction.Toestimate /afii9839,wecanuse /afii9839/p24/p58/p16/p27 /p15S/p19(t)dt thatis, /afii9839/p24isequaltotheareaundertheestimatedsurvivorshipfunction.Thus, if the times to death are ordered as t/p7/p16/p8/p45t/p7/p17/p8/p45···/p45t/p7/p75/p8(if there are m uncensored observations )andt/p7/p75/p8is the largest observation of all nobserva- tions[i.e., t/p7/p75/p8/p58t/p7/p76/p8whent/p7/p76/p8isanuncensoredobservation], /afii9839canbeestimated as /afii9839/p24/p581.000t/p7/p16/p8 /p59S/p19(t/p7/p16/p8)(t/p7/p17/p8 /p57t/p7/p16/p8)/p59S/p19(t/p7/p17/p8)(t/p7/p18/p8 /p57t/p7/p17/p8)/p59··· /p59S/p19(t/p7/p75/p92/p16/p8)(t/p7/p75/p8/p57t/p7/p75/p92/p16/p8) (4.1.10 ) whichisthesumoftheareasoftherectanglesunderthesurvivalcurveformed by the uncensored observations. Consider the data in Example 4.2: m/p586, t/p7/p16/p8 /p583.0,t/p7/p17/p8 /p586.5,t/p7/p18/p8 /p586.5,t/p7/p19/p8 /p5810,t/p7/p20/p8 /p5812, and t/p7/p21/p8 /p5815.The mean survivaltimeisestimatedusing (4.1.10 )as /afii9839/p24/p581.000/p593.0/p590.900 (6.5/p573.0)/p590.643 (10/p576.5) /p590.482 (12/p5710)/p590.241 (15/p5712) /p583.000 /p593.150 /p592.251 /p590.964 /p590.723 /p5810.088months However, if the largest observation in the data is censored and is used as t/p7/p75/p8in(4.10),/afii9839soobtainedmaybealowestimate.Insuchcases,Irwin (1949 ) suggeststhatinsteadofestimatingthemeansurvivaltime,oneshouldchooseatimelimit Landestimatethe‘‘meansurvivaltimelimitedtoatime L,’’say /afii9839/p9/p42/p10, byusing Lfort/p7/p75/p8in(4.1.10 ).For example,if inExample4.2 thelargest observationiscensored,thatis,15 /p59,andifwelet L/p5816,then /afii9839/p9/p16/p21/p10/p583.000 /p593.150; /p592.251 /p590.964 /p590.241 (16/p5712) /p5810.329months whichisthemeansurvivaltimelimitedto16months. Thevarianceof /afii9839/p24isestimatedby Var(/afii9839/p24)/p58/p26 /p80A/p17/p80(n/p57r)(n/p57r/p591)(4.1.11) whererrunsthroughthoseintegersfor which t/p80correspondsto adeath,and74       A/p80is theareaunderthecurve S/p19(t) to theright of t/p7/p80/p8. ThekthA/p80interms of themuncensoredobservationsis S/p19(t/p7/p73/p8)(t/p7/p73/p62/p16/p8 /p57k/p7/p73/p8)/p59S/p19(t/p7/p73/p62/p16/p8)(t/p7/p73/p62/p17/p8 /p57t/p7/p73/p62/p16/p8)/p59···/p59S/p19(t/p7/p75/p92/p16/p8)(t/p7/p75/p8/p57t/p7/p75/p92/p16/p8) (4.1.12) If thereare nocensoredobservations, (4.1.10 )reducesto thesamplemean t/p16/p58/afii9814t/p71/n,and (4.1.11 )reducesto Var(/afii9839/p24)/p58Var(t/p16)/p58/p26(t/p71/p57t/p16)/p17 n/p17(4.1.13) whichisnotanunbiasedestimate.KaplanandMeiersuggestthat (4.1.11 )and (4.1.13 )be multiplied by m/(m/p571) andn/(n/p571), respectively, to correct the bias. ConsiderthesurvivaltimesinExample4.1:Thesamplemeanis t/p16/p58/afii9839/p24/p588.2 months and the estimated variance of /afii9839/p24,b y (4.1.13 ), is 0.616. If the factor n/(n/p571)/p5810/9ismultiplied,theestimatedvarianceof /afii9839becomes0.684. Tocomputethevarianceof /afii9839/p24inExample4.2,wefirstcomputethefive A/p80’s: A/p16,A/p19,A/p20,A/p22,andA/p24.Thefirst A/p80is A/p16/p58S/p19(t/p7/p16/p8)(t/p7/p17/p8 /p57t/p7/p16/p8)/p59S/p19(t/p7/p17/p8)(t/p7/p18/p8 /p57t/p7/p17/p8)/p59···/p59S/p19(t/p7/p20/p8)(t/p7/p21/p8 /p57t/p7/p20/p8) /p583.150/p592.251/p590.964/p590.723/p587.088 Thesecond A/p80is A/p19/p58S/p19(t/p7/p17/p8)(t/p7/p18/p8 /p57t/p7/p17/p8)/p59···/p59S/p19(t/p7/p20/p8)(t/p7/p21/p8 /p57t/p7/p20/p8) /p582.251/p590.964/p590.723/p583.938 Thethird,fourth,andfifth A/p80’sare,respectively, A/p20/p582.251/p590.964/p590.723/p583.938 A/p22/p580.964/p590.723/p581.687 A/p24/p580.723 Thus, Var/p19(/afii9839/p24)/p58(7.088 )/p17 9/p5910/p59(3.938 )/p17 6/p597/p59(3.928 )/p17 5/p596/p59(1.687 )/p17 3/p594/p59(0.723 )/p17 1/p592/p581.942 The estimated standard error of /afii9839/p24is 1.394. If the factor m/(m/p571)/p586/5i s included,theseresultsbecome2.330and1.526,respectively. The Kaplan —Meier method provides very useful estimates of survival probabilitiesandgraphicalpresentationofsurvivaldistribution.Itisthemost-     75 Figure 4.3Kaplan—Meierestimateofmediansurvivaltime.widelyusedmethodinsurvivaldataanalysis.BreslowandCrowley (1974 )and Meier (1975b )have shown that under certain conditions, the estimate is consistent and asymptomatically normal. However, a few critical featuresshouldbementioned. 1. TheKaplan —Meierestimatesarelimitedtothetimeintervalinwhichthe observations fall. If the largest observation is uncensored, the PLestimate at that time equals zero. Although the estimate may not bewelcomed by physicians, it is correct since no one in the sample liveslonger.Ifthelargestobservationiscensored,thePLestimatecanneverequalzeroandisundefinedbeyondthelargestobservation. 2. The most commonly used summary statistic in survival analysis is the mediansurvivaltime.Asimpleestimateofthemediancanbereadfromsurvival curves estimated by the PL method as the time tat which S/p19(t)/p580.5. However, the solution may not be unique. Consider Figure 4.3a,wherethe survivalcurveis horizontalat S/p19(t)/p580.5;anytvalue in the interval t/p16tot/p17is a reasonableestimate of the median. A practical solutionistotakethemidpointoftheintervalasthePLestimateofthemedian.Figure4.3 bpresentsadifferentcaseinwhichthestraightforward estimate (t/p16)tendstooverestimatethemedian.Apracticalwaytohandle thisproblemistoconnectthepointsandlocatethemedian. 3. If less than 50% of the observations are uncensored and the largest observationiscensored,themediansurvivaltimecannotbeestimated.Apracticalwayto handlethesituationistouseprobabilitiesofsurvivinga given length of time, say 1, 3, or 5 years, or the mean survival timelimitedtoagiventime t. 4. ThePLmethodassumesthatthecensoringtimesareindependentofthe survivaltimes.Inotherwords,the reasonan observationis censoredis unrelatedtothecauseofdeath.Thisassumptionistrueifthepatientis76       still alive at the end of the study period. However, the assumption is violated if the patient develops severe adverse effects from the treat-ment and is forced to leave the study before death or if the patientdied of a cause other than the one under study (e.g., death due to automobile accidents in a cancer survival study ). When there is inap- propriate censoring, the PL method is not appropriate. In practice,one way to alleviate the problem is to avoid it or to reduce it to aminimum. 5. Similar to other estimators, the standard error (S.E.)of the Kaplan — Meierestimatorof S(t) givesanindicationofthepotentialerrorof S/p19(t). The confidence interval deserves more attention than just the pointestimate S/p19(t). A95%confidenceintervalfor S(t)i sS/p19(t)/p591.96S.E.[S/p19(t)]. 4.2 LIFE-TABLE ANALYSIS Thelife-table methodisone of theoldesttechniquesfor measuringmortality and describing the survival experience of a population. It has been used byactuaries, demographers, governmental agencies, and medical researchers instudies of survival, population growth, fertility, migration, length of marriedlife,lengthofworkinglife,andsoon.TherehasbeenadecennialseriesoflifetablesontheentireU.S.populationsince1900.Statesandlocalgovernmentsalsopublishlifetables.Theselifetables,summarizingthemortalityexperienceofaspecificpopulationfora specificperiodoftime,arecalled population life tables.As clinical and epidemiologic research become more common, the life-tablemethodhas been appliedto patientswitha given diseasewhohavebeen followed for a period of time. Life tables constructed for patients arecalledclinical life tables. Althoughpopulationandclinicallifetablesaresimilar incalculation,thesourcesofrequireddataaredifferent. 4.2.1 Population Life Tables Therearetwokindsofpopulationlifetables:thecohortlifetableandcurrent life table. The cohort life table describes the survival or mortality experience frombirthtodeathofaspecificcohortofpersonswhowerebornataboutthesametime,forexample,allpersonsbornin1950.Thecohorthastobefollowedfrom1950untilallofthemdie.Theproportionofdeath (survivor )isthenused toconstructlifetablesforsuccessivecalendaryears.Thistypeoftable,usefulinpopulationprojectionandprospectivestudies,isnotoftenconstructedsinceitrequiresalongfollow-upperiod. Thecurrent life table is constructed by applying the age-specific mortality rates of a population in a given period of time to a hypothetical cohort of100,000or1,000,000persons.Thestartingpointisbirthatyear0.Twosourcesofdataarerequiredforconstructingapopulationlifetable: (1)censusdataon-  77 thenumberoflivingpersons ateach agefor agiven yearat midyearand (2) vital statistics on the number of deaths in the given year for each age. Forexample, a current U.S. life table assumes a hypothetical cohort of 100,000persons that is subject to the age-specific death rates based on the observeddatafortheUnitedStatesinthe1990census.Thecurrentlifetable,basedonthelifeexperienceofanactualpopulationoverashortperiodoftime,givesagood summary of current mortality. This type of life table is regularlypublished by government agencies of different levels. One of the most often reported statistics from current life tables is the life expectancy. The termpopulation life table isoftenusedtorefertothecurrentlifetable. In the United States, the National Center for Health Statistics publishes detailed decennial life tables after each decennial census. These complete lifetables use one-year age groups. Between censuses, annual life tables are alsopublished.The annuallife tables are often seen in five-year age intervals andarecalled abridged life tables .Tables 4.4and 4.5 are,respectively,acomplete decennial life table for the total U.S. population for 1989 —1991 and an abridged life table for the same population for 1998. The abridged table inTable4.5wasconstructedbasedonacompletelifetable. Currentlifetablesusuallyhavethefollowingcolumns: 1.Age interval[xtox/p59t).Thisisthetimeintervalbetweentwoexactages xandx/p59t;tis the length of the interval. For example, the interval 20—21 includes the time interval from the 20th birthday up to the 21st birthday (butnotincludingthe21stbirthday ). 2.Proportionof persons aliveat beginning of age intervalbut dying during the interval(/p82q/p86). The information is obtained from census data. For example, (/p82q/p86)for age interval 20 —21 is the proportion of persons who diedonoraftertheir20thbirthdayandbeforetheir21stbirthday.Itisanestimateoftheconditionalprobabilityofdyingin theintervalgiventhepersonisaliveatage x.Thiscolumnisusuallycalculatedfromdata ofthedecennialcensusofpopulationanddeathsoccurringinthegiventimeinterval.Forexample,themortalityratesinTable4.4arecalculatedfromthedataofthe1990CensusofPopulationanddeathsoccurringinthe United States in the three years 1989 —1991. This column is the foundation of the life table from which all of the other columns arederived. 3.Number living at beginning of age interval (l/p86).Theinitialvalueof l/p86,the sizeofthehypotheticalpopulation,isusually100,000or1,000,000.Thesuccessivevaluesarecomputedusingtheformula l/p86/p58l/p86/p92/p16(1/p57/p82q/p86/p92/p82)( 4.2.1 ) where1 /p57/p82q/p86/p92/p82istheproportionofpersonswhosurvivedthepreviousage interval. For example, in Table 4.4, t/p581,l/p17/p15/p58l/p16/p24(1/p57/p16q/p16/p24)/p5878       98,314(1 /p570.00101) /p5898,215, which is the number of persons living at thebeginningofage20. 4.Number dying during age interval (/p82d/p86) /p82d/p86/p58l/p86(/p82q/p86)/p58l/p86/p57l/p86/p62/p16(4.2.2) For example, the number of persons dying during age interval 20 —21, /p82d/p17/p15/p5898,215(0.00104) /p58102(or/p16d/p17/p15/p5898,215 /p5798,113 /p58102). 5. Stationary population (/p82L/p86andT/p86).Here/p82L/p86isthetotalnumberofyears livedinthe ithageintervalorthenumberofperson-yearsthat l/p86persons, agedxexactly, live through the interval. For those who survive the interval,theircontributionto/p82L/p86isthelengthoftheinterval, t.Forthose whodieduringtheinterval,wemaynotknowexactlythetimeofdeathandthesurvivaltimemustbeestimated.Theconventionalassumptionisthattheyliveone-halfoftheintervalandcontribute t/2tothecalculation of/p82L/p86.Thus, /p82L/p86/p58t(l/p86/p62/p16/p59/p16/p17 /p82d/p86)( 4.2.3 ) Forexample,inTable4.4,/p16L/p17/p15/p5898,113 /p59102/2/p5898,164.Ifwedoknow the exact survival time of those who die in the interval,/p82L/p86should be computedaccordingly. Thesymbol T/p86isthetotalnumberofperson-yearslivedbeyondage t bypersonsaliveatthatage,thatis, T/p86/p58/p26 j/p46x/p82L/p72(4.2.4) and T/p86/p58/p82L/p86/p59T/p86/p62/p82(4.2.5) Forexample,inTable4.4, T/p15/p587,536,614,whichisthesumofall/p82L/p86values incolumn5,and T/p16/p587,437,356,whichis T/p15/p57/p16L/p15/p587,536,614 /p5799,258. 6.Average remaining lifetime or average number of years of life remaining at beginning of age interval (e /p4/p71).Thisisalsoknownasthe life expectancy at agivenage,whichisdefinedasthenumberofyears remainingtobelived bypersonsatage x: e /p4/p86/p58T/p86l/p86(4.2.6) Theexpectedageatdeathofapersonaged xisx/p59e /p4/p86.Thee /p4/p86atx/p580is the life expectancy at birth. For example, according to the U.S. life-  79 Table 4.4 Life Table for the Total Population, United States, 1989--1991 Proportion Of100,000 StationaryAverage Dying BornAlive PopulationRemainingLifetime Proportionof Average Age PersonsAlive Number Number InThis Numberof Interval atBeginningof Livingat Dying andAll YearsofLife AgeInterval Beginning During Inthe Subsequent Remainingat PeriodofLife DyingDuring ofAge Age Age Age BeginningofBetweenTwoAges Interval Interval Interval Interval Intervals AgeInterval xtox/p59t/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p15 Days 0—1 0.00351 100,000 351 274 7,536,614 75.37 1—7 0.00135 99,649 134 1,637 7,536,340 75.63 7—28 0.00104 99,515 104 5,722 7,534,703 75.71 28—365 0.00349 99,411 347 91,625 7,528,981 75.74 Years 0—1 0.00936 100,000 936 99,258 7,536,614 75.37 1—2 0.00073 99,064 72 99,028 7,437,356 75.08 2—3 0.00048 98,992 48 98,968 7,338,328 74.13 3—4 0.00037 98,944 37 98,926 7,239,360 73.17 4—5 0.00030 98,907 30 98,892 7,140,434 72.19 5—6 0.00027 98,877 27 98,863 7,041,542 71.22 6—7 0.00025 98,850 24 98,839 6,942,679 70.23 7—8 0.00023 98,826 23 98,814 6,843,840 69.25 8—9 0.00020 98,803 20 98,794 6,745,026 68.27 9—10 0.00018 98,783 17 98,774 6,646,232 67.28 80 10—11 0.00016 98,766 16 98,758 6,547,458 66.29 11—12 0.00016 98,750 16 98,742 6,448,700 65.30 12—13 0.00022 98,734 21 98,723 6,349,958 64.31 13—14 0.00032 98,713 32 98,697 6,251,235 63.33 14—15 0.00047 98,681 46 98,658 6,152,538 62.35 15—16 0.00063 98,635 62 98,604 6,053,880 61.38 16—17 0.00077 98,573 76 98,534 5,955,276 60.41 17—18 0.00089 98,497 88 98,453 5,856,742 59.46 18—19 0.00096 98,409 95 98,362 5,758,289 58.51 19—20 0.00101 98,314 99 98,265 5,659,927 57.57 20—21 0.00104 98,215 102 98,164 5,561,662 56.63 21—22 0.00109 98,113 107 98,060 5,463,498 55.69 22—23 0.00112 98,006 110 97,951 5,365,438 54.75 23—24 0.00114 97,896 112 97,840 5,267,487 53.81 24—25 0.00116 97,784 113 97,727 5,169,647 52.87 25—26 0.00117 97,671 115 97,614 5,071,920 51.93 26—27 0.00119 97,556 115 97,499 4,974,306 50.99 27—28 0.00121 97,441 119 97,381 4,876,807 50.05 28—29 0.00126 97,322 123 97,261 4,779,426 49.11 29—30 0.00133 97,199 129 97,135 4,682,165 48.17 30—31 0.00140 97,070 136 97,002 4,585,030 47.23 31—32 0.00147 96,934 143 96,862 4,488,028 46.30 32—33 0.00154 96,791 149 96,717 4,391,166 45.37 33—34 0.00162 96,642 157 96,563 4,294,449 44.44 34—35 0.00170 96,485 163 96,404 4,197,886 43.51 35—36 0.00178 96,322 172 96,236 4,101,482 42.58 36—37 0.00188 96,150 181 96,060 4,005,246 41.66 37—38 0.00198 95,969 189 95,874 3,909,186 40.73 38—39 0.00207 95,780 199 95,681 3,813,312 39.81 39—40 0.00217 95,581 208 95,477 3,717,631 38.90 (Continued overleaf ) 81 Table 4.4 Continued Proportion Of100,000 StationaryAverage Dying BornAlive PopulationRemainingLifetime Proportionof Average Age PersonsAlive Number Number InThis Numberof Interval atBeginningof Livingat Dying andAll YearsofLife AgeInterval Beginning During Inthe Subsequent Remainingat PeriodofLife DyingDuring ofAge Age Age Age BeginningofBetweenTwoAges Interval Interval Interval Interval Intervals AgeInterval xtox/p59t/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p15 40—41 0.00228 95,373 217 95,265 3,622,154 37.98 41—42 0.00240 95,156 228 95,042 3,526,889 37.06 42—43 0.00254 94,928 241 94,808 3,431,847 36.15 43—44 0.00271 94,687 257 94,559 3,337,039 35.24 44—45 0.00292 94,431 277 94,292 3,242,480 34.34 45—46 0.00318 94,154 299 94,005 3,148,188 33.44 46—47 0.00348 93,855 327 93,692 3,054,183 32.54 47—48 0.00380 93,528 355 93,350 2,960,491 31.65 48—49 0.00414 93,173 386 92,980 2,867,141 30.77 49—50 0.00449 92,787 417 92,579 2,774,161 29.90 50—51 0.00490 92,370 452 92,144 2,681,582 29.03 51—52 0.00537 91,918 494 91,671 2,589,438 28.17 52—53 0.00590 91,424 539 91,155 2,497,767 27.32 53—54 0.00647 90,885 588 90,591 2,406,612 26.48 54—55 0.00708 90.297 639 89.978 2,316,021 25.65 55—56 0.00773 89,658 693 89,311 2,226,043 24.83 82 56—57 0.00844 88,965 751 88,589 2,136,732 24.02 57—58 0.00926 88,214 817 87,806 2,048,143 23.22 58—59 0.01019 87,397 891 86,951 1,960,337 22.43 59—60 0.01120 86,506 679 86,021 1,873,386 21.66 60—61 0.01223 85,537 1,047 85,013 1,787,365 20.90 61—62 0.01328 84,490 1,122 83,930 1,702,352 20.15 62—63 0.01439 83,368 1,199 82,768 1,618,422 19.41 63—64 0.01560 82,169 1,282 81,527 1,535,654 18.69 64—65 0.01691 80,887 1,368 80,203 1,454,127 17.98 65—66 0.01827 79,519 1.453 78,793 1,373,924 17.28 66—67 0.01967 78,066 1,535 77,298 1,295,131 16.59 67—68 0.02121 76,531 1,624 75,719 1,217,833 15.91 69—69 0.02297 74,907 1,721 74,047 1,142,114 15.25 69—70 0.02499 73,186 1,829 72,272 1,068,067 14.59 70—71 0.02727 71,357 1,946 70,384 995,795 13.96 71—72 0.02979 69,411 2,067 68,377 925,411 13.33 72—73 0.03251 67,344 2,190 66,249 857,034 12.73 73—74 0.03534 65,154 2,302 64,003 790,785 12.14 74—75 0.03824 62,852 2,403 61,651 726,782 11.56 75—76 0.04126 60,449 2,494 59,201 665,131 11.00 76—77 0.04455 57,955 2,582 56,664 605,930 10.46 77—78 0.04819 55,373 2,669 54,039 549,266 9.92 78—79 0.05239 52,704 2,761 51,323 495,227 9.40 79—80 0.05723 49,943 2.859 48,514 443,904 8.89 80—81 0.06277 47,084 2,955 45,607 395,390 8.40 81—82 0.06885 44,129 3,038 42,609 349,783 7.93 82—83 0.07535 41,091 3,097 39,543 307,174 7.48 83—84 0.08207 37,994 3,118 36,435 267,631 7.04 84—85 0.08907 34,876 3,106 33,324 231,196 6.63 85—86 0.09705 31,770 3,083 30,228 197,872 6.23 (Continued overleaf ) 83 Table 4.4 Continued Proportion Of100,000 StationaryAverage Dying BornAlive PopulationRemainingLifetime Proportionof Average Age PersonsAlive Number Number InThis Numberof Interval atBeginningof Livingat Dying andAll YearsofLife AgeInterval Beginning During Inthe Subsequent Remainingat PeriodofLife DyingDuring ofAge Age Age Age BeginningofBetweenTwoAges Interval Interval Interval Interval Intervals AgeInterval xtox/p59t/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p15 86—87 0.10627 28,687 3,049 27,163 167,644 5.84 87—88 0.11625 25,638 2,980 24,148 140,481 5.48 88—89 0.12688 22,658 2,875 21,220 116,333 5.13 89—90 0.13834 19,783 2,737 18,415 95,113 4.81 90—91 0.15135 17,046 2,580 15,757 76,698 4.50 91—92 0.16591 14,466 2,400 13,266 60,941 4.21 84 92—93 0.18088 12,066 2,182 10,975 47,675 3.95 93—94 0.19552 9,884 1,933 8,918 36,700 3.71 94—95 0.21000 7,951 1,669 7,116 27,782 3.49 95—96 0.22502 6,282 1,414 5,575 20,666 3.29 96—97 0.24126 4,868 1,174 4,281 15,091 3.10 97—98 0.25689 3,694 949 3,219 10,810 2.93 98—99 0.27175 2,745 746 2,372 7,591 2.77 99—100 0.28751 1,999 575 1,711 5,219 2.61 100—101 0.30418 1,424 433 1,208 3,508 2.46 101—102 0.32182 991 319 832 2,300 2.32 102—103 0.34049 672 229 557 1,468 2.19 103—104 0.36024 443 159 364 911 2.05 104—105 0.38113 284 109 229 547 1.93 105—106 0.40324 175 70 140 318 1.81 106—107 0.42663 105 45 83 178 1.70 107—108 0.45137 60 27 46 95 1.59 108—109 0.47755 33 16 25 49 1.49 109—110 0.50525 17 8 13 24 1.39 Source: U.S. Decennial Life Tables for 1989—1991,V o l.1 ,No .1 , U.S. Life Tables , DHHSPublication PHS-98-1150-1, National Center for HealthStatistics, Washington,DC,1997. 85 Table 4.5 Abridged Life Table for the Total Population, United States, 1998 Stationary ProportionDying NumberLiving NumberDying Stationary PopulationinThis LifeExpectancy DuringAge atBeginningof DuringAge Populationinthe andAllSubsequent atBeginning Interval, AgeInterval, Interval, AgeInterval, AgeIntervals, ofAgeInterval, Age/p82q/p86l/p86/p82d/p86/p82L/p86T/p86e /p4/p86 0—1 0.00721 100,000 721 99,370 7,671,400 76.7 1—5 0.00139 99,279 138 396,786 7,572,030 76.3 5—10 0.00089 99,141 88 495,473 7,175,244 72.4 10—15 0.00110 99,053 109 495,057 6,679,771 67.4 15—20 0.00353 98,944 349 493,926 6,184,714 62.5 20—25 0.00476 98,595 469 491,820 5,690,788 57.7 25—30 0.00487 98,126 478 489,450 5,198,968 53.0 30—35 0.00600 97,648 586 486,840 4,709,518 48.2 35—40 0.00819 97,062 795 483,428 4,222,678 43.5 40—45 0.01176 96,267 1,132 478,670 3,739,250 38.8 45—50 0.01728 95,135 1,644 471,811 3,260,580 34.3 50—55 0.02564 93,491 2,397 461,839 2,788,769 29.8 55—60 0.04009 91,094 3,652 446,966 2,326,930 25.5 60—65 0.06302 87,442 5,511 424,280 1,879,964 21.5 65—70 0.09437 81,931 7,732 391,364 1,455,684 17.8 70—75 0.14239 74,199 10,565 345,660 1,064,320 14.3 75—80 0.20604 63,634 13,111 286,484 718,660 11.3 80—85 0.31641 50,523 15,986 213,526 432,176 8.6 85—90 0.46104 34,537 15,923 131,897 218,650 6.3 90—95 0.61502 18,614 11,448 62,020 86,753 4.7 95—100 0.75426 7,166 5,405 20,150 24,733 3.5 100/p59 1.00000 1,761 1,761 4,583 4,583 2.6 Source: U.S. Life Tables ,1998.NationalVitalStatisticsReports,Vol.48,No.18,NationalCenterforHealthStatistics,Washington,DC,2001. 86 tablefor1989 —1991thelifeexpectancyatbirthis75.37yearsandthatat age40is37.98years.Thismeansthataccordingtothemortalityratesof1989—1991newbornsareexpectedtolive75.37yearsandthoseatage40 are expected to live another 37.98 years. The life expectancy of apopulationisageneralindicationofthecapabilityofprolonginglife.Itisusedtoidentifytrendsandtocomparelongevity.Table4.5showsthataccordingtothemortalityratesof1998,thenewbornsandthoseatage40areexpectedtolive76.7and38.8years,respectively.Theoveralllife expectancy indicates an improvement in longevity in the United Statesoverthetimeperiod. Population life tables can be constructed for various subgroups. For example,therearepublishedlifetablesbygender,race,causeofdeath,aswellasthosewhicheliminatecertaincausesofdeath. 4.2.2 Clinical Life Tables The actuarial life table method has been applied to clinical data for many decades. Berkson and Gage (1950 )and Cutler and Ederer (1958 )give a life-table method for estimating the survivorship function; Gehan (1969 ) providesmethodsforestimatingallthreefunctions (survivorship,density,and hazard ). Thelife-tablemethodrequiresafairlylargenumberofobservations,sothat survival times can be grouped into intervals. Similar to the PL estimate, thelife-tablemethodincorporatesallsurvivalinformationaccumulateduptotheterminationofthe study.Forexample,incomputinga five-yearsurvivalrateof breast cancer patients, one need not restrict oneself only to those patientswhohaveenteredonstudyforfiveormoreyears.Patientswhohaveenteredfor four, three, two, and even one year contribute useful information to theevaluation of five-year survival. In this way, the life-table technique usesincomplete data such as losses to follow-up and persons withdrawn alive aswellascompletedeathdata. Table 4.6 shows the format of the clinical life table. The columns are describedbelow. 1.Interval[t/p71/p59t/p71/p62/p16). The first column gives the intervals intowhich the survival times and times to loss or withdrawal are distributed. Theinterval is from t/p71up to but notincluding t/p71/p62/p16,i/p581,...,s. The last intervalhasaninfinitelength.Theseintervalsareassumedtobefixed. 2.Midpoint (t/p75/p71).Themidpointofeachinterval,designated t/p75/p71,i/p581,..., s/p571,isincludedforconvenienceinplottingthehazardandprobability densityfunctions.Bothfunctionsareplottedas t/p75/p71. 3.Width (b/p71).Thewidth ofeach interval, b/p71/p58t/p71/p62/p16/p57t/p71,i/p581,...,,s/p571, isneededforcalculationofthehazardanddensityfunctions.Thewidth-  87 Table 4.6 Formatof a Life Table Number Number Number Number Conditional Conditional Cumulative Probability Lostto Withdrawn Number Entering Exposed Proportion Proportion Proportion Density Hazard Interval Midpoint Width Follow-up Alive Dying Interval toRisk Dying Surviving Surviving f(t/p75/p71)h/p19(t/p75/p71) t/p57t/p17t/p75/p16b/p16l/p16w/p16d/p16n/p30/p16n/p16q/p24/p16p/p24/p16S/p19(t/p16)/p581.00f/p19(t/p75/p16)h/p19(t/p75/p71) t/p17/p57t/p18t/p75/p17b/p17l/p17w/p17d/p17n/p30/p17n/p17q/p24/p17p/p24/p17S/p19(t/p17) f/p19(t/p75/p17)h/p19(t/p75/p17) /p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36/p36/p36/p36 /p36 t/p71/p58t/p71/p62/p16t/p75/p71b/p71l/p71w/p71d/p71n/p30/p71n/p16q/p24/p71p/p24/p71s/p19(t/p71) f/p19(t/p75/p71)h/p19(t/p75/p71) /p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36 /p36/p36/p36/p36 /p36 t/p81/p92/p16/p57t/p81t/p75/p11/p81/p92/p16b/p81/p92/p16l/p81/p92/p16w/p81/p92/p16d/p81/p92/p16n/p30/p81/p92/p16n/p81/p92/p16q/p24/p81/p92/p16p/p24/p81/p92/p16S/p19(t/p81/p92/p16)f/p19(t/p75/p11/p81/p92/p16)h/p19(t/p75/p11/p81/p92/p16) t/p81/p57/p45—— l/p81w/p81d/p81n/p30/p81n/p8110 S/p19(t/p81)—— 88 ofthelastinterval, b/p81,istheoreticallyinfinite;noestimateofthehazardor densityfunctioncanbeobtainedforthisinterval. 4.Number lost to follow-up (l/p71).Thisisthenumberofpeoplewhoarelost to observation and whose survival status is thus unknown in the ith interval (i/p581,...,s). 5. Number withdrawn alive (w/p71).Peoplewithdrawnaliveinthe ithinterval arethoseknowntobealiveattheclosingdateofthestudy.Thesurvival time recorded for such persons is the length of time from entrance totheclosingdateofthestudy. 6.Number dying (d/p71). This is the number of people who die in the ith interval.Thesurvivaltimeofthesepeopleisthetimefromentrancetodeath. 7.Number entering the i thinterval (n/p30/p71).Thenumberofpeopleenteringthe first interval n/p30/p16is the total sample size. Other entries are determined fromn/p30/p71/p58n/p30/p71/p92/p16/p57l/p71/p92/p16/p57w/p71/p92/p16/p57d/p71/p92/p16. That is, the number of persons enteringthe ithintervalisequaltothenumberstudiedatthebeginning of the preceding interval minus those who are lost to follow-up,withdrawnalive,orhavediedintheprecedinginterval. 8.Number exposed to risk (n/p71). This is the number of people who are exposedtoriskinthe ithintervalandisdefinedas n/p71/p58n/p30/p71/p57/p16/p17 (l/p71/p59w/p71). It is assumed that the times to loss or withdrawal are approximatelyuniformly distributed in the interval. Therefore, people lost or with-drawn in the interval are exposed to risk of death for one-half theinterval.Iftherearenolossesorwithdrawals, n/p71/p58n/p30/p71. 9.Conditional proportion dying (q/p24/p71). This is defined as q/p71/p58d/p71/n/p71for i/p581,...,s/p571, andq/p24/p81/p581. It is an estimate of the conditional probability of death in the ith interval given exposure to the risk of deathinthe ithinterval. 10.Conditional proportion surviving (q/p24/p71).Thisisgivenby p/p24/p71/p581/p57q/p24/p71,which is an estimate of the conditional probability of surviving in the ith interval. 11.Cumulative proportion surviving [S/p19(t/p71)]. This is an estimate of the survivorshipfunctionattime t/p71;itisoftenreferredtoasthe cumulative survival rate . Fori/p581,S/p19/p24(t/p16/p71)/p581 and for i/p582,...,s,S/p19(t/p71)/p58 p/p24/p71/p92/p16S/p19(t/p71/p92/p16). It is the usual life-table estimate and is based on the fact thatsurvivingtothestartofthe ithintervalmeanssurvivingtothestart ofandthenthroughthe (i/p571)th interval. 12.Estimated probability density function [f/p19(t/p75)]. This is defined as the probabilityofdying inthe ith intervalper unitwidth.Thus, anatural estimateatthemidpointoftheintervalis f/p19(t/p75)/p58S/p19(t/p71)/p57S/p19(t/p71/p92/p16) b/p71 /p58S/p19(t/p71)q/p24/p71b/p71i/p581,...,s/p571( 4 .2.7)-  89 13. Hazard function [h/p19(t/p75/p71)]. The hazard function for the ith interval, estimatedatthemidpoint,is h/p19(t/p75/p71)/p58d/p71b/p71(n/p71/p57/p16/p17d/p71)/p582q/p24/p71b/p71(1/p59p/p24/p71)i/p581,...,s/p571( 4.2.8) It is the number of deaths per unit time in the interval divided by the average number of survivors at the midpoint of the interval.That is, h/p19(t/p75/p71) is derived from f/p19(t/p75/p71)/S/p19(t/p75/p71) andS/p19(t/p75/p71)/p58/p16/p17[S/p19(t/p71/p62/p16) /p59S/p19(t/p71)] since S(t/p71) is defined as the probability of surviving at the beginning,notthemidpoint,ofthe ithinterval: h/p19(t/p75/p71)/p58f/p19(t/p75/p71) S/p19(t/p75/p71)/p58S/p19(t/p71)q/p24/p71/b/p71/p16/p17S/p19(t/p71)(p/p24/p71/p591)(4.2.9) whichreducesto (4.2.8 ). Sacher (1956 )derivesanestimateofthehazardfunctionbyassuming that hazard is constant within an interval but varies among intervals.Hisestimateis h/p19(t/p75/p71)/p58/p57logp/p24/p71b/p71 (4.2.10 ) InaMonteCarlostudy,GehanandSiddiqui (1973 )showthat (4.2.9)is lessbiasedthan (4.2.10 ). Thelarge-sampleapproximatevariancesoftheestimatedsurvivalfunctions, S/p19(t/p71),f/p19(t/p75/p71), andh/p19(t/p75/p71) intheithintervalare Var[S/p19(t/p71)]/p60[S/p19(t/p71)]/p17/p71/p92/p16/p26 /p72/p14/p16q/p24/p72n/p72p/p24/p72(4.2.11) Var[f/p19(t/p75/p71)]/p60[S/p19(t/p71)q/p24/p71]/p17 b/p71 /p1/p71/p92/p16/p26 /p72/p14/p16q/p24/p72n/p72p/p24/p72/p59p/p24/p71n/p72q/p24/p72/p2(4.2.12) and Var[h/p19(t/p75/p71)]/p60[h/p19(t/p75/p71)]/p17 n/p71q/p24/p71/p71/p57/p31 2h/p19(t/p75/p71)b/p71/p4/p17/p8(4.2.13) Equation (4.2.11 )isgivenbyGreenwood (1926 );Gehan (1969 )derived (4.2.12 ) and(4.2.13 ).Thesemaybeusedtoobtainapproximateconfidenceintervalsfor thevarioussurvivalfunctions. The graph of S/p19(t/p71) can be used to find an estimate of the median. Or let (t/p72,t/p72/p62/p16) be the interval such that S/p19(t/p72)/p460.5 andS/p19(t/p72/p62/p16)/p580.5. Then the90       mediansurvivaltime t/p75canbeestimatedbylinearinterpolation: t/p19/p75/p58t/p72/p59[S/p19(t/p72)/p570.5]b/p72S/p19(t/p72)/p57S/p19(t/p72/p62/p16)/p58t/p72/p59S/p19(t/p72)/p570.5 f/p19(t/p75/p72)(4.2.14) wheref/p19(t/p75/p72)isdefinedin (4.2.7 ). Anotherinterestingmeasurethatcanbeobtainedfromthelifetableisthe median remaining lifetime at timet/p72, denotedby t/p75/p80(i),i/p581,...,s/p571. If att/p71the proportion of individual survival is S/p19(t/p71), the proportion of individual survivalat t/p75/p80(i)i s/p16/p17S/p19(t/p71). Thatis,one-halfofthepeoplewhoarealiveattime t/p71areexpectedtobealiveattime t/p75/p80(i).Let (t/p72,t/p72/p62/p16) betheintervalinwhich /p16/p17S/p19(t/p71)falls;thatis, S/p19(t/p72)/p46/p16/p17S/p19(t/p71)andS/p19(t/p72/p62/p16)/p58/p16/p17S/p19(t/p71).Thenanestimateof t/p75/p80(i) is t/p19/p75/p80(i)/p58(t/p72/p57t/p71)/p59b/p72[S/p19(t/p72)/p57/p16/p17S/p19(t/p71)] S/p19(t/p72)/p57S/p19(t/p72/p62/p16)(4.2.15 ) HereS/p19(t/p72)istheestimatedproportionsurvivingbeyondthelowerlimitofthe intervalcontainingthemedian. Thevarianceof t/p75/p80(i) isapproximately Var[t/p19/p75/p80(i)]/p58[S/p19(t/p71)]/p17 4n/p71[f/p19(t/p75/p72)]/p17(4.2.16 ) Example 4.4 The following survival data for 2418 males with angina pectoris, originally reported by Parker et al. (1946 ), were also included in Gehan’s (1969 )paper. Survival time is computed from time of diagnosis in years.Thelifetableuses16intervalsofoneyear.Table4.7givesestimatesofthe various survival functions, the median remaining lifetime, and theirstandarderrors.Thesurvivorshipfunction, S/p19(t), isplottedat tandthehazard anddensityfunctions, h/p19(t) andf/p19(t), areplottedatthemidpointoftheinterval (Figure4.4 ). The graph of the estimated hazard function shows that the death rate is highestin the first year after diagnosis.From the end of the first year to thebeginning of the tenth year, the death rate remains relatively constant,fluctuatingbetween0.09and0.12.Thehazardrateisgenerallyhigherafterthetenth year. Hence, the prognosis for a patient who has survived one year isbetter than that for a newly diagnosed patient if factors such as age, gender,andracearenotconsidered.Asimilarinterpretationisreachedbyexaminingthe estimated median remaining lifetimes. Initially, the estimated medianremaininglifetimeis5.33years.Itreachesapeakof6.34yearsatthebeginningof the second year after diagnosis and then decreases. The median survivaltime,eitherreadfromthesurvivalcurveorusing (4.2.14 ),is5.33yearsandthe five-yearsurvivalrateis0.5193withastandarderrorof0.0103.-  91 Table 4.7 Life-Table Analysis of 2418 Males with Angina Pectoris Number Number Number Number Conditional Conditional Yearafter Lostto Withdrawn Number Entering Exposed Proportion Proportion Diagnosis Midpoint Width Follow-up Alive Dying Interval toRisk Dying Surviving S/p19(t/p71)f/p19(t/p75/p71)h/p19(t/p75/p71)/p40Var[S/p19(t/p71)]/p40Var[f/p19(t/p75/p71)]/p40Var[h/p19(t/p75/p71)]t/p19/p75/p80(i)/p40Var[t/p19/p75/p80(i)] 00 .51 .0 0 0 456 2418 2418 .00.1886 0 .8114 1 .0000 0.1886 0.2082 — 0 .0080 0 .0097 5 .33 0 .17 11 .51 .0 39 0 226 1962 1942 .50.1163 0 .8837 0 .8114 0.0944 0.1235 0 .0080 0 .0060 0 .0082 6 .35 0 .20 22 .51 .0 22 0 152 1697 1686 .00.0902 0 .9098 0 .7170 0.0646 0.0944 0 .0092 0 .0051 0 .0076 6 .34 0 .24 33 .51 .0 23 0 171 1523 1511 .50.1131 0 .8869 0 .6524 0.0738 0.1199 0 .0097 0 .0054 0 .0092 6 .23 0 .24 44 .51 .0 24 0 135 1329 1317 .00.1025 0 .8975 0 .5786 0.0593 0.1080 0 .0101 0 .0049 0 .0093 6 .22 0 .19 55 .51 .0 107 0 125 1170 1116 .50.1120 0 .8880 0 .5193 0.0581 0.1186 0 .0103 0 .0050 0 .0106 5 .91 0 .18 66 .51 .0 133 0 83 938 871 .50.0952 0 .9048 0 .4611 0.0439 0.1000 0 .0104 0 .0047 0 .0110 5 .60 0 .19 77 .51 .0 102 0 74 722 671 .00.1103 0 .8897 0 .4172 0.0460 0.1167 0 .0105 0 .0052 0 .0135 5 .17 0 .27 88 .51 .0 68 0 51 546 512 .00.0996 0 .9904 0 .3712 0.0370 0.1048 0 .0106 0 .0050 0 .0147 4 .94 0 .28 99 .51 .0 64 0 42 427 395 .00.1063 0 .8937 0 .3342 0.0355 0.1123 0 .0107 0 .0053 0 .0173 4 .83 0 .41 10 10 .51 .0 45 0 43 321 298 .50.1441 0 .8559 0 .2987 0.0430 0.1552 0 .0109 0 .0063 0 .0236 4 .69 0 .42 11 11 .51 .0 53 0 34 233 206 .50.1646 0 .8354 0 .2557 0.0421 0.1794 0 .0111 0 .0068 0 .0306 4 .00/p59— 12 12 .51 .0 33 0 18 146 129 .50.1390 0 .8610 0 .2136 0.0297 0.1494 0 .0114 0 .0067 0 .0351 3 .00/p59— 13 13 .51 .02 7 0 99 5 8 1 .50.1104 0 .8896 0 .1839 0.0203 0.1169 0 .0118 0 .0065 0 .0389 2 .00/p59— 14 14 .51 .02 3 0 65 9 4 7 .50.1263 0 .8737 0 .1636 0.0207 0.1348 0 .0123 0 .0080 0 .0549 1 .00/p59— 15 — — 0 0 0 30 30 .01.0000 0 .0000 0 .1429 — — 0 .0133 — — — — Source:Gehan (1969 ). 92 Figure 4.4Survivalfunctionsofmalepatientswithanginapectoris.-  93 Assume that survival time t(year)from each of 2418 males with angina pectorisinExample4.4hasthesameformatasthedatafile‘‘C: /p33D4d2.DAT’’ defined in Example 4.2 and is saved in ‘‘C: /p33D4d4.DAT’’. Then the following SAScodecanbeusedtoproduceaclinicallifetablesuchasTable4.7. dataw1; infile‘c: /p33d4d4.dat’missover; inputtcens; run;proclifetestdata /p58w1outsurv /p58wamethod /p58lifeintervals /p580to15by1; timet*cens (0); run;title‘Lifetableofthesurvivaltimes’; procprintdata /p58wa; run; IfBMDP1Lisused,therespectivecodeis /input file /p58‘c:/p33d4d4.dat’. variables /p582. format /p58free. /variable names /p58t,cens. /form unit /p58year. time /p58t. status /p58cens. response /p581. /estimate method /p58life. Print. /end IftheSPSSSURVIVALprocedureisused,therespectivecodeis datalistfile /p58‘c:/p33d4d4.dat’free /tcens. survivaltables /p58t /status /p58cens (1)fort /intervals /p58thru15by1 /print. 4.3 RELATIVE, FIVE-YEAR, AND CORRECTED SURVIVAL RATES Anotherapproachtolarge-scalesurvivaldataisthecalculationofthe relative survival rate or annual survival ratio. The relative survival rate evaluates the survivalexperienceofpatientsintermsofthegeneralpopulation.Greenwood(1926 )first suggested this approach for evaluating the efficacy of cancer treatment:Iftheaveragesurvivaltimeofthepatientstreatedequalsthatofa94       randomsampleofpersonsofthesameage,gender,occupation,andsoon,the patientscouldbeconsidered‘‘cured.’’Cutleretal. (1957,1959,1960 a,b,1967 ) adopted Greenwood’s idea of comparing the survival experience of cancerpatients with that of the general population to ascertain (1)the ratio of observedtoexpectedsurvivalratesand (2)whether,intime,themortalityrate declinestoa‘‘normal’’level. The relative survival rate is defined as the ratio of the survival rate (probabilityofsurvivingoneyear )forapatientunderstudy (observed rate )to someoneinthegeneralpopulationofthesameage,gender,andrace (expected rate)overaspecifiedperiodoftime.Toprovideamoreprecisemeasureofthe relationshipof the observedand expectedsurvivalrates,Cutler et al. suggestcomputingtheratioforeachindividualfollow-upyear.Arelativerateof100%meansthatduring a specific follow-upyear the mortalityratesin the patientand in the general population are equal. A relative rate of less than 100%meansthatthemortalityrateinthepatientsishigherthanthatinthegeneralpopulation.Cutleretal.usethesurvivalratesintheConnecticutandU.S.lifetablesforthegeneralpopulation. UsingthenotationsinTable4.6,thesurvivalrateobservedattime t/p71isp/p24/p71, theexpectedsurvivalratecanbecomputedasfollows:Supposethatattime t/p71there are n/p30/p71individuals alive for whom age, gender, race, and time of observationareknown.Let p*/p71/p72bethesurvivalrateofthe jthindividualfrom generalpopulationlifetables (withcorrespondingage,gender,andrace ).The expectedsurvivalrateis p*/p71/p581 n/p30/p71 /p76/p89/p71/p26 /p72/p14/p16p*/p71/p72(4.3.1 ) Thentherelativesurvivalrateattime t/p71isdefinedby r/p71/p58p/p24/p71p*/p71(4.3.2) Example4.5taken from Cutler et al. (1957)illustrates the interpretation of relative survival rates. Example 4.5 A total of 9121 breast cancer cases were diagnosed in Connecticuthospitalsfrom1935to1953.TheConnecticutlifetableforwhitefemales,1939 —1941,isusedincalculationoftheexpectedsurvivalrate.Table 4.8 gives the observed and expected survival rates as well as the relativesurvivalrates.Figure4.5 agraphicallyshowsthesedata:thesurvivalcurvesfor the breast cancer patients and the general population. The relative survivalratesareplottedinFigure4.5 b.Forthisgroupofpatients,therelativesurvival rates, although increasing during 13 successive years, are less than 100%throughout the 15 years of follow-up. During each of the 15 years, the,-,    95 Table 4.8 Relative Survival Rates of Breast Cancer Patients in Connecticut, 1935--1953 SurvivalRates (%)Relative Yearsafter SurvivalRate Diagnosis Observed Expected (%) 0—1 82.9 97.2 85 1—2 83.3 97.1 86 2—3 85.9 96.9 89 3—4 86.8 96.7 90 4—5 89.2 96.6 92 5—6 90.0 96.4 93 6—7 89.9 96.4 93 7—8 91.6 96.2 95 8—9 92.0 96.1 96 9—10 92.7 96.1 96 10—11 92.9 95.9 97 11—12 94.0 95.8 98 12—13 94.1 95.3 99 13—14 91.5 95.3 96 14—15 90.6 94.9 95 Source:Cutleretal. (1957 ). breast cancer patient mortality rate is greater than that of the general population. Othermeasuresofdescribingsurvivalexperienceofcancerpatientsarethe five-year survival rate and the corrected rate. The five-year survival rate is simply the cumulative proportion surviving at the end of the fifth year. Forexample, the five-year survival rate for the males with angina pectoris inExample 4.4 is 0.5193. The five-year survival rate is no longer a measure oftreatmentsuccessforpatientswithmanytypesofcancersincethesurvivalofcancerpatientshasimprovedconsiderablyinthelastfewdecades. Berkson (1942 )suggestsusinga corrected survival rate. Thisisthesurvival rate if the disease under study alone is the cause of death. In most survivalstudies, the proportion of patients surviving is usually determined withoutconsideringthecauseofdeath,whichmightbeunrelatedtothespecificillness.Ifp/p65denotesthesurvivalratewhencanceraloneisthecauseofdeath,Berkson proposesthat p/p65/p58p p/p15 (4.3.3 ) wherepistheobservedtotalsurvivalrateinagroupofcancerpatientsand p/p15is the survival rate for a group of the same age and gender in the general96       Figure 4.5SurvivalratesofbreastcancerpatientsinConnecticut,1935 —1953. population. Rate p/p65may be computed at any time after the initiation of follow-up;it provides a measure of the proportionof patientsthat escaped adeathfromcancerupto thatpoint.Ifa five-yearsurvivalrateis0.5anditiscorrectedfornoncancerdeathsandifwefindthatfive-yearsurvivalrateofthegeneralpopulationis0.9,thecorrectedsurvivalrateis0.5/0.9,or0.56. 4.4 STANDARDIZED RATES AND RATIOS Ratesandratiosareoftenusedindemographyandepidemiologyto describe the occurrence of a health-related event. For example, the standardized mortality (or morbidity )ratio (SMR )is frequently used in occupational epidemiology as a measure of risk, and the standardized death rate iscommonlyusedincomparingmortalityexperiencesofdifferentpopulationsorthesamepopulationatdifferenttimes. TheconceptoftheSMRisverysimilartothatoftherelativesurvivalrate described above. It is defined as the ratio of the observed and the expectednumberofdeathandcanbeexpressedas SMR /p58observed number of deaths in study population expected number of deaths in study population /p59100 (4.4.1 ) wherethe expectednumberof deathsisthe sumof the expecteddeathsfrom the same age, gender, and race groups in the general population. Thestandardized morbidity ratio can similarly be calculated simply by replacingtheword deathsbydisease cases in(4.4.1 ).Ifonlynewcasesareofinterest,we calltheratiothe standardized incidence ratio (SIR).    97 Table 4.9 Population and Deaths of Sunny City and Happy City by Age SunnyCity HappyCity Age-Specific Age-Specific Rates Rates Age Population Deaths (per1000 )Population Deaths (per1000 ) /p5825 25,000 25 1.00 55,000 110 2.0 25—44 40,000 50 1.25 20,000 50 2.5 45—64 20,000 200 10.00 21,000 315 15.0 /p4665 15,000 1,200 80.00 4,000 650 162.5 Total 100,000 1,475 100,000 1,125 Thestandardizeddeathrateisonlyoneofthemanyratesusedtodescribe thehealth status of a populationorto comparethe healthstatus of differentpopulations. If the populations are similar with respect to demographicvariablessuchasage,gender,orrace,the crude rate,orratioofthenumberof persons to whom the event under study occurred to the total number ofpersonsinthepopulation,cansafelybeusedforcomparison. Thelevelofthecruderateisaffectedbydemographiccharacteristicsofthe population for which the rate is computed. If populations have differentdemographiccompositions,acomparison ofthe cruderates may be mislead-ing.Asanexampleconsiderthetwohypotheticalpopulations,SunnyCityandHappy City, in Table 4.9. The crude death rate of Sunny City is 1000 (1475/ 100,000 )or14.7 per 1000.Thecrude death rate of HappyCity is1000 (1125/ 100,000 ),or11.25per1000,whichislowerthanthatofSunnyCityeventhough allage-specificratesinHappyCityarehigher.Thisismainlybecausethereisa large proportion of older people in Sunny City. A crude death rate of apopulation may be relatively high merely because the population has a highproportionofolderpeople;itmayberelativelylowbecausethepopulationhasa high proportion of younger people. Thus, one should adjust the rate toeliminate the effects of age, gender, or other differences. The procedure ofadjustmentiscalled standardization andtherateobtainedafterstandardization iscalledthe standardized rate . Themostfrequentlyusedmethodsforstandardizationarethedirectmethod andtheindirectmethod. Direct Method Inthis method a standardpopulationis selected. Thedistributionacross thegroups with different values of the demographic characteristic (e.g., different age groups )must be known. Let r/p16,...,r/p73, wherekis the number of groups, bethespecificratesofthedifferentgroupsforthepopulationunderstudy.Letp/p16,...,p/p73be the proportions of people in the kgroups for the standard population.Thedirectstandardizedrateisobtainedbymultiplyingthespecific98       ratesr/p71byp/p71ineachgroup.Theformulaforthedirectstandardizedrateis R/p3/p9/p18/p4/p2/p20/p58/p73/p26 /p71/p14/p16r/p71p/p71(4.3.2) As an example, consider the data in Table 4.9. If we choose a standard population whose distribution is shown in the second column of Table 4.10, thedirectstandardizeddeathrateforSunnyCityandHappyCityis,respect-ively,9.37and17.84per1000.Thesestandardizedratesaremorereliablethanthecruderatesforcomparisonpurposes. Indirect Method Ifthespecificrates r/p71ofthepopulationbeingstudiedareunknown,thedirect methodcannotbeapplied.Inthiscase,itispossibletostandardizetheratebyanindirectmethodifthefollowingareavailable: 1. Thenumberofpersonstowhomtheeventbeingstudiedoccurred (D)in thepopulation.Forexample,ifthedeathrateisbeingstandardized, Dis thenumberofdeaths. 2. The distribution across the various groups for the population being studied,denotedby n/p16,...,n/p73. 3. The specific rates of the selected standard population, denoted by s/p16,...,s/p73. 4. Thecruderateofthestandardpopulation,denotedby r. Theformulaforindirectstandardizationis R/p9/p14/p3/p9/p18/p4/p2/p20/p58D /p26/p73/p71/p14/p16n/p71s/p71 r (4.3.3) Thesummationin (4.3.3 )istheexpectednumberofpersonstowhomtheevent occurredonthebasisofthespecificratesofthestandardpopulation.Thus,theindirectmethodadjuststhecruderateofthestandardpopulationbytheratiooftheobservedtoexpectednumberofpersonstowhomtheeventoccurredinthepopulationunderstudy. Table 4.11 represents an example for the death rate in the states of OklahomaandArizonain1960 (dataarefromGroveandHetzel,1963 ).The U.S.populationin 1960is used as the standardpopulation.Thecrudedeathrate of Oklahoma (9.7 per thousand )is higher than that of Arizona (7.8 per thousand ). However, the indirect standardized rates show a reverse relation- ship (8.6 for Oklahoma and 9.6 for Arizona ). This, again, is because of the differencesinagedistribution.Thereisahigherproportionofpeoplebelowtheageof25inArizonaandahigherproportionofpeopleabovetheageof54inOklahoma.    99 Table 4.10 Standardized Death Rates by Direct Method for Sunny City and Happy City SunnyCity HappyCity Age-Specific Age-Standardized Age-Specific Age-Standardized Standard Proportion, DeathRates, DeathRates, DeathRates, DeathRates, Age Population p/p71r/p71p/p71r/p71r/p71p/p71r/p71 /p5825 420,000 0 .42 1 .00 0 .42 2 .00 .84 25—44 280,000 0.28 1.25 0.35 2.5 0.70 45—64 220,000 0.22 10.00 2.20 15.0 3.30 /p4665 80,000 0.08 80.00 6.40 162.5 13.00 Total 1,000,000 9.37 17.84 (R/p3/p9/p18/p4/p2/p20)( R/p3/p9/p18/p4/p2/p20) 100 Table 4.11 Standardized Death Rates by Indirect Method for Oklahoma and Arizona, 1960 Oklahoma Arizona StandardPopulation (U.S.Population,1960 ) Expected Expected Age-SpecificDeathRates, Population, Deaths, Population, Deaths, Age s/p71n/p71n/p71s/p71n/p71n/p71s/p71 /p5810 .0270 49,103 1,325 .78 34,599 934 .17 1—4 0.0011 193,644 213.01 132,367 145.60 5—14 0.0005 454,972 227.49 285,830 142.92 15—24 0.0011 329,230 362.15 186,789 205.47 25—34 0.0015 279,327 418.99 169,873 254.81 35—44 0.0030 287,994 863.98 173,029 519.09 45—54 0.0076 269,147 2,045.52 136,573 1,037.95 55—64 0.0174 216,036 3,759.03 92,871 1,615.96 65—74 0.0382 157,385 6,012.11 63,634 2,430.82 75—84 0.0875 74,848 6,549.20 22,499 1,968.66 85/p59 0.1986 16,598 3,296.36 4,092 812.67 Total 2,328,284 25,074 1,302,161 10,068 Cruderates 9.5 9.7 7.8 (perthousand ) Observeddeaths 22,584 10,157Expecteddeaths /p63 25,074 10,068 Standardizedrate /p122,584 25,074 /p29.5/p588.6 /p110,157 10,068 /p29.5/p589.6 (perthousand ) Source:DatafromGroveandHetzel (1963 ). /p63/afii9814n/p71s/p71. 101 Resultsfor the adjustedratesdependon thestandardpopulationselected. Hence,thisselectionshouldbedonecarefully.Whendiscussingdeathratebyage,Shryocketal. (1971 )suggestthatapopulationwithsimilaragedistribu- tion to the various populationsunder study be selected as a standard. If thedeathrateoftwopopulationsisbeingcompared,itisbesttousetheaverageofthetwodistributionsasastandard. Itshouldberememberedthatspecificratesarestillthemostaccurateand essential indicators of the variations among populations. No matter which methodis used, standardizedrates aremeaningful only whencompared withsimilarlycomputedrates.Kitagawa (1964 )alsocriticizesthestandardizedrate becauseifthespecificratesvaryindifferentwaysbetweenthetwopopulationsbeing compared, standardization will not indicate the differences and some-timeswill evenmaskthe differences.Nevertheless,ifthespecificratesarenotavailable, if a single rate for a population is desired, or if the demographiccomposition of the population being compared is different, the standardizedrateisuseful. Bibliographical Remarks KaplanandMeier’s (1958 )PLmethodisthemostcommonlyusedtechnique forestimatingthesurvivorshipfunctionforsamplesofsmallandmoderatesize.However,withtheaid ofacomputer,it isnotdifficult to usethe methodforlargesamplesizes. Berkson (1942 ), Berkson and Gage (1950 ), Cutler and Ederer (1958 ), and Gehan (1969 )have written classic reports on life-table analysis. Peto et al. (1976 )published an excellent reviewof some statistical methods related to clinicaltrials.Theterm life-tableanalysis thattheyuseincludesthePLmethod. Otherreferencesonlifetablesare,forexample,Armitage (1971 ),Shryocketal. (1971 ), Kuzma (1967 ), Chiang (1968 ), Gross and Clark (1975 ), and Elandt- JohnsonandJohnson (1980 ). RelativesurvivalratesandcorrectedsurvivalrateshavebeenusedbyCutler andco-workersinaseriesofsurvivalstudiesoncancerpatientsinConnecticutinthe1950sand1960s (Cutleretal.,1957,1959,1960 a,b,1967;Edereretal., 1961 ).DiscussionsofSMR,standardizedrates,andrelatedtopicscanbefound inmanystandardepidemiologytextbooks:forexample,MausnerandKramer(1985 ),Kahn (1983 ),Kelseyetal. (1986 ),Shryocketal. (1971 ),Chiang (1961 ), andMantelandStark (1968 ). EXERCISES 4.1Considerthesurvivaltimeofthe30melanomapatientsinTable3.1. (a)Compute and plot the PL estimates of the survivorship functions S/p19(t) ofthetwotreatmentgroupsandcheckyourresultswithTable 3.2andFigure3.1.102       Exercise Table 4.1 Number Timefrom NumberLost Withdrawn Number Number Diagnosis toFollow-up, Alive, Dying, Entering, (yr) l/p71w/p71d/p71n/p30/p71 0—5 18 0 731 949 5—10 16 0 52 200 10—15 8 67 14 132 15—20 0 33 10 43(b)Computethevarianceof S/p19(t) foreveryuncensoredobservation. (c)Estimatethemediansurvivaltimesofthetwogroups. 4.2Do the same as in Exercise 4.1 for the remission durations of the two treatmentgroupsinTable3.1. 4.3ComputeandplotthePLestimatesofthetumor-freetimedistributions for the saturated fat and unsaturated fat diet groups in Table 3.4.CompareyourresultswithFigure3.4. 4.4Consider the remission data of 42 patients with acute leukemia in Example3.3. (a)ComputeandplotthePLestimatesof S(t) ateverytimetorelapse forthe6-MPandplacebogroups. (b)Compute the variances of S/p19(10) in the 6-MP group and of S/p19(3) in theplacebogroup. (c)Estimatethemedianremissiontimesofthetwotreatmentgroups. 4.5 (a)ComputethesurvivaltimeforeachpatientinExerciseTable3.1. (b)Estimate and plot the overall survivorship function using the PL method.Whatisthemediansurvivaltime? (c)Divide the patients into two groups by gender. Compute and plot thePLestimatesofthesurvivorshipfunctionsforeachgroup.Whatisthemediansurvivaltimeforeach? 4.6ConsidertheskintestresultsinExerciseTable3.1.Foreachofthefive skintests: (a)Divide patients into two groups according to whether they had a positivereaction.Measurementslessthan10 /p5910(5/p595formumps ) areconsiderednegative. (b)Estimateandplotthesurvivorshipfunctionsofthetwogroups. (c)Can you tell from the plots if any skin tests might predict survival time? 4.7Consider the data of patients with cancer of the ovary diagnosed in Connecticutfrom1935to1944 (Cutleretal.1960b ).ExerciseTable4.1 103 Exercise Table 4.2 Survival Data of Female Patients with Angina Pectoris YearAfter NumberEntering NumberLostto Diagnosis Interval Follow-up NumberDying 0—1 555 0 82 1—2 473 8 30 2—3 435 8 27 3—4 400 7 22 4—5 371 7 26 5—6 338 28 25 6—7 285 31 20 7—8 234 32 11 8—9 191 24 14 9—10 153 27 13 10—11 113 22 5 11—12 86 23 5 12—13 58 18 5 13—14 35 9 2 14—15 24 7 3 15/p59 14 11 3 Source:R. L. Parker et al., JAMA,131(2),9 5—100 (1946 ). Copyright 1946. American Medical Association.reproducesthe data in life-table format.Provide a life-table like Table 4.5.Whatdoyoufindout? 4.8Doacompletelife-tableanalysisforthetwosetsofdatagiveninTable 3.5.Plotthethreesurvivalfunctions. 4.9Doacompletelife-tableanalysisofthedatagiveninExerciseTable4.2. Plotthethreesurvivalfunctions. 4.10ConsiderthesurvivaltimesofthemelanomapatientsinExerciseTable 3.4.Doacompletelife-tableanalysisofthesurvivaltime.Plotthethreesurvivalfunctions. 4.11Consider the data given in Exercise Table 4.3. Compute the direct standardizeddeathrateforthestatesofOklahomaandMontanausingtheU.S.populationof1960asthestandard. 4.12GiventhepopulationofJapanandChile (ExerciseTable4.4 ),compute theindirectstandardizeddeathrateforthetwocountriesusingtheU.S.deathrateof1960inTable4.11asthestandard.104       Exercise Table 4.3 OklahomaAverage MontanaAverage DeathRate DeathRate U.S.Population, Proportion, (per1000 )(per1000 ) Age 1960 (thousands ) p/p71r/p71r/p71 /p581 4,112 0 .023 25 .52 5 .8 1—4 16,209 0.091 1.2 1.2 5—14 35,465 0.198 0.5 0.5 15—24 24,020 0.134 1.2 1.6 25—34 22,818 0.127 1.6 1.8 35—44 24,081 0.134 2.9 3.1 45—54 20,486 0.114 6.9 7.5 55—64 15,572 0.087 14.8 16.3 65—74 10,997 0.061 32.4 37.3 75—84 4,634 0.026 79.0 87.3 85/p59 929 0.005 190.4 202.8 Total 179,323 1.000 Source:GroveandHetzel (1963 ). Exercise Table 4.4 Population (thousands ) Age Japan Chile /p581 1,577 228 1—4 6,268 876 5—14 20,223 1,817 15—24 17,627 1,323 25—34 15,727 1,034 35—44 11,057 779 45—54 9,018 603 55—64 6,573 395 65—74 3,724 212 75—84 1,438 83 /p4585 188 22———————Total 93,419 7,374 Observeddeaths 706,599 95,486 Source:Shryocketal. (1971 ). 105 CHAPTER 5 Nonparametric Methods for Comparing Survival Distributions The problem of comparing survival distributions arises often in biomedicalresearch.A laboratory researcher may want to compare the tumor-free timesof two or more groups of rats exposed to carcinogens.A diabetologist maywish to compare the retinopathy-free times of two groups of diabetic patients.A clinical oncologist may be interested in comparing the ability of two or moretreatments to prolong life or maintain health.Almost invariably, the disease-free or survival times of the different groups vary.These differences can beillustrated by drawing graphs of the estimated survivorship functions, but thatgives only a rough idea of the difference between the distributions.It does notreveal whether the differences are significant or merely chance variations.Astatistical test is necessary. In Section 5.1 we introduce five nonparametric tests that can be used for data with and without censored observations.Section 5. 2 is devoted to theMantel—Haenszel test, which is particularly useful in stratified analysis, a method commonly used to take account of possible confounding variables.InSection 5.3 we discuss the problem of comparing three or more survivaldistributions with or without censoring. 5.1 COMPARISON OF TWO SURVIVAL DISTRIBUTIONS Suppose that there are n/p16andn/p17patients who receive treatments 1 and 2, respectively.Let x/p16,...,x/p80/p129be ther/p16failure observations and x/p62/p80/p129/p62/p16,...,x/p62/p76/p129the n/p16/p57r/p16censored observations in group 1.In group 2, let y/p16,...,y/p80/p130be ther/p17failure observations and y/p62/p80/p130/p62/p16,...,y/p62/p76/p130then/p17/p57r/p17censored observations.That is, at the end of the study n/p16/p57r/p16patients who received treatment 1 and n/p17/p57r/p17patients who received treatment 2 are still alive.Suppose that the observations in group 1 are samples from a distribution with survivorshipfunction S/p16(t) and the observations in group 2 are samples from a distribution 106 with survivorship function S/p17(t).Then null hypothesis to consider is H/p15:S/p16(t)/p58S/p17(t) (treatments 1 and 2 are equally effective ) against the alternative H/p16:S/p16(t)/p57S/p17(t) (treatment 1 more effective than 2 ) or H/p17:S/p16(t)/p58S/p17(t) (treatment 2 more effective than 1 ) or H/p18:S/p16(t)/p34S/p17(t) (treatments 1 and 2 not equally effective ) When there are no censored observations, standard nonparametric tests can be used to compare two survival distributions.For example, the Wilcoxon(1945 )test or the Mann —Whitney (1947 )U-test can test the equality of two independent populations, and the sign test can be used for paired (or depend- ent)samples (Marascuilo and McSweeney, 1977 ).In the following we introduce five nonparametric tests: Gehan’s generalized Wilcoxon test (Gehan, 1965 a,b), the Cox—Mantel test (Cox 1959, 1972; Mantel, 1966 ), the logrank test (Peto and Peto, 1972 ), Peto and Peto’s generalized Wilcoxon test (1972 ), and Cox’s F-test (1964 ).All the tests are designed to handle censored data; data without censored observations can be considered a special case. 5.1.1 Gehan’s Generalized Wilcoxon Test In Gehan’s generalized Wilcoxon test every observation x/p71orx/p62/p71in group 1 is compared with every observation y/p72ory/p62/p72in group 2 and a score U/p71/p72is given to the result of every comparison.For the purpose of illustration, let us assumethat the alternative hypothesis is H/p16:S/p16(t)/p57S/p17(t), that is, treatment 1 is more effective than treatment 2. Define U/p71/p72/p58 /p7/p591i fx/p71/p57y/p72orx/p62/p71/p46y/p72 0i fx/p71/p58y/p72orx/p62/p71/p58y/p72ory/p62/p72/p58x/p71or (x/p62/p71,y/p62/p72) /p571i fx/p71/p58y/p72orx/p71/p45y/p62/p72 and calculate the test statistic W/p58/p76/p129/p26 /p71/p14/p16/p76/p130/p26 /p72/p14/p16U/p71/p72(5.1.1 ) where the sum is over all n/p16n/p16comparisons.Hence, there is a contribution to     107 the test statistic Wfor every comparison where both observations are failures (except for ties )and for every comparison where a censored observation is equal to or larger than a failure.The calculation of Wis laborious when n/p16andn/p17are large.Mantel (1967 )shows that it can be calculated in an alternative way by assigning a score to each observation based on its relative ranking.InGehan’s computation each observation in sample 1 is compared with each insample 2.If the two samples are combined into a single pooled sample ofn/p16/p59n/p17observations, it is the same as comparing each observation with the remaining n/p16/p59n/p17/p571.LetU/p71,i/p581,...,n/p16/p59n/p17, be the number of remaining n/p16/p59n/p17/p571 observations that the ith is definitely greater than minus the number that it is definitely less than.The n/p16/p59n/p17U/p71’s define a finite population with mean 0 and it is true that Gehan’s W/p58/p76/p129/p26 /p71/p14/p16U/p71(5.1.2 ) where summation is over the U/p71of sample 1 only.From either (5.1.1 )or(5.1.2 ), it is clear that Wwould be a large positive number if H/p16is true.Mantel also suggests that the permutational variance of Wbe used instead of the more complicated variance formula derived by Gehan.The permutational distribu-tion ofWcan be obtained by considering all /p16 /p1n/p16/p59n/p17n/p17/p2/p58(n/p16/p59n/p17)! n/p16!n/p17! ways of selecting n/p16of theU/p71at random.The test statistic WunderH/p15can be considered approximately normally distributed with mean 0 and variance /p17 Var(W)/p58n/p16n/p17/p76/p129/p62/p76/p130/p26 /p71/p14/p16U/p17/p71 (n/p16/p59n/p17)(n/p16/p59n/p17/p571)(5.1.3 ) SinceWis discrete, an appropriate continuity correction of 1 is ordinarily used when there are neither ties nor censored observations.Otherwise, a continuitycorrection of 0.5 would probably be appropriate. SinceWhas an asymptotically normal distribution with mean zero and variance in (5.1.3 ),Z/p58W//p40Var(W) has standard normal distribution.The rejection regions are Z/p57Z/p63forH/p16, andZ/p58/p57Z/p63forH/p17, and /p34Z/p34/p57Z/p63/p30/p17for H/p18whereP(Z/p57Z/p63/p34H/p15)/p58/afii9825. /p16n! is read n factorial: n !/p58n(n/p571)(n/p572)/p373.2.1. /p17This is called the permutational variance because it is obtained by considering the per mutational distribution of all ( n/p16/p59n/p17)!/n/p16!n/p17!W’s108       The number U/p71can be computed in two stages.For each observation, the first stage yields, unity plus the number of remaining observations that it isdefinitely larger than, that is, R/p16/p71.The second stage yields R/p17/p71, which is unity plus the number of remaining observations that the particular observation isdefinitely less than.Then U/p71/p58R/p16/p71/p57R/p17/p71.The computations of R/p16/p71andR/p17/p71can be accomplished systematically in steps, as illustrated in the following hypo-thetical example. Example 5.1 Ten female patients with breast cancer are randomized to receive either CMF (cyclic administration of cyclophosphamide, methatrexate, and fluorouracil )or no treatment after a radical mastectomy.At the end of two years, the following times to relapse (or remission times )in months are recorded: CMF (group 1 ): 23, 16 /p59,1 8/p59,2 0/p59,2 4/p59 Control (group 2 ): 15, 18, 19, 19, 20 The null hypothesis and the alternatives are H/p15:S/p16/p58S/p17(the two treatments are equally effective ) H/p16:S/p16/p57S/p17(CMF more efficient than no treatment ) The computations of R/p16/p71,R/p17/p71, andU/p71are given in Table 5.1. Thus, W/p581/p592/p595/p594/p596/p5818, Var (W)/p58(5)(5)(208) /[(10)(9)] /p5857.78, and Z/p5818//p4057.78 /p582.368.Suppose that the significance level used is /afii9825/p580.05, Z/p15/p13/p15/p20/p581.64; then the Zvalue computed is in the rejection region.Therefore, we reject H/p15at 0.05 level and conclude that the data show that CMF is more effective than no treatment.In fact, the approximate pvalue corresponding to Z/p582.368 is 0.009. Note that the sum of all n/p16/p59n/p17U/p71’s equals zero.This fact can be used to check the computation. 5.1.2 Cox--Mantel Test Lett/p7/p16/p8/p58···/p58t/p7/p73/p8be the distinct failure times in the two groups together and m/p7/p71/p8be the number of failure times equal to t/p71, or the multiplicity of t/p71, so that /p73/p26 /p71/p14/p16m/p7/p71/p8/p58r/p16/p59r/p17(5.1.4) Further, let R(t) be the set of people still exposed to risk of failure at time t, whose failure or censoring times are at least t.HereR(t) is called the riskset at timet.Letn/p16/p82andn/p17/p82be the number of patients in R(t) that belong to     109 Table 5.1 Mantel’s Procedure of Calculating Uifor Gehan’s Generalized Wilcoxon Test Observations of Two Samples in Ascending Order 15 16 /p62 18 18 /p62 19 19 20 20 /p62 23 24 /p62 Computation of R/p16/p71 Step 1.Rank from left to right, omitting censoredobservations 1 2 3 4 5 6 Step 2.Assign next-higher rank to censoredobservations 2 3 6 7 Step 3.Reduce the rank of tied observationsto the lower rankfor the value 3 Step 4.R/p16/p7112 2 3 3 3 5 6 6 7 Computation ofR/p17/p71 Step 5.Rank from right to left 10 9 8 7 6 5 4 3 2 1 Step 6.Reduce the rank of tied observations to the lowest rank for the value 5 Step 7.Reduce the rank of censoredobservations to 1 1 1 1 1 Step 8.R/p17/p711 0 18 15 5 4 1 2 1 U/p71/p58R/p16/p71/p57R/p17/p71/p5791 /p63 /p5762 /p63 /p572/p5721 5 /p63 4/p63 6/p63 /p63From group 1. treatment groups 1 and 2, respectively.The total number of observations, failure or censored in R(t/p7/p71/p8), isr/p7/p71/p8/p58n/p16/p82/p59n/p17/p82.Define U/p58r/p17/p57/p73/p26 /p71/p14/p16m/p7/p71/p8A/p7/p71/p8(5.1.5) I/p58/p73/p26 /p71/p14/p16m/p7/p71/p8(r/p7/p71/p8/p57m/p7/p71/p8) r/p7/p71/p8/p571A/p7/p71/p8(1/p57A/p7/p71/p8) (5.1.6 ) wherer/p7/p71/p8is the number of observations, failure or censored, in R(t/p7/p71/p8) andA/p7/p71/p8110       Table 5.2 Computations of Cox--Mantel Test Number in Risk Set of: Distinct Sample 1 Sample 2 Failure Time, t/p71m/p7/p71/p8n/p16/p82n/p17/p82r/p7/p71/p8A/p7/p71/p8 15 1 5 5 10 0 .5 18 1 4 4 8 0 .5 19 2 3 3 6 0 .5 20 1 3 1 4 0 .25 23 1 2 0 2 0is the proportion of r/p7/p71/p8that belong to group 2.An asymptotic two-sample test is thus obtained by treating the statistic C/p58U//p40Ias a standard normal variate under the null hypothesis (Cox, 1972 ).The following example illustrates the procedure. Example 5.2 Consider the remission data and the hypotheses in Example 5.1. There are k/p585 distinct failure times in the two groups, r/p16/p581 andr/p17/p585. To perform the Cox —Mantel test, Table 5.2 is prepared for convenience: U/p585/p57(0.5/p590.5/p592/p590.5/p590.25) /p585/p572.25 /p582.75 I/p581/p599 9(0.5/p590.5)/p591/p597 7(0.5/p590.5)/p592/p594 5(0.5/p590.5)/p591/p593 3(0.25/p590.75) /p580.25 /p590.25 /p590.4/p590.1875 /p581.0875 Therefore, C/p582.75//p401.0875 /p582.637/p57Z/p15/p13/p15/p20/p581.64 and we reject H/p15at 0.05 level and reach the same conclusion as in Example 5.1. The pvalue correspond- ing toZ/p582.637 is approximately 0.004. 5.1.3 Logrank Test Mantel’s (1966 )generalization of the Savage (1956 )test, often referred to as the logranktest (Peto and Peto, 1972 ), is based on a set of scores w/p71assigned to the observations.The scores are functions of the logarithm of the survival     111 function.Altshuler (1970 )estimates the log survival function at t/p7/p71/p8using /p57e(t/p7/p71/p8)/p58/p57 /p26 j/p45t/p7/p71/p8m/p7/p72/p8r/p7/p72/p8(5.1.7) wherem/p7/p72/p8andr/p7/p72/p8are as defined in Section 5.1.2. The scores suggested by Peto and Peto are w/p71/p581/p57e(t/p7/p71/p8) for an uncensored observation t/p7/p71/p8and /p57e(T) for an observation censored at T.In practice, for a censored observation t/p62/p71, w/p71/p58/p57e(t/p7/p72/p8), where t/p7/p72/p8is the largest uncensored observation that t/p7/p72/p8/p45t/p62/p71. Thus, the larger the uncensored observation, the smaller its score.Censoredobservations receive negative scores.The wscores sum identically to zero for the two groups together.The logrank test is based on the sum Sof thewscores of the two groups.The permutational variance of Sis given by Var(S)/p58n/p16n/p17/p26/p76/p129/p62/p76/p130/p71/p14/p16w/p17/p71(n/p16/p59n/p17)(n/p16/p59n/p17/p571)(5.1.8) which can be rewritten as V/p58/p3/p73/p26 /p72/p14/p16m/p7/p72/p8(r/p7/p72/p8/p57m/p7/p72/p8) r/p7/p72/p8 /p4n/p16n/p17(n/p16/p59n/p17)(n/p16/p59n/p17/p571)(5.1.9) The test statistic L/p58S//p40Var(S)has an asymptotically standard normal distribution under the null hypothesis.If Sis obtained from group 1, the critical region is L/p58/p57Z/p63, and ifSis obtained from group 2, the critical region is L/p57Z/p63, where /afii9825is the significance level for testing H/p15:S/p16/p58S/p17against H/p16:S/p16/p57S/p17.The following example illustrates the computational procedures. Example 5.3 Consider the data and hypotheses in Example 5.1. The test statistic of the logrank test can be computed by tabulating m/p7/p71/p8,r/p7/p71/p8,m/p7/p71/p8/r/p7/p71/p8, ande(t/p7/p71/p8) as in Table 5.3. Since every observation in the two samples, censored or not, is assigned a score, it is convenient to list them in column 1.Columns2 to 5 pertain only to the failure times; e(t/p7/p71/p8) is the cumulative value of m/p7/p71/p8/r/p7/p71/p8, Altshuler’s (1970 )estimate of the logarithm of the survivorship function multipled by /p571.For example, at t/p7/p71/p8/p5818,e(t/p7/p71/p8)/p580.100 /p590.125 /p580.225; at t/p7/p71/p8/p5819,e(t/p7/p71/p8)/p580.225/p590.333/p580.558.The last column, w/p71, gives the score for every observation.For an uncensored observation w/p71/p581/p57e(t/p7/p71/p8), for example, att/p71/p5818,w/p71/p581/p570.225/p580.775.Since e(t/p7/p71/p8) is an estimate of a function of the survivorship function, which we assume to be constant between twoconsecutive failures, e(t/p62/p71) is equal to e(t/p7/p72/p8) fort/p7/p72/p8/p45t/p62/p71.Thusw/p71for censored observations t/p62/p71equals /p57e(t/p7/p72/p8), where t/p7/p72/p8/p45t/p62/p71.For example, w/p71for 16 /p62is /p57e(15), or /p570.100, and that for 18 /p62is/p57e(18), or /p570.225. Tied observations like the two 19’s receive the same score: 0.442. The 10 scores w/p71sum to zero, which can be used to check the computation.112       Table 5.3 Computations of Logrank Test Remission Times in Both Samples, t/p71m/p7/p71/p8r/p7/p71/p8m/p7/p71/p8/r/p7/p71/p8e(t/p7/p71/p8)w/p71 15 1 10 0 .100 0 .100 0 .900 /p63 16/p59 —— — — /p570.100 18 1 8 0 .125 0 .225 0 .775 /p63 18/p59 —— — — /p570.225 19 2 6 0 .333 0 .558 0 .442 /p63 20 1 4 0 .250 0 .808 0 .192 /p63 20/p59 —— — — /p570.808 23 1 2 0 .500 1 .308 /p570.308 24/p59 —— — — /p571.308 /p63From sample 2. The statistic S/p580.900/p590.775/p590.442/p590.442/p590.192/p582.751.The vari- ance of S, computed by (5.8)is 1.210. Hence, the test statistic L/p582.751/ /p401.210/p582.5 and the pvalue is approximately 0.0064, data showing that CMF treatment is superior.The logrank statistic Scan be shown to equal the sum of the failures observed minus the conditional failures expected computed ateach failure time, or simply the difference between the observed and expectedfailures in one of the groups.A similar version of the logrank test is achi-square test which compares the observed number of failures to the expectednumber of failures under the hypothesis.Let O/p16andO/p17be the observed numbers and E/p16andE/p17the expected numbers of death in the two treatment groups.The test statistic X/p17/p58(O/p16/p57E/p16)/p17 E/p16 /p59(O/p17/p57E/p17)/p17 E/p17(5.1.10) has approximately the chi-square distribution with 1 degree of freedom.A large X/p17value (e.g.,/p46X/p17/p16/p11/p13/p15/p20)would lead to the rejection of the null hypothesis in favor of the alternative that the two treatments are not equally effective(/afii9825/p580.05). To compute E/p16andE/p17, we arrange all the uncensored observations in ascending order and compute the deaths expected at each uncensored time andsum them.The number of deaths expected at an uncensored time is obtainedby multiplying the deaths observed at that time by the proportion of patientsexposed to risk in the treatment group.Let d/p16be the number of deaths at time tandn/p16/p82andn/p17/p82be the numbers of patients still exposed to risk of dying at time up to tin the two treatment groups.The deaths expected for groups 1     113 Table 5.4 Computation of E1of Logrank Test Relapse time, td/p82n/p16/p82n/p17/p82e/p16/p82e/p17/p82 15 1 5 5 0.5 0.5 18 1 4 4 0.5 0.519 2 3 3 1.0 1.020 1 3 1 0.75 0.2523 1 2 0 1.0 0 Total 3.75 2.25and 2 at time tare e/p16/p82/p58n/p16/p82n/p16/p82/p59n/p17/p82/p59d/p82e/p17/p82/p58n/p17/p82n/p16/p82/p59n/p17/p82/p59d/p82(5.1.11 ) Then the total numbers of deaths expected in the two groups are E/p16/p58/p26 /p0/p12/p12/p82e/p16/p82E/p17/p58/p26 /p0/p12/p12/p82e/p17/p82 In practice, we only need to compute the total number of deaths expected in one of the two groups, for example, E/p16, sinceE/p17is the total observed number of deaths minus E/p16.The following example illustrates the calculation pro- cedure. Example 5.4 Let us use the hypothetical data in Example 5.1 again. The remission times in months are: CMF (group 1 ): 23, 16 /p59,1 8/p59,2 0/p59,2 4/p59 Control (group 2 ): 15, 18, 19, 19, 20. Consider the following null and alternative hypotheses: H/p15:S/p16/p58S/p17(the two treatments are equally effective ) H/p16:S/p16/p34S/p17(the two treatments are not equally effective ) Table 5.4 gives the calculation of E/p16.For example, at t/p5818, four patients in group 1 and four in group 2 are still exposed to the risk of relapse, and thereis one relapse.Thus, d/p82/p581,n/p16/p82/p58n/p17/p82/p584, ande/p16/p82/p580.5. The total number of relapses expected is E/p16/p583.75.Since there are a total of six deaths ( O/p16/p581,O/p17/p585)in the two groups, E/p17/p586/p573.75/p582.25.Using114       (5.1.10 ), we have X/p17/p58(1/p573.75)/p17 3.75/p59(5/p572.25)/p17 2.25/p585.378 Using Table C-2, the pvalue corresponding to this X/p17value is less 0.05 (p/p600.02).Therefore, we reach the same conclusion: that there is a significant difference in remission duration between the CMF and control groups. Computer software is available to perform a number of two-sample tests with censored observations.For example, SAS, SPSS, and BMDP provideprocedures for the logrank and Cox —Mantel tests.We use the remission time of the 10 breast cancer patients in Example 5.1 to illustrate the use of thesesoftware packages.To compare the two groups, we create the following threevariables: t, remission time; CENS /p580i ftis censored and 1 otherwise; and TREAT /p581 if receiving CMF and /p582 if no treatment.Assume that the data have been saved in ‘‘C: /p33D5d1.DAT’’ as a text file, which contains three columns, separated by a space (tis in the first column, CENS the second column, and TREAT the third column ), and the data in each row are for the same patient.The following SAS code can be used to perform the logrank test. data w1; infile ‘c: /p33d5d1.dat’ missover; input t cens treat; run; proc lifetest data /p58w1; time t*cens (0); strata treat; run; If BMDP procedure 1L is used, the following code can be used to perform the Cox—Mantel test. /input file /p58‘c:/p33d5d1.dat’ . variables /p583. format /p58free. /variable names /p58t, cens, treat. /form time /p58t. status /p58cens. response /p581. /group codes (treat )/p581, 2. Names (treat )/p58treated, control. /estimate method /p58product. Group /p58treat. Stat /p58mantel. /end     115 If procedure KM in SPSS is used, the following code can be used to perform the Cox—Mantel test. data list file /p58‘c:/p33d5d1.dat’ free / t cens treat. km t by treat /status /p58cens event (1) /test /p58logrank. These codes can be modified to perform tests comparing more than two groups simply by replacing TREAT in the codes with the group variable defined. 5.1.4 Peto and Peto’s Generalized Wilcoxon Test Another generalization of Wilcoxon’s two-sample rank sum test is described by Peto and Peto (1972 ).Similar to the logrank test, this test assigns a score to every observation.For an uncensored observation t, the score is u/p71/p58 S/p19(t/p59)/p59S/p19(t/p57)/p571, and for an observation censored at T, the score is u/p71/p58S/p19(T)/p571, where S/p19is the Kaplan —Meier estimate of the survival function. If we use the notation of Section 5.1.2, the score for an uncensored observationt/p7/p71/p8isu/p71/p58S/p19(t/p7/p71/p8)/p59S/p19(t/p7/p71/p92/p16/p8)/p571 andS/p19(t/p7/p15/p8)/p580 and that for a censored observa- tion ist/p62/p72isu/p72/p58S/p19(t/p7/p71/p8)/p571, where t/p7/p71/p8/p45t/p62/p72.These generalized Wilcoxon scores sum to zero.The test procedure after the scores are assigned is the same as forthe logrank test.The following example illustrates the computational pro-cedures. Example 5.5 Using the same data and hypotheses as in Example 5.1, the calculations of the scores u/p71for Peto and Peto’s generalized Wilcoxon test are given in Table 5.5. Using the scores of group 1, we obtain S/p58/p57 0.100/p570.212/p570.605/p570.408/p570.803/p58/p57 2.128 Var(S)/p58(5)(5)(0.9)/p17/p59/p37/p59(/p570.803) /p17 10/p599 /p580.765 Thus,Z/p58/p57 2.128//p400.765/p58/p57 2.433/p58/p57Z/p15/p13/p15/p20/p58/p57 1.64.We reject H/p15at the 0.05 level and reach the same conclusion as in the last three examples: that thedata show that CMB is more effective than no treatment. 5.1.5 Cox’s F-test Cox’sF-test (Cox, 1964 )is based on ordered scores from the exponential distribution.It is for singly censored or complete samples; it is not applicable to progressively censored data.The procedure is as follows:116       Table 5.5 Computations of Peto and Peto’s Generalized Wilcoxon Test t/p7/p71/p8S/p19(t) u/p71 15 0.900 1 /p590.900 /p571/p580.900 16/p59 — 0.900 /p571/p58/p57 0.100 /p63 18 0.788 0.900 /p590.788 /p571/p580.688 18/p59 — 0.788 /p571/p58/p57 0.212 /p63 19 0.657 0.788 /p590.657 /p571/p580.445 19 0.526 0.526 /p590.657 /p571/p580.183 20 0.395 0.395 /p590.526 /p571/p58/p57 0.079 20/p59 — 0.395 /p571/p58/p57 0.605 /p63 23 0.197 0.197 /p590.395 /p571/p58/p57 04.08 /p63 24/p59 — 0.197 /p571/p58/p57 0.803 /p63 /p63Group 1. 1.Rank the observations in the combined sample. 2.Replace the ranks by the corresponding expected order statistics in sampling the unit exponential distribution [ f(t)/p58e/p92/p82].Denote by t/p80/p76the expected value of the rth observation in increasing order of magnitude, t/p80/p76/p581 n/p59/p37/p591 n/p57r/p591r/p581,...,n (5.1.12) wherenis the total number of observations in the two samples.In particular, t/p16/p76/p581 n t/p17/p76/p581 n/p591 n/p571 /p36 t/p76/p76/p581 n/p591 n/p571/p59/p37/p591(5.1.13) Fornnot too large, they can easily be computed by using tables of reciprocals.When two or more observations are tied, the average of thescores is used. 3.For data without censored observations, the entire set of nobservations is replaced by the set of scores /p43t/p80/p76/p44so obtained.The sample mean scores denoted by t/p16/p16andt/p16/p17of the two samples with n/p16,n/p17observations are then computed.The ratio t/p16/p16/t/p16/p17has been shown to follow an Fdistribution with (2n/p16,2n/p17)degrees of freedom.Critical regions for testing H/p15:S/p16/p58S/p17     117 againstH/p16(S/p16/p57S/p17),H/p17(S/p16/p58S/p17), andH/p18(S/p16/p34S/p17) are, respectively, t/p16/p16/t/p16/p17/p57F/p17/p76/p129/p11/p17/p76/p130/p11/p63,t/p16/p16/t/p16/p17/p58F/p17/p76/p129/p17/p76/p130/p11/p16/p92/p63, andt/p16/p16/t/p16/p17/p57F/p17/p76/p129/p11/p17/p76/p130/p11/p63/p30/p17ort/p16/p16/ t/p16/p17/p58F/p17/p76/p129/p11/p17/p76/p130/p11/p16/p92/p63/p30/p17. 4.The calculation of Fis slightly different for singly censored data.Let r/p16andr/p17be the number of failures and n/p16/p57r/p16andn/p17/p57r/p17the number of censored observations in the two samples.Then there are p/p58r/p16/p59r/p17failures in the combined sample and n/p57pcensored observations.Cox (1964 )suggests using the scores t/p16/p76,...,t/p78/p76as before for the failures and t/p7/p78/p62/p16/p8/p76for all censored observations.The mean score, for example, for the first group is t/p16/p16/p58r/p16t/p16/p30/p16/p59(n/p16/p57r/p16)t/p7/p78/p62/p16/p8/p76r/p16(5.1.14 ) wheret/p16/p30/p16is the mean score of the failures.The mean score for the second group is calculated in a similar way.The F-statistic t/p16/p16/t/p16/p17, has an approximate F-distribution with (2r/p16,2r/p17)degrees of freedom. This test is for the hypothesis that the two samples are from populations with equal means.It can also determine if the second population mean is k times the first population mean, for a given k, by dividing the observations in the second sample by kbefore ranking and applying the test.The set of all valuesknot rejected in such a significance test forms a confidence interval.The following example illustrates the computation. Example 5.6 In an experiment comparing two treatments (A and B )for solid tumor, suppose that the question is whether treatment B is better thantreatment A.Six mice are assigned to treatment A and six to treatment B.Theexperiment is terminated after 30 days.The following survival times in days arerecorded.Our null and alternative hypotheses are H/p15:S/p31/p58S/p32and H/p16:S/p31/p58S/p32. Treatment A: 8, 8, 10, 12, 12, 13 Treatment B: 9, 12, 15, 20, 30 /p59,3 0/p59 That is, all the mice receiving treatment A die within 13 days and two mice receiving treatment B are still alive at the end of the study.Do the data providesufficient evidence that treatment B is more effective than treatment A? To compute the test statistic, it is convenient to set up a table like Table 5.6. The first column lists all the observations in the two samples.The secondcolumn contains the ordered exponential scores t/p80/p76.In this case, n/p16/p586,n/p17/p586, n/p5812,r/p16/p586, andr/p17/p584.The scores are computed following (5.1.12 )and (5.1.13 ).For example, t/p80/p76fort/p71/p5810 is equal to 1/12 /p591/11 /p591/10 /p591/9 or simply the previous t/p80/p76plus 1/9, that is, 0.274 /p591/9/p580.385. The tied observations receive an average score: for example, for t/p71/p5812,118       Table 5.6 Computations of Cox’s F-Test for Data in Example 5.6 t/p80/p76of t/p80/p76of t/p71t/p80/p76Sample A Sample B 8 /p16/p16/p17/p580.0831 0.129 — 8 /p16/p16/p17/p59/p16/p16/p16/p580.174/p440.1290.129 — 9 /p16/p16/p17/p59/p16/p16/p16/p59/p16/p16/p15/p580.174 /p590.100 /p580.274 — 0.274 10 0.274 /p59/p16/p24/p580.385 0.385 — 12 0.385 /p59/p16/p23/p580.510 0.661 — 12 0.510 /p59/p16/p22/p580.661/p440.661 0.661 — 12 0.653 /p59/p16/p21/p580.820 — 0.661 13 0.820 /p59/p16/p20/p581.020 1.020 — 15 1.020 /p59/p16/p19/p581.270 — 1.270 20 1.270 /p59/p16/p18/p581.603 — 1.603 30/p59 1.603 /p59/p16/p17/p582.103 — 2.103 30/p59 2.103 — 2.103 ——— ——— 2.985 8.014 t/p80/p76/p58/p16/p18(0.510/p590.653/p590.820) /p580.661.The last two columns of Table 5. 6 give the scores for the two samples and the sums are entered at the bottom.Thust/p16/p31/p582.985/6/p580.498 and t/p16/p32/p588.014/4/p582.004 according to (5.1.14 )and F/p58t/p16/p31t/p16/p32/p580.498 2.004/p580.249 with (12, 8 )degrees of freedom.The critical region is F/p58F/p16/p17/p11/p23/p11/p15/p13/p24/p20/p58 1/F/p23/p11/p16/p17/p11/p15/p13/p15/p20/p581/2.8486 /p580.351 for /afii9825/p580.05. /p18Hence, the data provide strong evidence that treatment B is superior to treatment A. 5.1.6 Comments on the Tests The tests presented in Sections 5.1.1 to 5.1.5 are based on rank statistics obtained from scores assigned to each observation.The first four tests areapplicable to data with progressive censoring.They can be further groupedinto two categories: generalization of the Wilcoxon test (Gehan’s and Peto and Peto’s )and the non-Wilcoxon test (Cox—Mantel and logrank test ).In the logrank test, if the statistic Sis the sum of wscores in group 2, it is the same asUof the Cox —Mantel test.This can be seen in Examples 5. 2 (U/p582.75) and 5.3(S/p582.751); the small discrepancy is due to rounding-off errors. /p18F/p80/p129/p11/p80/p130/p11/p63/p581/F/p80/p130/p11/p80/p129/p11/p16/p92/p63.     119 The only reason to choose one test over another in a given circumstance is if it will be more powerful, that is, more likely to reject a false hypothesis.Whensample sizes are small (n,n/p17/p4550), Gehan and Thomas (1969 )show that Cox’s F-test is more powerful than Gehan’s generalized Wilcoxon test if samples are from exponential or Weibull distributions and if there are no censoredobservations or the observations are singly censored.Comparisons of Gehan’sWilcoxon test to several other tests are reported by Lee et al. (1975 ).They show that when samples are from exponential distributions, with or without censoring, the Cox —Mantel and logrank tests are more powerful and more efficient than the generalized Wilcoxon tests of Gehan and Peto and Peto.There is little difference between the Cox —Mantel and logrank tests and between the two generalized Wilcoxon tests.When the samples are taken fromWeibull distributions with a constant hazard ratio (i.e., the ratio of the two hazard functions does not vary with time ), the results are essentially the same as in the exponential case.However, when the hazard ratio is nonconstant,the two generalizations of the Wilcoxon test have more power than the othertests.Thus, the logrank test is more powerful than the Wilcoxon tests indetecting departures when the two hazard functions are parallel (proportional hazards )or when there is random but equal censoring and when there is no censoring in the samples (Crowley and Thomas, 1975 ).The generalized Wilcoxon tests appear to be more powerful than the logrank test for detectingmany other types of differences, for example, when the hazard functions are notparallel and when there is no censoring and the logarithm of the survival timesfollow the normal distribution with equal variance but possibly differentmeans. The generalized Wilcoxon tests give more weight to early failures than later failures, whereas the logrank test gives equal weight to all failures.Therefore,the generalized Wilcoxon tests are more likely to detect early differences in thetwo survival distributions, whereas the logrank test is more sensitive todifferences in the right tails.Prentice and Marek (1979 )show that Gehan’s Wilcoxon statistic is subject to a serious criticism when censoring rates arehigh.If heavy censoring exists, the test statistic is dominated by a small numberof early failures and has very low power. There are situations in which neither the logrank nor Wilcoxon test is very effective.When the two distributions differ but their hazard functions orsurvivorship functions cross, neither the Wilcoxon nor logrank test is verypowerful, and it will be sensible to consider other tests.For example, Taroneand Ware (1977 )discuss general statistics of similar form (using scores )and Fleming and Harrington (1979 )and Fleming et al. (1980 )present a two-sample test based on the maximum of a Smirnov-type statistic designed to measure themaximum distance between estimates of two distributions.The latter approachis shown to be more effective than the logrank or Wilcoxon tests when twosurvival distributions differ substantially for some range of tvalues, but not necessarily elsewhere.These statistics have not been widely applied.Interestedreaders are referred to the original papers.120       5.2 MANTEL--HAENSZEL TEST The Mantel —Haenszel (1959 )test is particularly useful in comparing survival experience between two groups when adjustments for other prognostic factorsare needed.The test has been used in many clinical and epidemiological studiesas a method of controlling the effects of confounding variables.For example,in comparing two treatments for malignant melanoma, it would be importantto adjust the comparison for a possible confounding variable such as stage of the disease.In studying the association of smoking and heart disease, it wouldbe important to control the effects of age.To use the Mantel —Haenszel test, the data are stratified by the confounding variable and cast into a sequence of2/p592 tables, one for each stratum. Letsbe the number of strata, n/p72/p71be the number of individuals in group j, j/p581, 2, and stratum i,i/p581,...,s, andd/p72/p71be the number of deaths (or failures ) in group jand stratum i.For each of the sstrata, the data can be represented by a 2/p592 contingency table: Number of Number of Group Deaths Survivors Total 1 d/p16/p71n/p16/p71/p57d/p16/p71n/p16/p712 d/p17/p71n/p17/p71/p57d/p17/p71n/p17/p71Total D/p71S/p71T/p71 The null hypothesis to be tested can be stated as H/p15:p/p16/p16/p58p/p16/p17 p/p17/p16/p58p/p17/p17 /p36 p/p81/p16/p58p/p81/p17 wherep/p71/p72/p58P(death /p34groupj, stratum i).Thus, the test permits simultaneous comparison over all the scontingency tables of the difference in survival or death probabilities for the two groups. The chi-square test statistic without continuity correction /p19is given by X/p17/p58[/afii9814/p81/p71/p14/p16d/p16/p71/p57/afii9814/p81/p71/p14/p16(d/p16/p71)]/p17 /afii9814/p81/p71/p14/p16Var(d/p16/p71)(5.2.1) /p19According to Grizzle (1967 ), the distribution of X/p17without continuity correction is closer to the chi-square distribution than the X/p17with continuity correction.His simulations show that the probability of Type I error (rejecting a true hypothesis )is better controlled without the continuity correction at /afii9825/p580.01, 0.05.—  121 where E(d/p16/p71)/p58n/p16/p71D/p71T/p71(5.2.2 ) Var(d/p16/p71)/p58n/p16/p71n/p17/p71D/p71S/p71T/p17/p71(T/p71/p571)(5.2.3 ) are the mean and variance, respectively, of the number of deaths in group i computed conditionally on the contingency table marginal totals.This statisticfollows the chi-square distribution with 1 degree of freedom.Thus, a computedchi-square value larger than the table chi-square value for the significance levelchosen indicates a significant difference in survival between the two groups.The following two examples illustrate the use of the test. Example 5.7 Five hundred and ninety-five persons participate in a case control study of the association of cholesterol and coronary heart disease(CHD ).Among them, 300 persons are known to have CHD and 295 are free of CHD.To find out if elevated cholesterol is significantly associated withCHD, the investigator decides to control the effects of smoking.The studysubjects are then divided into two strata: smokers and nonsmokers. The following tables give the data for smokers: Elevated Cholesterol? With CHD Without CHD Total Yes 120 20 140No 80 60 140 Total 200 80 280 and for nonsmokers: Elevated Cholesterol? With CHD Without CHD Total Yes 30 60 90 No 70 155 255 Total 100 215 315122       Using (5.2.2 )and (5.2.3 ), we obtain E(d/p16/p16)/p58140/p59200 280/p58100E(d/p16/p17)/p5890/p59100 315/p5828.571 Var(d/p16/p16)/p58140/p59140/p59200/p5980 (280) /p17(280 /p571)/p5814.337 Var(d/p16/p17)/p5890/p59225/p59100/p59215 (315) /p17(315 /p571)/p5813.974 Using (5.2.1 )andd/p16/p16/p58120,d/p16/p17/p5830, we have X/p17/p58(150 /p57128.571) /p17 14.337/p5913.974/p5816.220 which is significant at the 0.001 level. Thus, elevated cholesterol is significantly associated with CHD after adjusting for the effects of smoking. Example 5.8 Table 5.7 gives survival data in life-table format of male cases with localized cancer of the rectum in Connecticut for 1935 —1944 and 1945—1954.We use Mantel and Haenszel’s chi-square test to see if the survival distribution of patients diagnosed in 1935 —1944 is the same as for patients diagnosed in 1945 —1954.The null hypothesis is that the two survival distribu- tions are the same.It is not necessary to set up 10 contingency tables for the10 intervals.The chi-square value is easily calculated by constructing columns7 to 12 directly from the life table.Using the sums in columns 1, 10, and 12,we obtain X/p17/p58(330.0/p57246.50)/p17 132.491 /p5852.624 which is significant at the 0.001 level. Thus, the data show a significant difference between the survival distributions of patients diagnosed in 1935 — 1944 and 1945 —1954. It should be noted that this chi-square test statistic, when applied to life tables, gives more weight to those deaths that occur in an early time intervalrather than later.That is, if the two groups are subject to the same probabilityof surviving through the entire study period, (5.2.1 )—(5.2.3 )will give high mortality for the group in which early deaths occur.Mantel (1966 )gives the following illustration. Consider two groups of 100 persons each.Both have 50 deaths.In group 1 all deaths occur in the first interval, and in group 2 all deaths occur in the—  123 Table 5.7 Computational Procedure for Comparing Two Survival Distributions: Male Cases with Localized Cancer of Rectum in Connecticut (1) (2) (3) (4) (5) (6) (7) (8) (9) (10) (11) (12) Combined 1935—1944 1945 —1954 Time Periods E(d/p16/p71) n/p17/p71S/p71/T/p71Var(d/p16/p71) Deaths, Survivors, Total, Deaths, Survivors, Total, Deaths, Survivors, Total, Internal d/p16/p71n/p16/p71/p57d/p16/p71n/p16/p71d/p17/p71n/p17/p71/p57d/p17/p71n/p17/p71D/p71S/p71T/p71(3)/p59(7) (9)(6)/p59(8) (9)(10)/p59(11) (9)/p571.0 1 167 220 .0 387 .0 185 559 .0 744 .0 352 779 .01 1 3 1 .01 2 0 .45 512 .45 54 .624 2 45 173 .5 218 .5 88 461 .0 549 .0 133 634 .57 6 7 .53 7 .86 453 .86 22 .418 3 45 127 .5 172 .5 55 396 .0 451 .0 100 523 .56 2 3 .52 7 .67 378 .67 16 .832 4 19 108 .0 127 .0 43 343 .0 386 .0 62 451 .05 1 3 .01 5 .35 339 .35 10 .174 51 7 9 1 .0 108 .0 32 299 .0 331 .0 49 390 .04 3 9 .01 2 .05 294 .05 8 .090 61 1 7 9 .59 0 .5 31 235 .0 266 .0 42 314 .53 5 6 .51 0 .66 234 .66 7 .036 78 7 1 .07 9 .0 20 170 .0 190 .0 28 241 .02 6 9 .08 .22 170 .22 5 .221 85 6 6 .07 1 .0 7 132 .0 139 .0 12 198 .02 1 0 .04 .06 131 .06 2 .546 96 5 9 .56 5 .5 6 101 .5 107 .5 12 161 .01 7 3 .04 .54 100 .04 2 .641 10 7 52 .05 9 .06 7 1 .07 7 .0 13 123 .01 3 6 .05 .64 69 .64 2 .909—— ——— ——— —330 246.50 132 .491 124 second interval.The contingency table for the first interval is: Group Deaths Survivors Total 1 50 50 100 2 0 100 100 Total 50 150 200 and for the second interval is: Group Deaths Survivors Total 1 0 50 50 2 50 50 100 Total 50 100 150 From these two tables we have E(d/p16/p16)/p58100/p5950/200/p5825 and E(d/p16/p17)/p58 100/p5950/150/p5816.67.The total deaths expected is 25 /p5916.67 /p5841.67, so the 50 deaths in group 1 is 20% larger than expected.Thus, a significant chi-squarevalue may be obtained if early survival patterns differ significantly in the twogroups. 5.3 COMPARISON OF K( K /p572) SAMPLES In this section the two-sample problem is generalized to a situation in which the data consist of K(K/p572)samples, one sample from each of the K treatment populations.The problem is to decide whether the Kindependent samples can be regarded as coming from the same population, or in practicalterms, to see if the survival data from patients receiving the Ktreatments provide enough evidence to conclude that the Ktreatments are not equally effective.This problem has been considered by many statisticians: for example,Kruskal and Wallis (1952 ), Mantel and Haenszel (1959 ), Breslow (1970 ), and Peto and Peto (1972 ).In this section two nonparametric tests for the problem are presented.The first is Kruskal and Wallis’s (1952 )H-test for uncensored data.The second is a generalization of the H-test for censored data (Peto and Peto, 1972 ).Both use ranks instead of the original observations and are simple to apply. 5.3.1 Kruskal--Wallis Test The Kruskal —WallisH-test (Kruskal and Wallis, 1952; Hollander and Wolfe, 1973; Marascuilo and McSweeney, 1977 ), analogous to the F-test in the usual analysis of variance, uses ranks rather than original observations; it is alsocalled the Kruskal —Wallis one-way analysis of variance by ranks.It assumesCOMPARISON OF K(K/p572) SAMPLES 125 that the variable (survival time )under study has an underlying continuous distribution. LetNbe the total number of independent observations in the Ksamples, n/p72the number of observations in the jth sample, j/p581,...,K, andt/p71/p72theith observation in the jth sample.The null hypothesis H/p15states that the Ksamples come from the same population (or clinically, the Ktreatments are equally effective ). In computation of the Kruskal —WallisH-test, we first rank all Nobserva- tions from smallest to largest.Let r/p71/p72be the rank of t/p71/p72.Compute, for j/p581,...,K, R/p72/p58/p76/p72/p26 /p71/p14/p16r/p71/p72R/p16/p72/p58R/p72n/p72R/p16/p581 2(N/p591) (5.3.1 ) whereR/p72andR/p16/p72are, respectively, the sum of the ranks and the average rank of thejth treatment, and R/p16is the overall average rank.Then the Kruskal — WallisH-statistic is H/p5812 N(N/p591)/p41/p26 /p72/p14/p16n/p72(R/p16/p72/p57R/p16)/p17 (5.3.2) /p5812 N(N/p591)/p41/p26 /p72/p14/p16R/p17/p72n/p72/p573(N/p591) (5.3.3 ) Under the null hypothesis, Hhas an asymptotic (n/p72’s approaching infinity or n/p72’s are large )chi-square distribution with K/p571 degrees of freedom.Thus, for largen/p72’s, the approximate test procedure at the /afii9825level is to reject H/p15if H/p46/afii9851/p17/p7/p73/p92/p16/p8/p11/p63.WhenK/p583 and the number of observations in each of the three samples is 5 or fewer, the chi-square approximation is not sufficiently close.Forsuch cases, exact permutational distributions of Hare available and percentage points /afii9851/p73are given in Table B-4 of Appendix B.The test procedure is to reject H/p15ifH/p46/afii9851/p73/p11/p63, where /afii9851/p73/p11/p63satisfies the equation P(H/p46/afii9851/p73/p11/p63/p34H/p15)/p58/afii9825. When there are tied observations, each is assigned the average of the ranks. To correct for the effects of ties, His computed by (5.3.3 )and then divided by 1/p571 N/p18/p57N/p69/p26 /p72/p14/p16T/p72(5.3.4 ) wheregis the number of tied groups, and T/p72/p58t/p18/p72/p57t/p72, witht/p72being the number of tied observations in a tied group.In counting g, an untied observation is considered as a tied group of size 1.Thus, a general expression of Hcorrected for ties is H/p58[12/N(N/p591)]/afii9814/p73/p72/p14/p16(R/p17/p72/n/p72)/p573(N/p591) 1/p57/afii9814/p69/p72/p14/p16T/p72/(N/p18/p57N)(5.3.5)126       Table 5.8 Cholesterol Values of 12 Subjects on Three Different Diets Diet 1 Diet 2 Diet 3 229 145 231176 181 208187 147 217 208 187 199 Table 5.9 Computation of Hfor Data in Example 5.9 Ordered Ranks of Ranks of Ranks of Observations Ranks Diet 1 Diet 2 Diet 3 145 1 — 1 147 2 — 2176 3 3181 4 — 4 187 5.5 — 5.5 187 5.5 5.5199 7 — — 7208 8.5 8.5208 8.5 — — 8.5217 10 — — 10 229 11 11 231 12 — — 12 R/p7228 12.5 37.5Note that when there are no ties, g/p58N,t/p72/p581 for allj, andT/p72/p580, and (5.3.5 ) reduces to (5.3.3 ).The following example illustrates the use of the test. Example 5.9 In a study of the relationship between cholesterol level and diet, three diets are given randomly to 12 men whose initial cholesterol levelsare almost the same.Table 5. 8 shows the cholesterol levels of the 12 peopleafter having their assigned diet for a given period of time.The purpose of thestudy is to decide if the three diets are equally effective in controllingcholesterol level. The null hypothesis H/p15states that there is no difference in cholesterol level of men having the three diets, and the alternative H/p16says that the cholesterol levels of men having the three different diets are different.To compute theH-statistic, we first rank the observations as in Table 5.9 and compute R/p72.In this case N/p5812,n/p16/p58n/p17/p58n/p18/p584,g/p5810, and T/p72/p580 except for j/p585, 7.COMPARISON OF K(K/p572) SAMPLES 127 Hence, /afii9814/p69/p72/p14/p16T/p72/p582(8/p572)/p5812, and by (5.3.5 ), H/p58[12/(12/p5913)](784/4/p59156.25/4/p591406.25/4)/p573(13) 1/p5712/(1728 /p5712)/p586.168 From Table B-4 we find that P(H/p466.038/p34H/p15)/p580.037 and P(H/p466.269/p34H/p15) /p580.033; we reject H/p15at the /afii9825/p32 0.035 level. There is evidence of significant differences among the diets. 5.3.2 Multiple Comparisons Based on the Kruskal--Wallis Test If the null hypothesis that the Ksamples are from the same distribution is rejected, we might ask which particular samples are from different distribu-tions.In Example 5. 9 we reject the null hypothesis that the three diets aresimilar.The investigator may also be interested in knowing which particulardiets differ from one another.In this section we introduce some nonparametricmethods for multiple comparison based on Kruskal —Wallis rank sums.An excellent treatment of multiple comparisons is given by Miller (1966 ). To decide which treatments differ from one another, there are /p16/p17 K(K/p571) decisions to make, one for each pair of treatments.The null hypothesis can bewritten as H/p15: samples iandjare from the same population for i/p581,...,K/p571,j/p58i/p591,...,K,i/p58j Let the probability of at least one wrong decision when H/p15is true be controlled by/afii9825and the probability of making all correct decisions when H/p15is true be 1/p57/afii9825.To make the /p16/p17K(K/p571) decisions, we introduce the following compari- son procedures. 1.When sample sizes are equal, that is, n/p16/p58n/p17/p58/p37/p58n/p41/p58n, andnis small, we reject the hypothesis that the ith andjth samples, i/p58j, are from the same distribution if /p34R/p71/p57R/p72/p34/p46y(/afii9825,K,n) (5.3.6 ) wherey(/afii9825,K,n)satisfies the equation P(/p34R/p71/p57R/p72/p34/p46y(/afii9825,K,n)/p34H/p15,i/p58j)/p581/p57/afii9825 (5.3.7) andR/p16,R/p17,...,R/p41are given in (5.3.1 ).Some approximate values of yare given in Table B-5. 2.When sample sizes are equal to nandnis large, we introduce Miller’s (1966 )procedure, that is, to reject the hypothesis that the ith andjth128       Table 5.10 Multiple Comparisons for Data in Example 5.9 ij /p34R/p71/p57R/p72/p34 Decision 12 /p3428/p5712.5/p34/p5815.5 Not significant 13 /p3428/p5737.5/p34/p589.5 Not significant 23 /p3412.5 /p5737.5/p34/p5825.0 Significant samples, i/p58j, are from the same distribution if /p34R/p16/p71/p57R/p16/p72/p34/p46q(a,K)[/p16/p16/p17K(Kn/p591)]/p16/p30/p17 (5.3.8) whereR/p16/p16,...,R/p16/p41are given in (5.3.1 )andq(/afii9825,K)is the upper /afii9825percentile point of the range of Kindependent standard normal variables.Table B-6 gives the q(/afii9825,K)values for some Kand/afii9825. 3.For cases of small unequal sample sizes n/p16,...,n/p41, a conservative procedure is to reject the hypothesis that the ith andjth samples, i/p58j, are from the same distribution if /p34R/p16/p71/p57R/p16/p72/p34/p46(x/p63/p11/p41)/p16/p30/p17/p31 12N(N/p591)/p4/p16/p30/p17/p11 n/p71/p591 n/p72/p2/p16/p30/p17(5.3.9) whereNis the total number of observations.Values of x/p63/p11/p41are given in Table B-4. 4.When n/p16,...,n/p41are large, Dunn (1964 )suggests the following procedure. Reject the hypothesis that the ith andjth samples, i/p58j, are from the same distribution if /p34R/p16/p71/p57R/p16/p72/p34/p46Z/p63/p30/p9/p41/p7/p41/p92/p16/p8/p10/p31 12N(N/p591)/p4/p16/p30/p17/p11 n/p71/p591 n/p72/p2/p16/p30/p17(5.3.10) whereZ/p63/p30/p9/p41/p7/p41/p92/p16/p8/p10is the upper 100 /afii9825/[K(K/p571)] percentage point of the standard normal distribution (see Table B-1 ). Example 5.10 Let us use the data in Example 5.9. To examine which particular diets differ from one another, we apply (5.3.6 ).SinceK/p583, there are three possible comparisons.The calculation is shown in Table 5. 10.For K/p583, n/p584, and from Table B-5, y(0.045, 3, 4) /p5824; hence for ( i,j)/p58(2, 3), (5.3.6 )is satisfied.Thus, at /afii9825/p600.045, we conclude that diets 2 and 3 are significantly different.COMPARISON OF K(K/p572) SAMPLES 129 Table 5.11 Initial Remission Times of Leukemia Patients 12 3 4, 5, 9, 10, 12, 13, 10, 8, 10, 10, 12, 14, 8, 10, 11, 23, 25, 25, 23, 28, 28, 28, 29 20, 48, 70, 75, 99, 103, 28, 28, 31, 31, 40,31, 32, 37, 41, 41, 162, 169, 195, 220, 48, 89, 124, 143,57, 62, 74, 100, 139, 161 /p59, 199 /p59, 217 /p59,1 2 /p59, 159 /p59, 190 /p59, 196 /p59, 20/p59, 258 /p59, 269 /p59 245/p59 197/p59, 205 /p59, 219 /p595.3.3 Test for Censored Data In Section 5.1 we introduced three nonparametric tests based on scores for comparing two samples with censored observations; Gehan’s generalizedWilcoxon test (if Mantel’s procedure is used ), Peto and Peto’s generalized Wilcoxon test, and the logrank test.The K-sample test discussed in this section can be considered an extension of these tests and the Kruskal —Wallis test. Suppose that we have a set of Nscoresw/p16,w/p17,...,w/p44obtained according to the manner of scoring in one of the three tests mentioned above.The sumof theNscores is zero.Let S/p72be the sum of the scores in the jth sample.The null hypothesis H/p15states that the Ksamples are from the same distribution. To testH/p15we calculate X/p17/p58/afii9814/p41/p72/p14/p16(S/p17/p72/n/p72) s/p17 (5.3.11) where s/p17/p58/afii9814/p44/p71/p14/p16w/p17/p71N/p571(5.3.12) Under the null hypothesis X/p17has approximately chi-square distribution with K/p571 degrees of freedom (Peto and Peto, 1972 ).Thus, we reject H/p15ifX/p17 exceeds the upper 100 /afii9825percentage point of the chi-square distribution with K/p571 degrees of freedom, that is, if X/p17/p46/afii9851/p17/p7/p73/p92/p16/p8/p11/p63. Example 5.11, using the scoring method of Mantel (1967 )for Gehan’s generalized Wilcoxon test, illustrates the K-sample test for censored data. Example 5.11 Consider the initial remission times of leukemia patients (in days )induced by three treatments as given in Table 5.11. In this case, K/p583, N/p5866,n/p16/p5825,n/p17/p5819, andn/p18/p5822.A table similar to Table 5. 1 may be set up to compute the score for every observation.The computation is left to thereader as an exercise.The sums of scores in the three samples are S/p16/p58/p57 273, S/p17/p58170, and S/p18/p58103.The sum of squares of the scores /afii9814w/p17/p71/p5889,702. Hence, from (5.3.11 )and (5.3.12 ),X/p17/p58 3.612.From Table B-2, /afii9851/p17/p17/p11/p15/p13/p15/p20/p585.991;130       thus we do not reject H/p15.The data do not show significant differences among the three initial treatments. Bibliographical Remarks Gehan’s test was first proposed in 1965.The Cox —Mantel test was first discussed by Cox in 1959, then by Mantel in 1966, and finally, by Cox againin 1972.The scores for the logrank test was proposed by Peto and Peto in 1972 along with another generalization of the Wilcoxon test.In the same paper, theyalso discuss the K-sample test for censored data.The logrank test is also discussed in Peto et al. (1977 ).Cox’sF-test was developed in 1964 and the Mantel—Haenszel chi-square test can be found in Mantel and Haenszel (1959 ) and Mantel (1966 ).The Kruskal —Wallis one-way analysis of variance can be found in most standard textbooks under nonparametric methods.Readers whoare interested in the theoretical development or more properties of these testsshould read the original papers cited above.Applications of these tests aregiven in the original papers or can easily be found in medical and epidemiologi-cal journals. EXERCISES The first five exercises are continuations of Exercises 4.1 to 4.4 and 4.6.5.1 For the survival times given in Table 3.1, compare the survival distribu- tions of the two treatment groups using: (a) Gehan’s generalized Wilcoxon test (b) The Cox —Mantel test 5.2 For the remission data given in Table 3.1, compare the remission time distributions of the two treatment groups using: (a) The logrank test (b) Peto and Peto’s generalized Wilcoxon test 5.3 For the data given in Table 3.4, compare the tumor-free time distribu- tions of the three diet groups. 5.4 For the remission data of 42 leukemia patients given in Example 3.3, use the two generalized Wilcoxon tests to see if 6-MP is more effective thanplacebo in prolonging remission time. 5.5 For the first four skin tests given in Exercise Table 3.1, use the Cox—Mantel and logrank tests to see if there is a significant difference in survival between patients with positive (/p465 mm for mumps, /p4610 mm for others )and negative (/p585 mm for mumps, /p5810 mm for others )reactions. 131 Exercise Table 5.1 Percentage of Male Female Standard BMI /p63 Case Control Case Control /p58140 130 160 60 65 /p46140 85 55 55 50 /p63Standard BMI: male, 22.1; female, 20.6. Percentage of standard BMI /p58(observed BMI/standard BMI )/p59100. Exercise Table 5.2 12 3 10.5 10.0 12.0 9.0 12.0 13.0 9.5 12.5 15.59.0 11.0 14.08.5 12.0 12.5 10.0 10.5 15.05.6 Compute the test statistic Wof Gehan’s generalized Wilcoxon test by using (5.1.1 )for the data in Example 5.1. Do you get the same result as in Example 5.1? 5.7 Consider the data in Example 5.11. Use Mantel’s procedure for Gehan’s generalized Wilcoxon test to compute a score for each observation andthe sum of scores for each of the three treatment groups. 5.8 Using the data in Table 3.1, compare the age distributions of the two treatment groups using Cox’s F-test. 5.9 Consider the data in Exercise Table 5.1. Is elevated percent standard BMI associated with renal cell carcinoma after controlling the effects ofgender? 5.10 Consider the survival data of men with angina pectoris in Table 4.6 and women with the same disease in Exercise Table 4.2. Is there a significantdifference between the survival distributions of men and women? 5.11 In a study of noise level and efficiency, 18 students were given a very simple test under three different noise levels.It is known that under132       Exercise Table 5.3 12 2 3 41 3 5 54 7 1 59 9 14 20 12 12 20 31 20/p59 15 27 39 25 23 30 4730/p59 30 32 /p59 55/p59 50/p59 67/p59normal conditions, they should be able to finish the test in 10 minutes. The students were randomly assigned to the three levels.Exercise Table5.2 gives the time required to finish the test. Are the three noise levelssignificantly different? If they are, determine which levels differ from oneanother. 5.12 Exercise Table 5.3 gives the survival time in weeks of 30 brain tumor patients receiving four different treatments.Are the four treatments equally effective? 133 CHAPTER6 SomeWell-KnownParametric SurvivalDistributionsandTheirApplications Usually,therearemanyphysicalcausesthatleadtothefailureordeathofapersonataparticulartime.Itisverydifficult,ifnotimpossible,toisolatethesephysicalcausesandaccountmathematicallyforallofthem.Therefore,choos-ingatheoreticaldistributiontoapproximatesurvivaldataisasmuchanartasascientifictask.Inthischapter,severaltheoreticaldistributionsthathavebeenusedwidelytodescribesurvivaltimearediscussed,theircharacteristicssummarized,andtheirapplicationsillustrated. 6.1 EXPONENTIAL DISTRIBUTION The simplest and most important distribution in survival studies is the exponential distribution. In the late 1940s, researchers began to choose theexponentialdistributiontodescribethelifepatternofelectronicsystems.Davis(1952 )givesanumberofexamples,includingbankstatementandledgererror, payroll check errors, automatic calculating machine failure, and radar setcomponent failure, in which the failure data are well described by theexponentialdistribution.EpsteinandSobel (1953 )reportwhytheyselectthe exponentialdistributionoverthepopularnormaldistributionandshowhowtoestimatetheparameterwhendataaresinglycensored.Epstein (1958 )also discussesinsomedetailthejustificationfortheassumptionofanexponentialdistribution.Theexponentialdistributionhassincecontinuedtoplayaroleinlifetimestudiesanalogoustothatofthenormaldistributioninotherareasofstatistics. Theexponentialdistributionisoftenreferredtoasapurelyrandomfailure pattern.Itisfamousforitsunique‘‘lackofmemory,’’whichrequiresthattheage of the animal or person does not affect future survival. Although many 134 Figure 6.1Exponentialdistribution: (a)survivorshipfunction;( b) probabilitydensity function;( c) hazardfunction. survivaldatacannotbedescribedadequatelybytheexponentialdistribution, anunderstandingofitfacilitatesthetreatmentofmoregeneralsituations. Theexponentialdistributionischaracterizedbyaconstanthazardrate /afii9838, itsonlyparameter.Ahigh /afii9838valueindicateshighriskandshortsurvival;alow /afii9838valueindicateslowriskandlongsurvival.Figure6.1depictsthesurvivorship function, the density function, and the hazard function of the exponentialdistributionwithparameter /afii9838.When /afii9838/p581,thedistributionisoftenreferredto astheunit exponential distribution. When the survival time Tfollows the exponential distribution with a parameter /afii9838,theprobabilitydensityfunctionisdefinedas f(t)/p58/p7/afii9838e/p92/p72/p82 0t/p460,/afii9838/p570 t/p580(6.1.1 ) Thecumulativedistributionfunctionis F(t)/p581/p57e/p92/p72/p82t/p460 (6.1.2 ) andthesurvivorshipfunctionisthen S(t)/p58e/p92/p72/p82t/p460 (6.1.3 ) Sothat,by (2.2.1 ),thehazardfunctionis h(t)/p58/afii9838t/p460( 6 .1.4) aconstant,independentof t.Figure6.1givesthegraphicalpresentationofthe threefunctions. Becausetheexponentialdistributionischaracterizedbyaconstanthazard rate,independentoftheageoftheperson,thereisnoagingorwearingout,  135 and failure or death is a random event independent of time. When natural logarithmsofthesurvivorshipfunctionaretaken,log S(t)/p58/p57/afii9838t,whichisa linearfunctionof t.Thus,itiseasytodeterminewhetherdatacomefroman exponentialdistributionbyplottinglog S/p19(t) against t,where S/p19(t) isanestimate ofS(t). A linear configurationindicates that the data follow an exponential distributionandtheslopeofthestraightlineisanestimateofthehazardrate /afii9838. Themeanandvarianceoftheexponentialdistributionwithparameter /afii9838are, respectively,1/ /afii9838and1/ /afii9838/p17.Themedianis (1//afii9838)log2.Thecoefficientofvariation is1. A moregeneral formof the exponentialdistributionis the two-parameter exponential distribution withprobabilitydensityfunction f(t)/p58/p7/afii9838e/p92/p72/p7/p82/p92/p37/p8 0t/p46G t/p58G(6.1.5 ) Then F(t)/p58/p71/p57e/p92/p72/p7/p82/p92/p37/p8 0t/p46G t/p58G(6.1.6 ) S(t)/p58/p7e/p92/p72/p7/p82/p92/p37/p8 0t/p46G t/p58G(6.1.7) and h(t)/p58/p70 /afii98380/p45t/p58G t/p58G(6.1.8) Theterm Gisaguarantee time withinwhichnodeathsorfailurescanoccur, or a minimum survival time. If G/p580, (6.1.5)—(6.1.8 )reduce to (6.1.1 )— (6.1.4 )for the one-parameter exponential. The mean and the median of the two-parameter exponential distribution are, respectively, G/p591//afii9838and (log2 /p59/afii9838G)//afii9838. Example 6.1 In a study of new anticancer drugs in the L1210 animal leukemiasystem,Zelen (1966 )usedtheexponentialdistributionsuccessfullyas themodelforsurvivaltime.Thesystemconsistsofinjectingatumorinoculuminto inbred mice. These tumor cells then proliferate and eventually kill theanimal, but survival time may be prolonged by an active drug. Figure 6.2shows the survival curve in a semilogarithmic scale of the untreated miceinoculatedatdifferentcelldilutions.Twenty-fivemicewereinoculatedateachdilution. The reasonably linear configurations suggest that the survival dis-tributionsfollowtheexponentialdistributionquitewell.Thefourstraightlinesfitted to the points are almost parallel, indicating that the hazard rate wasindependentoftheinoculumsize.Table6.1givestheestimatedvaluesofthe136 -    Figure 6.2Survivalcurvesofuntreatedmiceinoculatedwithserial10-folddilutionsof leukemiaL1210: (/p42)10/p20cells; (I)10/p19cells; (/p41)10/p18cells; (G)10/p17cells. (FromZelen, 1966. ) Table 6.1 Estimates of G,/afii9838, and Mean Survival Time for Untreated Mice with Serial Leukemia Dilutions Dilution G/p19/afii9838 /p19MeanSurvivalTime 10/p208.0 0.78 9.3 10/p1910.0 0.78 11.3 10/p1811.9 0.76 13.2 10/p1713.9 0.67 15.4 Source:Zelen (1966 ).guaranteetimes G,hazardrates /afii9838,andthemeansurvivaltimes (indays )for thevariousdilutions. (EstimationproceduresarediscussedinChapter7. )Note thattheestimatedhazardrates, /afii9838,areveryclose. Theprobabilitythat amousereceiving10 /p20cells ofinoculumwill survive morethan20daysis,from (6.1.7 ), S(20)/p58e/p92/p15/p13/p22/p23/p7/p17/p15/p92/p23/p8 /p600.0001 andthemediansurvivaltimeis8.9days.  137 Figure 6.3Survival curves of mice treated with cyclophosphamide on day 3 after inoculationof10 /p20tumorcells: (/p42)control; (I)80mg/kg; (G)160mg/kg. (FromZelen, 1966. ) Figure6.3givesthesurvivalcurvesofmicetreatedwithdifferentdosesof cyclophosphamideonday3afterreceivinga10 /p20tumorinoculum.Table6.2 gives the estimates of G,/afii9838, and the mean survival time. Mice receiving 160 mg/kg of the drug show a remarkably improved surivival pattern over thecontrolgroup. Theprobabilitythatamousereceiving80mg/kgofcyclophosphamidewill survivemorethan20daysis,from (6.1.7 ), S(20)/p58e/p92/p15/p13/p17/p24/p7/p17/p15/p92/p16/p20/p13/p21/p8 /p58 0.279 Themediansurvivaltimeisapproximately18days. 6.2 WEIBULL DISTRIBUTION The Weibull distribution is a generalization of the exponential distribution. However,unliketheexponentialdistribution,it does notassumea constanthazard rate and therefore has broader application. The distribution wasproposedbyWeibull (1939 )anditsapplicabilitytovariousfailuresituations discussedagainbyWeibull (1951 ).Ithasthenbeenusedinmanystudiesof reliabilityandhumandiseasemortality. TheWeibulldistributionischaracterizedbytwoparameters, /afii9828and/afii9838.The valueof /afii9828determinesthe shapeof thedistributioncurveandthevalueof /afii9838138 -    Table 6.2 Estimates of G,/afii9838, and Mean Survival Time for Treated Mice on Day 3 Dose (mg/kg )G/p19/afii9838 /p19MeanSurvivalTime Control 8.7 1.12 9.6 80 15.6 0.29 19.0 160 21.5 0.10 31.5 Source:Zelen (1966 ). Figure 6.4HazardfunctionsofWeibulldistributionwith /afii9838/p581.determinesits scaling. Consequently, /afii9828and/afii9838are called the shapeandscale parameters,respectively.Therelationshipbetweenthevalueof /afii9828andsurvival timecanbeseenfromFigure6.4,whichshowsthehazardrateoftheWeibulldistributionwith /afii9828/p580.5,1,2,4.When /afii9828/p581,thehazardrateremainsconstant astimeincreases;thisistheexponentialcase.Thehazardrateincreaseswhen/afii9828/p571anddecreaseswhen /afii9828/p581astincreases.Thus,theWeibulldistribution maybeusedtomodelthesurvivaldistributionofapopulationwithincreasing,decreasing, or constant risk. Examples of increasing and decreasing hazardrates are, respectively, patients with lung cancer and patients who undergosuccessfulmajorsurgery. Theprobabilitydensityfunctionandcumulativedistributionfunctionsare, respectively, f(t)/p58/afii9838/afii9828(/afii9838t)/p65/p92/p16e/p92/p7/p72/p82/p8 /p65t/p460,/afii9828,/afii9838/p570 (6.2.1 )  139 Figure 6.5DensitycurvesofWeibulldistributionwith /afii9838/p581.and F(t)/p581/p57e/p92/p7/p72/p82/p8/p65(6.2.2) Thesurvivorshipfunctionis,therefore, S(t)/p58e/p92/p7/p72/p82/p8/p65(6.2.3) andthehazardfunction,theratioof (6.2.1 )to(6.2.3 ),is h(t)/p58/afii9838/afii9828(/afii9838t)/p65/p92/p16 (6.2.4) Figure6.5givestheWeibulldensityfunctionwithscaleparameter /afii9838/p581and severaldifferentvaluesoftheshapeparameter /afii9828. Forthesurvivalcurve,itissimpletoplotthelogarithmof S(t), logS(t)/p58/p57(/afii9838t)/p65 (6.2.5) Figure6.6giveslog S(t) for/afii9838/p581and /afii9828/p581,/p571,/p581.When /afii9828/p581isastraight line with negative slope. When /afii9828/p581, negative aging, log S(t) decreasesvery slowly from 0 and then approaches a constant value. When /afii9828/p571, positive aging,log S(t)decreasessharplyfrom0as tincreases.Equation (6.2.5 )canalso140 -    Figure 6.6Curvesoflog/p67S(t) ofWeibulldistributionwith /afii9838/p581. bewrittenas log[/p57logS(t)]/p58/afii9828log/p67/afii9838/p59/afii9828log/p67t (6.2.6 ) ThemeanoftheWeibulldistributionis /afii9839/p58/afii9772(1/p591//afii9828) /afii9838(6.2.7 ) andthevarianceis /afii9846/p17/p581 /afii9838/p17/p3/afii9772/p11/p592 /afii9828/p2/p57/afii9772/p17/p11/p591 /afii9828/p2/p4(6.2.8 ) where /afii9772(/afii9828)isthewell-knowngammafunctiondefinedas /afii9772(/afii9828)/p58/p16/p27 /p15x/p65/p92/p16e/p92/p86dx /p58(/afii9828/p571)! when /afii9828isapositiveinteger (6.2.9 ) Valuesof /afii9772(/afii9828)canbefoundinAbramowitzandStegun (1964 ).Thecoefficient ofvariationisthen CV/p58/p3/afii9772(1/p592//afii9828) /afii9772/p17(1/p591//afii9828)/p571/p4/p16/p30/p17(6.2.10 ) The Weibull distribution can also be generalized to take into account a guarantee time Gduring which no deaths or failures can occur. The three-  141 Figure 6.7SurvivalcurvesofratsexposedtocarcinogenDMBA. (FromPike,1966. ReproducedwithpermissionoftheBiometricsSociety. )parameterWeibullprobabilitydensityfunctionis f(t)/p58/afii9838/p65/afii9828(t/p57G)/p65/p92/p16exp[ /p57/afii9838/p65(t/p57G)/p65] (6.2.11 ) Consequently, S(t)/p58exp[ /p57/afii9838/p65(t/p57G)/p65]( 6 .2.12) and h(t)/p58/afii9838/p65/afii9828(t/p57G)/p65/p92/p16 (6.2.13) Example 6.2 Pike (1966 )appliedtheWeibulldistributiontoatwo-group experimentonvaginalcancerinratsexposedtothecarcinogenDMBA.Thetwogroupsweredistinguishedbypretreatmentregime.Thetimesindays,afterthestartoftheexperiment,atwhichthecarcinomawasdiagnosedforthetwogroupsofratswereasfollows: Group1: 143,164,188,188,190,192,206,209,213,216,220,227,230, 234,246,265,304,216 /p59,244/p59 Group2: 142,156,173,198,205,232,232,233,233,233,233,239,240, 261,280,280,296,296,323204 /p59,344/p59 Assumingthat G/p58100and /afii9828/p583,Pikeobtained /afii9838/p19/p16/p58(4.51/p5910/p92/p22)/p16/p30/p18for group1and /afii9838/p19/p17/p58(2.38/p5910/p92/p22)/p16/p30/p18forgroup2.Analyticalestimationprocedure isdiscussedinChapter7.Figure6.7plotsthesurvivalcurvesofthetwogroups.ThestepfunctionsarenonparametricestimatessimilartotheKaplan —Meier142 -    Table 6.3 Calculation of Survivorship Functions for Group 2 of Rats Exposed to DMBA Kaplan—Modified Meier Kaplan —Meier S/p19(t) Obtained Estimates, Estimates /p63, fromWeibull Time rn /p57r/p591n/p57r n/p57r/p591 S/p19(t) S/p19(t) Plotted Fit 142 1 21 0.9524 0.9524 0.9546 0.9825 156 2 20 0.9500 0.9048 0.9091 0.9590 163 3 19 0.9474 0.8572 0.8637 0.9421 198 4 18 0.9444 0.8095 0.8182 0.7990204/p59— 17 1.0000 0.8095 0.8182 0.7647 205 6 16 0.9375 0.7589 0.7700 0.7588232 7 15 0.9333 0.7083 0.7218 0.5778232 8 14 0.9286 0.6577 0.6737 0.5778 232 9 13 0.9231 0.6071 0.6255 0.5706 233 10 12 0.9167 0.5565 0.5773 0.5706233 11 11 0.9091 0.5059 0.5292 0.5706233 12 10 0.9000 0.4553 0.4810 0.5706239 13 9 0.8888 0.4047 0.4320 0.5271240 14 8 0.8750 0.3541 0.3847 0.5198 261 15 7 0.8571 0.3035 0.3365 0.3697 280 16 6 0.8333 0.2529 0.2883 0.2489280 17 5 0.8000 0.2023 0.2402 0.2489296 18 4 0.7500 0.1517 0.1920 0.1660296 19 3 0.6667 0.1011 0.1439 0.1660323 20 2 0.5000 0.0506 0.0958 0.0710 344/p59— 1 1.0000 0.0506 0.0958 0.0313 /p63Insteadof( n/p57r)/(n/p57r/p591), Pikeuses( n/p57r/p591)/(n/p57r/p592) intheKaplan —Meierproduct- limitestimatetoavoid( n/p57r)/(n/p57r/p591)/p580. Source:Pike (1966 ).ReproducedwithpermissionoftheBiometricSociety. estimate. The smooth curves are obtained from the Weibull fits. Table 6.3 showsthecalculationoftheplottingpointsforgroup2. It is obvious that the Weibull distributions with G/p58100, /afii9828/p583, /afii9838/p19/p16/p58(4.51/p5910/p92/p22)/p16/p30/p18, and /afii9838/p19/p16/p58(2.38/p5910/p92/p22)/p16/p30/p18fitthecarcinoma-freetimeof thetwogroupsofratsverywell. 6.3 LOGNORMAL DISTRIBUTION Initssimplestformthelognormaldistributioncanbedefinedasthedistribu- tionofavariablewhoselogarithmfollowsthenormaldistribution.Itsorigin maybetracedasfarbackas1879,whenMcAlister (1879 )describedexplicitly atheoryofthedistribution.Mostofitsaspectshavesincebeenunderstudy.Gaddum (1945a,b)gave a review of its application in biology, followed by  143 Figure 6.8Hazardofthelognormaldistributionwithdifferentparameters.Boag’s (1949 )applicationsincancerresearch.Itshistory,properties,estimation problems,andusesineconomicshavebeendiscussedindetailbyAitchisonandBrown (1957 ).Later,otherinvestigatorsalsoobservedthattheageatonsetof Alzheimer’sdiseaseandthedistributionofsurvivaltimeofseveraldiseasessuchasHodgkin’sdisease and chronic leukemia couldbe rather closelyapproxi-matedbyalognormaldistributionsincetheyaremarkedlyskewedtotherightandthelogarithmsofsurvivaltimesareapproximatelynormallydistributed. Considerthesurvivaltime Tsuchthatlog Tisnormallydistributedwith mean /afii9839andvariance /afii9846/p17.Wethensaythat Tislognormallydistributedand writeTas/afii9806(/afii9839,/afii9846/p17).Itshouldbenotedthat /afii9839and/afii9846/p17arenotthemeanand varianceofthelognormaldistribution.Figure6.8givesthehazardfunctionofthelognormaldistributionwithdifferentvaluesfortheparameters.Thehazardfunctionincreasesinitiallytoamaximumandthendecreases (almostassoon asthemedianispassed )tozeroastimeapproachesinfinity (WatsonandWells 1961 ). Therefore, the lognormal distribution is suitable for survival patterns withaninitiallyincreasingandthendecreasinghazardrate.Byacentrallimittheorem,itcanbeshownthatthedistributionoftheproductof nindependent positive variates approaches a lognormal distribution under very generalconditions: for example, the distribution of the size of an organism whosegrowth is subject to many small impulses, the effect of each of which isproportionaltothemomentarysizeoftheorganism. Thepopularityofthelognormaldistributionisdueinparttothefactthat the cumulative values of y/p58logtcan be obtained from the tables of the standardnormaldistributionandthecorrespondingvaluesof tarethenfound bytakingantilogs.Thus,thepercentilesofthelognormaldistributionareeasytofind.144 -    Theprobabilitydensityfunctionandsurvivorshipfunctionare,respectively, f(t)/p581 t/afii9846/p402/afii9843exp/p3/p571 2/afii9846/p17(logt/p57/afii9839)/p17/p4t/p570,/afii9846/p572 (6.3.1 ) and S(t)/p581 /afii9846/p402/afii9843/p16/p27 /p821 xexp/p3/p571 2/afii9846/p17(logx/p57/afii9839)/p17/p4dx (6.3.2 ) Leta/p58exp(/p57/afii9839).Then /p57/afii9839/p58loga,(6.3.1 )and (6.3.2 )canbewrittenas f(t)/p581 t/afii9846/p402/afii9843exp/p3/p571 2/afii9846/p17(logat)/p17/p4(6.3.3) and S(t)/p581 /afii9846/p402/afii9843/p16/p27 /p821 xexp/p3/p571 2/afii9846/p17(logax)/p17/p4dx(6.3.4) /p581/p57G/p1logat /afii9846/p2(6.3.5) whereG(y) is the cumulative distribution function of a standard normal variable G(y)/p581 /p402/afii9843/p16/p87 /p15e/p92/p83/p130/p30/p17du (6.3.6) Thelognormaldistributionisspecifiedcompletelybythetwoparameters /afii9839 and/afii9846/p17.TimeTcannotassumezerovaluessincelog Tisnotdefinedfor T/p580. Figure6.9givesthelognormalfrequencycurvesfor /afii9839/p580,/afii9846/p17/p580.1,0.5,2,from which an idea of the flexibility of the distribution may be obtained. It isobviousthatthedistributionispositivelyskewedandthatthegreaterthevalueof/afii9846/p17, the greater the skewness. Figure 6.10 shows the frequency curves for /afii9846/p17/p580.5,/afii9839/p580, 0.5, 1. It is obvious that /afii9839and/afii9846/p17are, respectively, scale parametersand notlocationandscaleparametersasinthenormaldistribu- tion.Thehazardfunction,from (6.3.3 )and (6.3.5 ),hastheform h(t)/p58(1/t/afii9846/p402/afii9843 )exp[ /p57(logat)/p17/2/afii9846/p17] 1/p57G(logat//afii9846)(6.3.7 ) andisplottedinFigure6.8.  145 Figure 6.9Lognormaldensitycurveswith /afii9839/p580. Figure 6.10Lognormaldensitycurveswith /afii9846/p17/p580.5.The meanand variance of the two-parameterlognormaldistribution are, respectively, exp (/afii9839/p59/p16/p17/afii9846/p17)and [exp (/afii9846/p17)/p571]exp (2/afii9839/p59/afii9846/p17). The coefficient of variationofthedistributionisthen[exp (/afii9846/p17)/p571]/p16/p30/p17.Themedianis e/p73andthe modeisexp (/afii9839/p59/afii9846/p17). The two-parameter lognormal distribution can also be generalized to a three-parameter distribution by replacing twitht/p57Gin(6.3.1 ). In other words,thesurvivaltimelog( T/p57G)followsthenormaldistributionwithmean /afii9839andvariance /afii9846/p17.Incertainsituationsthevalueof Gmaybedetermineda priori and should not be regarded as an unknown parameter that requiresestimation.Ifthisisso,thevariable T/p57Gmaybeconsideredinplaceof T and the distribution of T/p57Ghas all the properties of the two-parameter lognormaldistribution.However,theestimationproceduresdevelopedforthe two-parametercasearenotdirectlyapplicabletothedistributionof T/p57G.146 -    Figure 6.11Lognormalprobabilityplotofthesurvivaltimeof234malepatientswith chroniclymphocyticleukemia. (FromFeinleibandMacMahon,1960.Reproducedby permissionofthepublisher. ) Example 6.3 Inastudyofchroniclymphocyticandmyelocyticleukemia, FeinleibandMacMahon (1960 )appliedthelognormaldistributiontoanalyze survivaldataof649whiteresidentsofBrooklyndiagnosedfrom1943to1952.Theanalysisofseveralsubgroupsofpatientsfollows.Thesurvivaltimeofeachpatientiscomputedfromthedateofdiagnosisinmonths.Analyticalmethodisusedtofitthelognormaldistributiontothedata.ThemethodisdiscussedinChapters7and8. Figure 6.11 gives the probability plot of the survival time of 234 male patientswithchroniclymphocyticleukemia,inwhichthehorizontalaxisfor the survival time is in logarithmic scale and the vertical axis is in normalprobabilityscale.Whenplotting1 /p57S(t) onthisgraphpaper,astraightlineis obtained when the data follow a two-parameter lognormal distribution. Aninspection of the graph shows that the distribution is concave. Gaddum(1945a,b)haspointedoutthatsuchadeviationcanbecorrectedbysubtracting an appropriate constant from the survival times. In other words, the three-parameterlognormaldistributioncanbeused.Figure6.12givesasimilarplotinwhichthesurvivaltimeofeverypatientplus4isplotted.Theconfigurationis linear and hence empiricallyit seems valid to assume that the lognormaldistributionisappropriate. Similargraphsformalepatientswithchronicmyelocyticleukemiaandfor femalepatientswithchroniclymphocyticormyelocyticleukemiaaregiveninFigures6.13and6.14.Parametersofthelognormaldistributionareestimated.FeinleibandMacMahonreportthattheagreementbetweentheobservedandcalculated distributions is striking for each group except for women withchroniclymphocyticleukemia.Thecorresponding pvaluesforthechi-square  147 Figure 6.12Lognormalprobabilityplotofthesurvivaltimeinmonthsplus4of234 male patients with chronic lymphocytic leukemia. (From Feinleib and MacMahon, 1960.Reproducedbypermissionofthepublisher. ) goodness-of-fittestareasfollows: ChronicMyelocytic ChronicLymphocytic Male 0.86 0.73 Female 0.57 0.016 Since a large pvalue indicates close agreement, it is concluded that the three-parameterlognormaldistribution adequatelydescribesthe distributionofsurvivaltimesforeachsubgroupexceptwomenwithchroniclymphocyticleukemia.Theshapeoftheobserveddistributionforthelattergroupsuggeststhat itmight actuallybe composedof two dissimilargroups,each of whosesurvivaltimesmightfitalognormaldistribution. 6.4 GAMMA AND GENERALIZED GAMMA DISTRIBUTIONS The gamma distribution, which includes the exponential and chi-square distribution,wasusedalongtimeagobyBrownandFlood (1947 )todescribe the life of glass tumblers circulating in a cafeteria and by Birnbaum andSaunders (1958 )asastatisticalmodelforlifelengthofmaterials.Sincethen, thisdistributionhasbeenusedfrequentlyasamodelforindustrialreliabilityproblemsandhumansurvival.148 -    Figure 6.13Lognormalprobabilityplotofthesurvivaltimeinmonthsplus4of162 malepatientswithchronicmyelocyticleukemia. (FromFeinleibandMacMahon,1960. Reproducedbypermissionofthepublisher. ) Figure 6.14Lognormalprobabilityplotofthesurvivaltimeinmonthsplus4offemale patientswithtwotypesofleukemia. (FromFeinleibandMacMahon,1960.Reproduced bypermissionofthepublisher. )Suppose that failure or death takes place in nstages or as soon as n subfailureshavehappened.Attheendofthefirststage,aftertime T/p16,thefirst subfailureoccurs;afterthatthesecondstagebeginsandthesecondsubfailureoccursaftertime T/p17;andsoon.Totalfailureordeathoccursattheendofthe nth stage, when the nth subfailure happens. The survival time, T, is then T/p16/p59T/p17/p59/p37/p59T/p76.Thetimes T/p16,T/p17,...,T/p76spentineachstageareassumedto     149 Figure 6.15Gammahazardfunctionswith /afii9838/p581.be independently exponentially distributed with probability density function /afii9838exp(/p57/afii9838t/p71),i/p581,...,n. That is, the subfailures occur independently at a constantrate /afii9838.Thedistributionof Tisthencalledthe Erlangian distribution . There is no need for the stages to have physical significance since we canalways assume that death occurs in the n-stage process just described. This idea, introduced by A. K. Erlang in his study of congestion in telephonesystems,hasbeenusedwidelyinqueuingtheoryandlifeprocesses. A natural generalization of the Erlangian distribution is to replace the parameter nrestrictedtotheintegers1,2,...byaparameter /afii9828takinganyreal positivevalue.Wethenobtainthe gamma distribution. Thegammadistributionischaracterizedbytwoparameters, /afii9828and/afii9838.When 0/p58/afii9828/p581,thereisnegativeagingandthehazardratedecreasesmonotonically from infinity to /afii9838as time increases from 0 to infinity. When /afii9828/p571, there is positiveagingandthehazardrateincreasesmonotonicallyfrom0to /afii9838astime increasesfrom0toinfinity.When /afii9828/p581,thehazardrateequals /afii9838,aconstant, asintheexponentialcase.Figure6.15illustratesthegammahazardfunctionfor/afii9838/p581 and /afii9828/p581,/afii9828/p581, 2, 4. Thus, the gamma distribution describes a different type of survival pattern where the hazard rate is decreasing orincreasingtoaconstantvalueastimeapproachesinfinity. Theprobabilitydensityfunctionofagammadistributionis f(t)/p58/afii9838 /afii9772(/afii9828) (/afii9838t)/p65/p92/p16e/p92/p72/p82t/p570,/afii9828/p570,/afii9838/p570 (6.4.1 ) where /afii9772(/afii9828)is defined as in (6.2.9 ). Figures 6.16 and 6.17 show the gamma densityfunctionwithvariousvaluesof /afii9828and/afii9838.Itisseenthatvarying /afii9828changes the shape of the distribution while varying /afii9838changes only the scaling. Consequently, /afii9828and/afii9838are shape and scale parameters, respectively. When /afii9828/p571,thereisasinglepeakat t/p58(/afii9828/p571)//afii9838.150 -    Figure 6.16Gammadensityfunctionswith /afii9838/p581. Figure 6.17Gammadensityfunctionswith /afii9828/p583.Thecumulativedistributionfunction F(t) hasamorecomplexform: F(t)/p58/p16/p82 /p15/afii9838 /afii9772(/afii9828)(/afii9838x)/p65/p92/p16e/p92/p72/p86dx (6.4.2) /p581 /afii9772(/afii9828)/p16/p72/p82 /p15u/p65/p92/p16e/p92/p83du /p58I(/afii9838t,/afii9828)( 6 .4.3) where I(s,/afii9828)/p581 /afii9772(/afii9828)/p16/p81 /p15u/p65/p92/p16e/p92/p83du (6.4.4) knownasthe incomplete gamma function ,istabulatedinPearson (1922,1957 ).     151 FortheErlangiandistribution,itcanbeshownthat F(t)/p581/p57/p76/p92/p16/p26 /p73/p14/p15e/p92/p72/p82(/afii9838t)/p73 k!(6.4.5 ) Thus,thesurvivorshipfunction1 /p57F(t)i s S(t)/p58/p16/p27 /p82/afii9838 /afii9772(/afii9828)(/afii9838x)/p65/p92/p16e/p92/p72/p86dx (6.4.6) forthegammadistributionor S(t)/p58e/p92/p82/p76/p92/p16/p26 /p73/p14/p15(/afii9838t)/p73 k!(6.4.7 ) fortheErlangiandistribution. Sincethehazardfunctionistheratioof f(t)toS(t),itcanbecalculatedfrom (6.4.1 )and (6.4.7 ).When /afii9828isaninteger n, h(t)/p58/afii9838(/afii9838t)/p76/p92/p16 (n/p571)!/afii9814/p76/p92/p16/p73/p14/p15(1/k!)(/afii9838t)/p73(6.4.8) When /afii9828/p581,thedistributionisexponential.When /afii9838/p58/p16/p17and/afii9828/p58/p16/p17/afii9840,where /afii9840 isaninteger,thedistributionischi-squarewith /afii9840degreesoffreedom.Themean andvarianceofthestandardgammadistributionare,respectively, /afii9828//afii9838and/afii9828//afii9838/p17, sothatthecoefficientofvariationis1/ /p40/afii9828. Manysurvivaldistributionscanberepresented,atleastroughly,bysuitable choiceoftheparameters /afii9838and/afii9828.Inmanycases,thereisanadvantageinusing theErlangiandistribution,thatis,intaking /afii9828integer. Theexponential,Weibull,lognormal,andgammadistributionsarespecial casesofageneralizedgammadistributionwiththreeparameters, /afii9838,/afii9825,and /afii9828, whosedensityfunctionisdefinedas f(t)/p58/afii9825/afii9838/p63/p65 /afii9772(/afii9828)t/p63/p65/p92/p16exp[ /p57(/afii9838t)/p63]t/p570,/afii9828/p570,/afii9838/p570,/afii9825/p570( 6.4.9) It is easily seen that this generalized gamma distribution is the exponential distribution if /afii9825/p58/afii9828/p581, the Weibull distribution if /afii9828/p581; the lognormal distributionif /afii9828/p59/p45,andthegammadistributionif /afii9825/p581. In later chapters (e.g., Chapters 7 and 9 ), we discuss several parametric proceduresforestimationandhypothesistesting.TouseavailablecomputersoftwaresuchasSAStocarryoutthecomputation,weusethedistributionsadoptedbythesoftware.OneoftheveryfewsoftwarepackagesthatincludethegammaorgeneralizedgammadistributionisSAS.InSAS,thegeneralized152 -    Table 6.4 Lifetimes of 101 Strips of Aluminum Coupon 370 706 716746785797 844 855 858886886930960 988 990 1000101010161018 1020 10551085 1102 1102110811151120 1134 1140 11991200120012031222 1235 12381252125812621269 1270 12901293 1300 1310131313151330 1355 1390 14161419142014201450 1452 14751478148114851502 1505 15131522 1522 1530154015601567 1578 1594 16021604160816301642 1674 17301750175017631768 1781 17821792 1820 1868188118901893 1895 1910 19231940194520232100 2130 221522682440 Source:BirnbaumandSaunders (1958 ). gammadistributionisdefinedashavingthefollowingdensityfunction: f(t)/p58/p34/afii9825/p34/afii9828/p65/afii9838/p63/p65 /afii9772(/afii9828)t/p63/p65/p92/p16exp[ /p57/afii9828(/afii9838t)/p63],t/p570,/afii9828/p570,/afii9838/p570( 6.4.10) To differentiatethisformof thegeneralizedgammadistributionfromthe generalized gamma in (6.4.9 ), we refer to this distribution as the extended generalized gamma distribution .Itcanbeshownthattheextendedgeneralized gammadistributionreducestotheWeibulldistributionwhen /afii9825/p570and /afii9828/p581, thelognormaldistributionwhen /afii9828/p59/p45,thegammadistributionwhen /afii9825/p581, andtheexponentialdistributionwhen /afii9825/p58/afii9828/p581. Example 6.4 BirnbaumandSaunders (1958 )reportanapplicationofthe gammadistributiontothelifetimeofaluminumcoupon.Intheirstudy,17setsofsixstripswereplacedinaspeciallydesignedmachine.Periodicloadingwasappliedto the stripswith a frequencyof 18 hertzand a maximumstress of21,000poundspersquareinch.The102stripswererununtilallofthemfailed.One of the 102 strips tested had to be discarded for an extraneous reason,yielding101observations.ThelifetimedataaregiveninTable6.4inascendingorder. From the data the two parameters of the gamma distribution were     153 Figure 6.18Graphical comparison of observed and fitted cumulative distribution functionsofdatainExample6.4. (FromBirnbaumandSaunders,1958. ) estimated (estimation methods are discussed in Chapter 7 ). They obtained /afii9828/p24/p5811.8and /afii9838/p19/p581/(118.76/p5910/p18). Agraphicalcomparisonoftheobservedandfittedcumulativedistribution function is given in Figure 6.18, which shows very good agreement. Achi-squaregoodness-of-fittest (discussedinChapter9 )yieldeda /afii9851/p17valueof 4.49for6degreesoffreedom,correspondingtoasignificancelevelbetween0.5 and0.6.Thus,itwasconcludedthatthegammadistributionwasanadequatemodelforthelifelengthofsomematerials. 6.5 LOG-LOGISTIC DISTRIBUTION The survival time Thas a log-logistic distribution if log( T) has a logistic distribution.The density, survivorship,hazard, and cumulative hazard func-tionsofthelog-logisticdistributionare,respectively, f(t)/p58/afii9825/afii9828t/p65/p92/p16 (1/p59/afii9825t/p65)/p17 (6.5.1) S(t)/p581 1/p59/afii9825t/p65(6.5.2)154 -    h(t)/p58/afii9825/afii9828t/p65/p92/p16 1/p59/afii9825t/p65(6.5.3) H(t)/p58log(1/p59/afii9825t/p65)( 6 .5.4) t/p460,/afii9825/p570,/afii9828/p570 Thelog-logisticdistributionischaracterizedbytwoparameters /afii9825,and /afii9828.The medianofthelog-logisticdistributionis /afii9825/p92/p16/p30/p65.Figure6.19 (a)to(c)showthe log-logistichazard,density,andsurvivorshipfunctionswith /afii9825/p581andvarious valuesof /afii9828/p582.0,1,and0.67. When /afii9828/p571,thelog-logistichazardhasthevalue0attime0,increasestoa peakat t/p58(/afii9828/p571)/p16/p30/p65//afii9825/p16/p30/p65,andthendeclines,whichissimilartothelognormal hazard.When /afii9828/p581,thehazardstartsat /afii9825/p16/p30/p65andthendeclinesmonotonically. When /afii9828/p581,thehazardstartsatinfinityandthendeclines,whichissimilarto the Weibull distribution. The hazard function declines toward 0 as tap- proachesinfinity.Thus,thelog-logisticdistributionmaybeusedtodescribeafirst increasing and then decreasing hazard or a monotonically decreasinghazard. Example 6.5 Byers et al. (1988 )used the log-logistic distribution to describetherateofspreadofHIVbetween1978and1986.Between1978and1980,over6700homosexualandbisexualmeninSanFranciscowereenrolledinstudiesoftheprevalenceandincidenceofsexuallytransmittedhepatitisBvirus (HBV )infections.Bloodspecimenswerecollectedfromtheparticipants. Four hundred and eighty-eight men who were HBV-seronegative were ran-domly selected to participate in a study of HIV infection later. These menagreed to allow the investigators to test the specimens collected previouslytogether with a current specimen. For those who convert to positive, theinfectiontimeisonlyknowntohaveoccurredbetweenthepreviousnegativetestandthetimeofthefirstpositiveone.Theexacttimeisunknown.Thetimetoinfectionisthereforeintervalcensored.Theinvestigatorstriedtofitseveraldistributions to the interval-censored data, including the Weibull and log-logisticbymaximumlikelihoodmethods (discussedinChapter7 ).Basedon the Akaike information criterion (discussed in Chapter 9 ), the log-logistic distribution was found to provide the best fit to the data. The maximumlikelihoodestimatesofthetwoparametersare /afii9825/p24/p580.003757and /afii9828/p24/p581.424328. Basedonthelog-logisticmodel,themedianinfectiontimeisestimatedtobe50.4months,andthehazardfunctionapproachesitspeakat27.6months. 6.6 OTHER SURVIVAL DISTRIBUTIONS Manyotherdistributionscanbeusedasmodelsofsurvivaltime,threeofwhich wediscussbrieflyinthissection:thelinearexponential,theGompertz (1825 ),   155 (a) (b) (c) Figure 6.19(a) Hazardfunctionofthelog-logisticdistribution;( b) densityfunctionof thelog-logisticdistribution;( c) Survivorshipfunctionofthelog-logisticdistribution. 156 Figure 6.20Hazardfunctionoflinear-exponentialmodel.andadistributionwhosehazardrateisastepfunction.Thelinear-exponential model and the Gompertz distribution are extensions of the exponentialdistribution.Bothdescribesurvivalpatternsthathaveaconstantinitialhazardrate. The hazard rate varies as a linear function of time or age in thelinear-exponentialmodelandasanexponentialfunctionoftimeorageintheGompertzdistribution. Indemonstratingtheuseofthelinear-exponentialmodel,Broadbent (1958 ), uses as an example the service of milk bottles that are filled in a dairy,circulatedtocustomers,andreturnedemptytothedairy.ThemodelwasalsousedbyCarboneetal. (1967 )todescribethesurvivalpatternofpatientswith plasmacyticmyeloma.Thehazardfunctionofthelinear-exponentialdistribu-tionis h(t)/p58/afii9838/p59/afii9828t (6.6.1) where /afii9838and/afii9828can be values such that h(t) is nonnegative.The hazardrate increasesfrom /afii9838withtimeif /afii9828/p570,decreasesif /afii9828/p580,andremainsconstant (an exponentialcase )if/afii9828/p580,asdepictedinFigure6.20. Theprobabilitydensityfunctionandthesurvivorshipfunctionare,respec- tively, f(t)/p58(/afii9838/p59/afii9828t)exp[ /p57(/afii9838t/p59/p16/p17 /afii9828t/p17)] (6 .6.2) and S(t)/p58exp[ /p57(/afii9838t/p59/p16/p17/afii9828t/p17)] (6 .6.3) Themeanofthelinear-exponentialdistributionis /p57(/afii9838//afii9828)/p59(/afii9828/2)/p92/p16/p30/p17L(/afii9838/p17/2/afii9828), where L(x)/p58e/p86/p16/p27 /p86y/p16/p30/p17e/p92/p87dy   157 Table 6.5 Values of L(x) and G(x) xL (x) G(x) 00 .886 /p45 0.10 .951 2 .015 0.21 .012 1 .493 0.31 .067 1 .223 0.41 .119 1 .048 0.51 .168 0 .923 0.61 .214 0 .828 0.71 .258 0 .753 0.81 .300 0 .691 0.91 .341 0 .640 11 .381 0 .596 21 .712 0 .361 31 .987 0 .262 Source:Broadbent (1958 ). Figure 6.21Gompertzhazardfunction.istabulatedinTable6.5.Aspecialcaseofthelinear-exponentialdistribution, theRayleighdistribution,isobtainedbyreplacing /afii9828by/p16/p17/afii9828(Kodlin,1967 ).That is,thehazardfunctionoftheRayleighdistributionis h(t)/p58/afii9838/p59/p16/p17/afii9828t. TheGompertzdistributionisalsocharacterizedbytwoparameters, /afii9838and /afii9828.Thehazardfunction, h(t)/p58exp(/afii9838/p59/afii9828t)( 6 .6.4) isplottedinFigure6.21.When /afii9828/p570,thereispositiveagingstartingfrom e/p72; when /afii9828/p580,thereisnegativeaging;andwhen /afii9828/p580,h(t)reducestoaconstant, e/p72.ThesurvivorshipfunctionoftheGompertzdistributionis S(t)/p58exp/p3/p57e/p72 /afii9828(e/p65/p82/p571)/p4(6.6.5)158 -    Figure 6.22Stephazardfunction.andtheprobabilitydensityfunction,from (6.6.4 )and (2.2.5 ),isthen f(t)/p58exp/p3(/afii9838/p59/afii9828t)/p571 /afii9828(e/p72/p62/p65/p82/p57e/p72)/p4(6.6.6) ThemeanoftheGompertzdistributionis G(e/p72//afii9828)/e/p72, where G(x)/p58e/p86/p16/p27 /p86y/p92/p16e/p92/p87dy istabulatedinTable6.5. Finally,weconsideradistributionwherethehazardrateisastepfunction: h(t)/p58/p7a/p16 a/p17 /p36 a/p73/p92/p16 a/p730/p45t/p58t/p16 t/p16/p45t/p58t/p17 t/p73/p92/p17/p45t/p58t/p73/p92/p16 t/p46t/p73/p92/p16(6.6.7) wheret/p16,t/p17,...,t/p73aredifferenttimepoints.Figure6.22showsatypicalhazard functionofthisnaturefor k/p585.Using (2.2.4 ),thesurvivorshipfunctioncan bederived: S(t)/p58/p7exp(/p57a/p16t)0 /p45t/p58t/p16 exp[ /p57a/p16t/p16/p57a/p17(t/p57t/p16)] t/p16/p45t/p58t/p17 /p36 exp[ /p57a/p16t/p16/p57a/p17(t/p17/p57t/p16)/p57/p37/p57a/p73(t/p57t/p73/p92/p16)t/p46t/p73/p92/p16(6.6.8)   159 Theprobabilitydensityfunction f(t) canthen be obtained from (6.6.7 )and (6.6.8 )using (2.2.5 ): f(t)/p58/p7a/p16exp(/p57a/p16t)0 /p45t/p58t/p16 a/p17exp[ /p57a/p16t/p16/p57a/p17(t/p57t/p16)] t/p16/p45t/p58t/p17 /p36 a/p73exp[ /p57a/p16t/p16/p57a/p17(t/p17/p57t/p16)/p57/p37/p57a/p73(t/p57t/p73/p92/p16)t/p46t/p73/p92/p16(6.6.9) One application of this distribution is the life-table analysis discussed in Chapter4.Inalife-tableanalysis,timeisdividedintointervalsandthehazardrateisassumedtobeconstantineachinterval.However,theoverallhazardrateisnotnecessarilyconstant. The nine distributions described above are, among others, reasonable modelsforsurvivaltimedistribution.Allhavebeendesignedbyconsideringabiologicalfailure,adeathprocess,oranagingproperty.Theymayormaynotbe appropriate for many practical situations, but the objective here is toillustratethevariouspossibletechniques,assumptions,andargumentsthatcanbeusedtochoosethemostappropriatemodel.Ifnoneofthesedistributionsfitsthedata,investigatorsmighthavetoderiveanoriginalmodeltosuittheparticulardata,perhapsbyusingsomeoftheideaspresentedhere. Bibliographical Remarks Inadditiontothepapersonthedistributionscitedinthischapter,Mannetal. (1974 ), Hahn and Shapiro (1967 ), Johnson and Kotz (1970a,b), Elandt- JohnsonandJohnson (1980 ),Lawless (1982 ),Nelson (1982 ),CoxandOakes (1984 ), Gertsbakh (1989 ), and Klein and Moeschberger (1997 )also discuss statisticalfailuremodels,includingtheexponential,Weibull,gamma,lognor-mal,generalizedgamma,andlog-logisticdistributions.Applicationsofsurvivaldistributionscanbefoundeasilyinmedicalandepidemiologicaljournals.Thefollowing are a few examples: Dharmalingam et al. (2000 ), Riffenburgh and Johnstone (2001 ),andMafartetal. (2002 ). EXERCISES 6.1Summarize the distributions discussed in this chapter, answering the followingquestions. (a)Whatdistributionsdescribeconstanthazardrates?Givetherangeof parametervalues. (b)Whatdistributionsdescribeincreasinghazardrates?Iftherearemore thanone,discussthedifferencesbetweenthem. (c)Whatdistributionsdescribedecreasinghazardrates?Iftherearemore thanone,discussthedifferencesbetweenthem.160 -    6.2Supposethatthesurvivaldistributionofagroupofpatientsfollowsthe exponentialdistributionwith G/p580(year),/afii9838/p580.65. Plot the survivor- shipfunctionandfind: (a)Themeansurvivaltime (b)Themediansurvivaltime (c)Theprobabilityofsurviving1.5yearsormore 6.3Supposethatthesurvivaldistributionofagroupofpatientsfollowsthe exponentialdistributionwith G/p585(years )and/afii9838/p580.25.Plotthesurviv- orshipfunctionandfind: (a)Themeansurvivaltime (b)Themediansurvivaltime (c)Theprobabilityofsurviving6yearsormore 6.4ConsiderthefollowingtwoWeibulldistributionsassurvivalmodels: (i)G/p580,/afii9838/p581,/afii9828/p580.5 (ii)G/p580,/afii9838/p580.5,/afii9828/p582 For each distribution, plot the survivorship function and the hazard functionandfind: (a)Themean (b)Thevariance (c)Thecoefficientofvariation Whichdistributiongivesthelargerprobabilityofsurvivingatleast3units oftime? 6.5Suppose that the survival timefollows the lognormaldistribution with /afii9839/p581and /afii9846/p580.5.Find: (a)Themeansurvivaltime (b)Thevariance (c)Thecoefficientofvariation (d)Themedian (e)Themode 6.6Supposethatpainrelieftimefollowsthegammadistributionwith /afii9838/p581, /afii9828/p580.5.Find: (a)Themean (b)Thevariance (c)Thecoefficientofvariation 6.7Suppose that the survival distribution is (1)Gompertz and (2)linear- exponential,and /afii9838/p581,/afii9828/p582.0.Plotthehazardfunctionandfind: (a)Themean (b)Theprobabilityofsurvivinglongerthan1unitoftime 6.8ConsiderthesurvivaltimesofhypernephromapatientsgiveninExercise Table 3.1. From the plot you obtained in Exercise 4.5, suggest adistributionthatmightfitthedata. 161 CHAPTER 7 Estimation Procedures for Parametric Survival Distributionswithout Covariates In this chapter we discuss some analytical procedures for estimatingthe mostcommonlyusedsurvival distributionsdiscussed in Chapter6. We introducethemaximumlikelihoodestimates (MLEs )oftheparametersof thesedistributions. The general asymptotic likelihood inference results that are most widely usedfor these distributions are given in Section 7.1. We begin to used the generalsymbol b/p58(b/p16,b/p17,...,b/p78) to denote a set of parameters. For example, in discussingthe Weibull distribution, b/p16could be /afii9838andb/p17could be /afii9828, andp/p582. bis called avectorin linear algebra. Readers who are not familiar with linear algebra or are not interested in the mathematical details may skip this sectionand proceed to Section 7.2 without loss of continuity. In Sections 7.2 to 7.7 weintroducethe MLEs for the parameters of the exponential, Weibull, lognormal,gamma, log-logistic, and Gompertz distributions for data with and withoutcensored observations. The related BMDP or SAS programming codes thatmay be used to obtain the MLE are given in the respective sections. 7.1 GENERAL MAXIMUM LIKELIHOOD ESTIMATION PROCEDURE 7.1.1 Estimation Procedures for Data with Right-Censored Observations Suppose that persons were followed to death or censored in a study. Let t/p16, t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76be the survival times observed from the nindividuals, withrexact times and (n/p57r)right-censored times. Assume that the survival times follow a distribution with the density function f(t,b)and survivorship functionS(t,b), where b/p58(b/p16,...,b/p78) denotes unknown pparameters b/p16,...,b/p78in the distribution. As shown in Chapter 6, an exponential distribu- tion has one (p/p581)parameter /afii9838, the Weibull distribution has two (p/p582) 162 parameters /afii9838and /afii9828, and so on. If the survival time is discrete (i.e., it is observed atdiscretetime only ),f(t,b)representstheprobabilityof observing tandS(t,b) represents the probability that the survival or event time is greater than t.I n other words, f(t,b)andS(t,b)represent the information that can be obtained from an observed uncensored survival time and an observed right-censoredsurvival time, respectively. Therefore, the product /afii9811/p76/p71/p14/p16f(t/p71,b)represents the joint probability of observingthe uncensored survival times, and/afii9811/p76/p71/p14/p80/p62/p16S(t/p62/p71,b)represents the joint probability of those right-censored survival times. The product of these two probabilities, denoted by L(b), L(b)/p58/p80/p147 /p71/p14/p16f(t/p71,b)/p76/p147 /p71/p14/p80/p62/p16S(t/p62/p71,b) represents the joint probability of observing t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76. A similar interpretation applies to continuous survival. L(b)is called the likelihood functionofb, which can also be interpreted as a measure of the likelihood of observinga specific set of survival times t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76, given a specific set of parameters b. The method of the MLE is to find an estimator of bthat maximizes L(b), or in other words, which is ‘‘most likely’’ to have produced the observed data t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76. Take the logarithm of L(b)and denote it by l(b), l(b)/p58logL(b)/p58/p80/p26 /p71/p14/p16log[f(t/p71,b)]/p59/p76/p26 /p71/p14/p80/p62/p16log[S(t/p62/p71,b)] (7.1.1 ) Then the MLE b/p19ofbis the set ofb/p19/p16,...,b/p19/p78that maximizes l(b): l(b/p19)/p58max /p0/p12/p12 b(l(b)). It is clear that b/p19is a solution of the followingsimultaneous equations, which are obtained by takingthe derivative of l(b)with respect to each b/p72: /p42l(b) /p42b/p72/p580j/p581, 2,...,p (7.1.2) The exact forms of (7.1.2 )for the parametric survival distributions discussed in Chapter 6 are given in Sections 7.2 to 7.7. Often, there is no closed solution forthe MLE b/p19from (7.1.2 ). To obtain the MLE b/p19, one can use a numerical method. A commonly used numerical method is the Newton —Raphson iter- ative procedure, which can be summarized as follows. 1. Let the initial values of b/p16,...,b/p78be zero; that is, let b/p7/p15/p8 /p580     163 2. The changes for bat each subsequentstep, denoted by /afii9773/p7/p72/p8, is obtainedby takingthe second derivative of the log -likelihood function: /afii9773/p7/p72/p8/p58/p3/p57/p42/p17l(b/p7/p72/p92/p16/p8) /p42b/p42b/p30/p4/p92/p16/p42l(b/p7/p72/p92/p16/p8) /p42b(7.1.3 ) 3. Using /afii9773/p7/p72/p8, the value of b/p7/p72/p8atjth step is b/p7/p72/p8/p58b/p7/p72/p92/p16/p8 /p59 /afii9773/p7/p72/p8j/p581, 2,... The iteration terminatesat, say, the mth step if /p35/afii9773/p7/p75/p8/p35/p58/afii9829, where /afii9829is a given precision, usually a very small value, 10 /p92/p19or 10 /p92/p20. Then the MLE b/p19is defined as b/p19/p58b/p7/p75/p92/p16/p8 (7.1.4 ) The estimated covariance matrix of the MLE b/p19is given by V/p19(b/p19)/p58Co/p19v(b/p19)/p58/p3/p57/p42/p17l(b/p19) /p42b/p42b/p30/p4/p92/p16(7.1.5 ) One of the good properties of a MLE is that if b/p19is the MLE of b, theng(b/p19)is the MLE of g(b)ifg(b)is a finite function and need not be one-to-one. The concept of the Newton —Raphson method for p/p581 is illustrated in detail in Appendix A. The estimated 100 (1/p57/afii9825)% confidence interval for any parameter b/p71is (b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71,b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71)( 7.1.6 ) wherev/p71/p71is theith diagonal element of V/p19(b/p19)andZ/p63/p30/p17is the 100 (1/p57/afii9825/2) percentile point of the standard normal distribution [ P(Z/p57Z/p63/p30/p17)/p58/afii9825/2]. For a finite function g(b/p71)o fb/p71, the estimated 100 (1/p57/afii9825)% confidence interval for g(b/p71) is its respective range Ron the confidence interval (7.1.6 ), that is, R/p58/p43g(b/p71):b/p71/p43(b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71,b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71)/p44 (7.1.7 ) In caseg(b/p71) is monotone in b/p71, the estimated 100 (1/p57/afii9825)%confidence interval forg(b/p71)i s [g(b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71),g(b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71)] (7.1.8 )164       7.1.2 Estimation Procedures for Data with Right-, Left-, and Interval-Censored Observations If the survival times t/p16,t/p17,...,t/p76observed for the npersons consist of uncen- sored, left-, right-, and interval-censored observations, the estimation pro-ceduresaresimilar.Assumethatthesurvivaltimesfollowadistributionwiththedensityfunction f(t,b)and thesurvivorshipfunction S(t,b), wherebdenotesall unknown parameters of the distribution. Then the log-likelihood function is l(b)/p58logL(b)/p58/p26log[f(t/p71,b)]/p59/p26log[S(t/p71,b)] /p59/p26log[1 /p57S(t/p71,b)]/p59/p26log[S(v/p71,b)/p57S(t/p71,b)] (7.1.9 ) where the first sum is over the uncensored observations, the second sum over the right-censored observations, the third sum over the left-censored observa-tions, and the last sum over the interval-censored observations, with v/p71as the lower end of a censoringinterval. The other steps for obtainingthe MLE b/p19of bare similar to the steps shown in Section 7.1.1 by substitutingthe log - likelihood function defined in (7.1.1 )with the log-likelihood function in (7.1.9 ). The computation of the MLE b/p19and its estimated covariance matrix is tedious. The following example gives the general procedure for using SAS tocarry out the computation. Example 7.1 If the survival time observed contains uncensored, right-, left-, and interval-censored observations, one needs to create a new data setfrom the observed data to use SAS to obtain the estimates of the parametersin the distribution. For an observed survival time t(uncensored, right-, or left-censored ),wedefinetwovariablesLBand UBas follows:If tis uncensored, take LB /p58UB/p58t;i ftis left-censored, LB /p58. and UB /p58t; and iftis right-censored, then LB /p58tand UB /p58., where ‘‘.’’ means ‘‘missing’’ in SAS. If a survival time is interval-censored, [i.e., one observed two numbers t/p16andt/p17, t/p16/p58t/p17and the survivaltime is in theinterval (t/p16,t/p17)], let LB /p58t/p16and UB /p58t/p17. Assume that the new data set (in terms of LB and UB )has been saved in ‘‘C:/p33EXAMPLEA.DAT’’ as a text file, which contains two columns (LB in the first column and UB the second column )separated by a space. As an example, the followingSAS code can be used to obtain the estimated covariance matrix defined in (7.1.5 )and the MLE of the parameters of the Weibull distribution for the survival data observed in the text file ‘‘C: /p33EXAM- PLEA.DAT’’. One can replace d /p58weibull in the followingcode with the respective distribution in Sections 7.2 to 7.6 (see the SAS code in these sections for details )to obtain the estimate. data w1; infile ‘c: /p33examplea.dat’ missover; input lb ub; run; proc lifereg; model (lb,ub )/p58/covb d /p58weibull; run;     165 7.2 EXPONENTIAL DISTRIBUTION 7.2.1 One-Parameter Exponential Distribution The one-parameterexponential distribution has the followingdensity function; f(t)/p58/afii9838e/p92/p72/p82 (7.2.1) survivorship function; S(t)/p58e/p92/p72/p82 (7.2.2) and hazard function; h(t)/p58/afii9838 (7.2.3) wheret/p460,/afii9838/p570. Obviously, the exponential distribution is characterized by one parameter, /afii9838. The estimation of /afii9838by maximum likelihood methods for data without censoredobservationswill be given first followedby the case withcensored observations. Estimation of /afii9838for Data without Censored Observations Supposethat thereare npersons in thestudy and everyoneis followedto death or failure. Let t/p16,t/p17,...,t/p76be the exact survival times of the npeople. The likelihood function, using (7.2.1 )and (7.1.1 ),i s L/p58/p76/p147 /p71/p14/p16/afii9838e/p92/p72/p82/p71 and the log-likelihood function is l(/afii9838)/p58nlog/afii9838/p57/afii9838/p76/p26 /p71/p14/p16t/p71(7.2.4) From (7.1.2 ), the MLE of /afii9838is /afii9838/p19/p58n /afii9814/p76/p71/p14/p16t/p71 (7.2.5) Since the mean /afii9839of the exponential distribution is 1/ /afii9838and a MLE is invariant under an one-to-one transformation, the MLE of /afii9839is /afii9839/p24/p581 /afii9838/p19/p58/afii9814/p76/p71/p14/p16t/p71n/p58t/p16 (7.2.6 )166       It can be shown 2 n/afii9839/p19//afii9839has an exact chi-square distribution with 2 ndegrees of freedom (Epstein and Sobel, 1953 ). Since /afii9838/p581//afii9839and /afii9838/p19/p581//afii9839/p24, an exact 100(1/p57/afii9825)%confidence interval for /afii9838is /afii9838/p19 /afii9851/p17/p17/p76/p11/p16/p92/p63/p30/p172n/p58/afii9838/p58/afii9838/p19 /afii9851/p17/p17/p76/p11/p63/p30/p172n(7.2.7) where /afii9851/p17/p17/p76/p11/p63is the 100 /afii9825percentage point of the chi-square distribution with 2 n degrees of freedom, that is, P(/afii9851/p17/p17/p76/p57/afii9851/p17/p17/p76/p11/p63)/p58/afii9825(Table B-2 ). Whennis large (n/p4625, say ),/afii9838/p19is approximately normally distributed with mean /afii9838and variance /afii9838/p17/n. Thus, an approximate 100 (1/p57/afii9825)%confidence interval for /afii9838is /afii9838/p19/p57/afii9838/p19Z/p63/p30/p17 /p40n/p58/afii9838/p58/afii9838 /p19/p59/afii9838/p19Z/p63/p30/p17 /p40n(7.2.8) whereZ/p63/p30/p17is the 100 /afii9825/2 percentage point, P(Z/p57Z/p63/p30/p17)/p58/afii9825/2, of the standard normal distribution (Table B-1 ). Since 2n/afii9839/p24//afii9839has an exact chi-square distribution with 2 ndegrees of freedom, an exact 100 (1/p57a)% confidence interval for the mean survival time is 2n/afii9839/p24 /afii9851/p17/p17/p76/p11/p63/p30/p17/p58/afii9839/p582n/afii9839/p24 /afii9851/p17/p17/p76/p11/p16/p92/p63/p30/p17(7.2.9) The followingexample illustrates the procedures. Example 7.2 Consider the followingremission times in weeks for 21 patients with acute leukemia: 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 8, 8, 9, 10, 10, 12, 14, 16,20, 24, and 34. Assume that remission duration follows the exponentialdistribution. Let us estimate the parameter /afii9838by usingthe formulas g iven above. Accordingto (7.2.5 ), the MLE of the relapse rate, /afii9838,i s /afii9838/p19/p5821 198 /p580.106 per week The mean remission time /afii9839is then 198/21 /p589.429 weeks. Usingthe analytical procedures given above, confidence intervals for /afii9838and /afii9839can also be obtained. A 95% confidence interval for the relapse rate /afii9838, following (7.2.7 ),i s approximately (0.106 )(24.433 ) 42/p58/afii9838/p58(0.106 )(59.342 ) 42 or(0.062, 0.150 ). A 95% confidence interval for the mean remission time,  167 following (7.2.9 ),i s (42)(9.429 ) 59.342/p58/afii9839/p58(42)(9.429 ) 24.433 or(6.673, 16.208 ). Once the parameter /afii9838is estimated, other estimates can be obtained. For example,the probability of stayingin remission for at least 20 weeks, estimatedfrom (7.2.2 ),i sS/p19(20)/p58exp[/p570.106 (20)]/p580.120. Any percentile of survival timet/p78may be estimated by equating S(t)t opand solvingfor t/p78, that is, t/p78/p58/p57logp//afii9838/p19. For example, the median (50th percentile )survival time can be estimated by t/p15/p13/p20/p58/p57log0.5/ /afii9838/p19/p586.539 weeks. Estimation of /afii9838for Data with Censored Observations We first consider singly censored and then progressively censored data.Suppose that without loss of generality, the study or experiment begins at time0 with a total of nsubjects. Survival times are recorded and the data become available when the subjects die one after the other in such a way that theshortest survival time comes first, the second shortest second, and so on.Suppose that the investigator has decided to terminate the study after rout of thensubjects have died and to sacrifice the remaining n/p57rsubjects at that time. Then the survival times for the nsubjects are t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p58t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8 wherea superscriptplus indicates a sacrificedsubject,and thus t/p62/p7/p71/p8is a censored observation. In this case, nandrare fixed values and all of the n/p57rcensored observations are equal. The likelihood function, using (7.1.1 ),(7.2.1 ), and (7.2.2 ),i s L/p58n! (n/p57r)! /p80/p147 /p71/p14/p16/afii9838e/p92/p72/p82/p7/p71/p8(e/p92/p72/p82/p7/p80/p8)/p76/p92/p80 and from (7.1.2 ), the MLE of /afii9838is /afii9838/p19/p58r /afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p7/p71/p8(7.2.10) The mean survival time /afii9839/p581//afii9838can then be estimated by /afii9839/p24/p581 /afii9838/p19/p58/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p7/p71/p8r(7.2.11 ) It is shown by Halperin (1952 )that 2r/afii9838//afii9838/p19has a chi-square distribution with 2 r168       degrees of freedom. The mean and variance of /afii9838/p19arer/afii9838/(r/p571) and /afii9838/p17/(r/p571), respectively. The 100 (1/p57/afii9825)% confidence interval for /afii9838is /afii9838/p19/afii9851/p17/p17/p80/p11/p16/p92/p63/p30/p172r/p58/afii9838/p58/afii9838/p19 /afii9851/p17/p17/p80/p11/p63/p30/p172r(7.2.12) Whennis large, the distributionof /afii9838/p19is approximatelynormal with mean /afii9838and variance /afii9838/p17/(r/p571). An approximate 100/ (1/p57/afii9825)% confidence interval for /afii9838is then /afii9838/p19/p57/afii9838/p19Z/p63/p30/p17 /p40r/p571/p58/afii9838/p58/afii9838 /p19/p59/afii9838/p19Z/p63/p30/p17 /p40r/p571(7.2.13) Epstein and Sobel (1953 )show that 2r/afii9839/p24//afii9839has a chi-square distribution with 2rdegrees of freedom. Thus a 100/ (1/p57/afii9825)%confidence interval for /afii9839(see also Epstein, 1960b )is 2r/afii9839/p24 /afii9851/p17/p17/p80/p11/p63/p30/p17/p58/afii9839/p582r/afii9839/p24 /afii9851/p17/p17/p80/p11/p16/p92/p63/p30/p17(7.2.14) They also develop test procedures for the hypothesis H/p15:/afii9839/p58/afii9839/p15against the alternativeH/p16:/afii9839/p58/afii9839/p15. One of their rules of action is to accept H/p15if/afii9839/p24/p57cand rejectH/p15if/afii9839/p24/p58c, wherec/p58(/afii9839/p15/afii9851/p17/p17/p80/p11/p63)/2rand /afii9825is the significancelevel. Or if the estimated mean survival time calculated from (7.2.11 )is greater then c, the hypothesisH/p15is rejected at the /afii9825level. The followingexample illustrates the procedure. Example 7.3 Suppose that in a laboratory experiment 10 mice are exposed to carcinogens. The experimenter decides to terminate the study after half ofthemice are dead and to sacrifice the other halfat that time. The survival timesof the five dead mice are 4, 5, 8, 9, and 10 weeks. The survival data of the 10mice are 4, 5, 8, 9, 10, 10 /p59,1 0/p59,1 0/p59,1 0/p59, and 10 /p59. Assumingthat the failureof these mice follows an exponentialdistribution, the survival rate /afii9838and mean survival time /afii9839are estimated, respectively, accordingto (7.2.10 )and (7.2.11 )by /afii9838/p585 36/p5950 /p580.058 per week and /afii9839/p24/p581/0.058 /p5817.241 weeks. A 95%confidence interval for /afii9838by(7.2.12 )is (0.058 )(3.247 ) (2)(5)/p58/afii9838/p58(0.058 )(20.483 ) (2)(5)  169 or(0.019, 0.119 ). A 95% confidence interval for /afii9839following (7.2.13 )is 2(5)(17.241 ) 20.483/p58/afii9839/p582(5)(17.241 ) 3.247 or(8.417, 53.098 ). The probability of survivinga g iven time for the mice can be estimated from (7.2.2 ). For example, the probability that a mouse exposed to the same carcinogen will survive longer than 8 weeks is S/p19(8)/p58exp[/p570.058 (8)]/p580.629 The probability of dyingin 8 weeks is then 1 /p570.629 /p580.371. A slightly different situation may arise in laboratory experiments. Instead of terminatingthe study after the rth death, the experimenter may stop after a period of time T, which may be six months or a year. If we denote the number of deaths between 0 and Tasr, the survival data may look as follows: t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p45t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8/p58T Mathematical derivations of the MLE of /afii9838and /afii9839are exactly the same and (7.2.10 )can still be used. The samplingdistribution of /afii9839/p24for singly censored data is also discussed by Bartholomew (1963 ). Progressively censored data come more frequently from clinical studies where patients are entered at different times and the study lasts a predeter-mined period of time. Suppose that the study begins at time 0 and terminatesat timeTand there are a total of npeople entered. Let rbe the number of patients who die before or at time Tandn/p57rthe number of patients who are lost to follow-up duringthe study period or remain alive at time T. The data look as follows: t/p16,t/p17,...,t/p80,t/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76. Orderingthe runcensored observations accordingto their mag nitude, we have t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p80,t/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76 The likelihood function, using (7.1.1 ),(7.2.1 ), and (7.2.2 ),i s L/p58/p80/p147 /p71/p14/p16/afii9838e/p92/p72/p82 /p7/p71/p8/p76/p147 /p71/p14/p80/p62/p16e/p92/p72/p82/p71/p62 and the log-likelihood function is l(/afii9838)/p58n/afii9838/p57/afii9838/p80/p26 /p71/p14/p16t/p71/p57/afii9838/p76/p26 /p71/p14/p80/p62/p16t/p62/p71(7.2.15)170       and from (7.1.2 ), the MLE of the parameter /afii9838is /afii9838/p19/p58r /afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p71(7.2.16) Consequently, /afii9839/p24/p581 /afii9838/p19/p58/afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p71r(7.2.17) is the MLE of the mean survival time. The sum of all of the observations, censored and uncensored, divided by the number of uncensored observations,gives the MLE of the mean survival time. To overcome the mathematicaldifficulties arisingwhen all of the observations are censored (r/p580), Bar- tholomew (1957 )defines /afii9839/p24/p58/p76/p26 /p71/p14/p16t/p62/p71(7.2.18) In practice, this estimate has little value. Distributions of the estimators are discussed by Bartholomew (1957 ). The distributionof /afii9838/p19forlargenis approximatelynormalwith mean /afii9838and variance: Var(/afii9838/p19)/p58/afii9838/p17 /p76/p26 /p71/p14/p16(1/p57e/p92/p72/p50 /p71)(7.2.19) whereT/p71is the time that the ith person is under observation. In other words, T/p71is computed from the time the ith person enters the study to the end of the study. If the observation times T/p71are not known, the followingquick estimate of Var (/afii9838/p19)can be used: Var/p19(/afii9838/p19)/p58/afii9838/p19/p17 r(7.2.20 ) Thus an approximate 100 (1/p57/afii9838)% confidence interval for /afii9838is, by (7.1.6 ), /afii9838/p19/p57Z/p63/p30/p17/p40Var/p19(/afii9838/p19)/p58/afii9838/p58/afii9838 /p19/p59Z/p63/p30/p17/p40Var/p19(/afii9838/p19) (7.2.21 ) The distribution of /afii9839/p24is approximately normal with mean /afii9839and variance: Var(/afii9839/p24)/p58/afii9839/p17 /afii9814/p76/p71/p14/p16(1/p57e/p92/p72/p50/p71)(7.2.22)  171 Again, a quick estimate is Var/p19(/afii9839/p24)/p58/afii9839/p24/p17 r(7.2.23) An approximate 100 (1/p57/afii9825)% confidence interval for /afii9839is then, by (7.1.6 ), /afii9839/p24/p57Z/p63/p30/p17/p40Var/p19(/afii9839/p24)/p58/afii9839/p58/afii9839 /p24/p59Z/p63/p30/p17/p40Var/p19(/afii9839/p24) (7.2.24 ) The exact distribution of /afii9839/p24derived by Bartholomew (1963 )is too cumbersome for general use and thus is not included here. Example 7.4 Consider the remission duration of the 21 leukemia patients receiving6-MP in Example 3.3. The remission times in weeks were 6, 6, 6, 7,10, 13, 16, 22, 23, 6 /p59,9/p59,1 0/p59,1 1/p59,1 7/p59,1 9/p59,2 0/p59,2 5/p59,3 2/p59,3 2/p59,3 4/p59, and 35 /p59. The hazard plot given in Figure 3.6 shows that the exponential distribution fits the data very well. Maximum likelihood estimates of therelapse rate and the mean remission time can be obtained, respectively, from(7.2.16 )and (7.2.17 ): /afii9838/p19/p589 109/p59250 /p580.025 per week /afii9839/p24/p581 0.025/p5840 weeks The graphical estimate of /afii9838obtained in Example 3.3 is 0.027, which is very close to the MLE. Thus, the remission duration of leukemia patients receiving6-MP can be described by an exponential distribution with a constant weeklyrelapse rate of 2.5% and a mean remission time of 40 weeks. The probabilityof stayingin remission for one year (or 52 weeks )or more is estimated by S/p19(52)/p58exp[/p570.025 (52)]/p580.273 Using (7.2.20 )and (7.2.23 )for the variance of /afii9838/p19and /afii9839/p24, the 95% confidence intervals for /afii9838and /afii9839are, respectively, (0.009, 0.041 )and (13.867, 66.133 ). Example 7.5 The results in Examples 7.2 to 7.4 can also be obtained by usingavailable statistical software. Let tdenote the observed survival time (exact or censored )and CENS be an index (or dummy )variable with CENS /p580i ftis censored and 1 otherwise. Assume that the data have been saved in ‘‘C: /p33EXAMPLE.DAT’’ as a text file, which contains two columns (t in the first column and CENS in second column for the same study subject ), separated by a space. The followingSAS code for procedure LIFEREG can be used to obtain the estimated covariance matrix defined in (7.1.5 )and the MLE of the parameter of the exponential distribution for the observed survival data in‘‘C:/p33EXAMPLE.DAT’’.172       data w1; infile ‘c: /p33example.dat’ missover; input t cens; run; proc lifereg; model t*cens (0)/p58/covb d /p58exponential; run; The respective BMDP code for program 2L is /input file /p58‘c:/p33example.dat’ . variables /p582. format /p58free. /print level /p58brief. cova. survival. /variable names /p58t, cens. /form time /p58t. status /p58cens. response /p581. /regress accel /p58exponential. /end If SAS is used, the estimated parameter of the exponential distribution can be obtained by /afii9838/p19/p58exp(/p57INTERCEPT ), where INTERCEPT is the name of output estimated parameter in SAS procedure LIFEREG. In BMDP 2L, /afii9838/p19/p58exp(/p57CONSTANT ) where CONSTANT is given by the program. 7.2.2 Two-Parameter Exponential Distribution In the case where a two-parameter exponential distribution is more appropri- ate for the data (Zelen, 1966 ), the density and survivorship functions are defined, respectively, as f(t)/p58/p7/afii9838e/p92/p72/p7/p82/p92/p37/p8 0t/p46G/p460, /afii9838/p460 t/p58G(7.2.25 ) and S(t)/p58/p7e/p92/p72/p7/p82/p92/p37/p8 1t/p46G/p460, /afii9838/p460 t/p58G(7.2.26) whereGis called the guaranteetime , the minimum survival time before which no deaths occur.  173 Estimation of /afii9838and G for Data without Censored Observations Ift/p16,t/p17,...,t/p76are the survival times of the npatients, using (7.1.1 ),(7.1.2 ), (7.2.25 ), and (7.2.26 ), the MLE of /afii9838is /afii9838/p19/p58n /p76/p26 /p71/p14/p16(t/p71/p57G/p19)(7.2.27) whereG/p19is an estimate of Gthat is the smallest observation in the data, G/p19/p58min(t/p16,t/p17,...,t/p76)( 7 .2.28) and the mean survival time is estimated by /afii9839/p24/p58G/p19/p591//afii9838/p19. Example 7.6 Consider the survival times in months of 11 patients following initial pulmonarymetastasis from ostenogenicsarcoma consideredby Burdetteand Gehan (1970 ). The data were 11, 13, 13, 13, 13, 13, 14, 14, 15, 15, and 17. Suppose that the two-parameter exponential distribution is selected. Theguarantee time Gis estimated by the smallest observation (i.e.,G/p19/p5811), and the hazard rate /afii9838/p19estimated by (7.2.27 )is /afii9838/p19/p5811 (11/p5711)/p59(13/p5711)/p59/p37/p59(17/p5711) /p580.367 Thus, the exponential model tells us that the minimum survival time is 11 months, and after that the chance of death per month is 0.367. Similarly, theprobability of survivinga g iven amount of time can then be estimated from(7.2.26 ). For example, the estimated probability of surviving18 months or longer is S/p19(18)/p58exp[/p570.367 (18/p5711)]/p580.077 Estimation of /afii9838and G for Data with Censored Observations We first consider singly censored data. Suppose that an experiment begins withnanimals and terminates as soon as the first rdeaths occur. For this case, we introduce the estimation procedures derived by Epstein (1960a ). Let the first rsurvival times be t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8and letT*be the total survival observed between the first and the rth death: T*/p58(n/p571)(t/p7/p17/p8/p57t/p7/p16/p8)/p59(n/p572)(t/p7/p18/p8/p57t/p7/p17/p8)/p59/p37/p59(n/p57r/p591)(t/p7/p80/p8/p57t/p7/p80/p92/p16/p8) /p58/p57(n/p571)t/p7/p16/p8/p59t/p7/p17/p8/p59t/p7/p18/p8/p59/p37/p59t/p7/p80/p92/p16/p8/p59(n/p57r/p591)t/p7/p80/p8 /p58/p80/p26 /p71/p14/p16t/p7/p71/p8/p57nt/p7/p16/p8/p59(n/p57r)t/p7/p80/p8(7.2.29 )174       The best estimates for Gand /afii9839in the sense that they are unbiased and have minimum variance are given by G/p19/p58t/p7/p16/p8/p57/afii9839/p24 n(7.2.30) and /afii9839/p24/p58T* r/p571(7.2.31) Then /afii9838can then be estimated by /afii9838/p19/p581//afii9839/p24. Confidence intervals for the mean survival time /afii9839are easy to obtain from the fact that 2 (r/p571)/afii9839/p24//afii9839/p582T*//afii9839has a chi-square distribution with 2( r/p571) degreesoffreedom.Thus,for r/p571,the100 (1/p57/afii9825)%confidenceintervalfor /afii9839is 2(r/p571)/afii9839/p24 /afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p63/p30/p17/p58/afii9839/p582(r/p571)/afii9839/p24 /afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p16/p92 /p63/p30/p17(7.2.32) or 2T* /afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p63/p30/p17/p58/afii9839/p582T* /afii9851/p17/p17/p7/p80/p92/p16/p8/p11/p16/p92 /p63/p30/p17(7.2.33) To find confidence intervals for G, we use the fact that x/p16/p582n(t/p7/p16/p8/p57G)//afii9839 andx/p17/p582(r/p571)/afii9839/p24//afii9839are independent and have a chi-square distribution with 2 and 2 (r/p571) degrees of freedom, respectively. Thus the ratio Y/p58x/p16/2 x/p17/2(r/p571)/p58n(t/p7/p16/p8/p57G) /afii9839/p24/p58n(r/p571)(t/p7/p16/p8/p57G) T*(7.2.34) follows the F-distribution with 2 and 2( r/p571) degrees of freedom. Let F/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63be the 100 /afii9825percentage point of the F/p17/p11/p17/p7/p80/p92/p16/p8distribution [i.e., P(Y/p46F/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63)/p58/afii9825](Table B-3 in Appendix B ), and then a 100 (1/p57/afii9825)% confidence interval for Gis t/p7/p16/p8/p57/afii9839/p24 nF/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63/p58G/p58t/p7/p16/p8(7.2.35 ) or t/p7/p16/p8/p57T* n(r/p571)F/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63/p58G/p58t/p7/p16/p8(7.2.36) Epstein and Sobel (1953 )show that this interval is the shortest in the class of intervalsbeingused. If for someparticularvaluesof rand /afii9825thevalueF/p17/p11/p17/p7/p80/p92/p16/p8/p11/p63is not tabulated in the F-table, Epstein (1960a )suggests using the following  175 confidence intervals for G: t/p7/p16/p8/p57/afii9839/p24(r/p571) ng/p16/p92/p63/p58G/p58t/p7/p16/p8(7.2.37) or t/p7/p16/p8/p57T* ng/p16/p92/p63/p58G/p58t/p7/p16/p8(7.2.38) where g/p16/p92/p63/p58/p11 /afii9825/p2/p16/p30/p7/p80/p92/p16/p8/p571( 7 .2.39) is computable for any /afii9825andr. Example 7.7 illustrates the procedures. Example 7.7 In a laboratoryexperiment20 mice areinjectedwith a tumor inoculum. These tumor cells multiply and eventually kill the animal. Supposethat the investigator decides to terminate the experiment after 10 deaths. Thefirst occurs 30 days after the experiment starts. The total survival observedbetween the time when the first and tenth deaths occur is 600 animal days.Assumingthat the survival distribution of these mice is exponential, theshortest 95% confidence interval for Gcan be obtained by (7.2.36 ). Since F/p17/p11/p16/p23/p11/p15/p13/p15/p20/p583.555, the interval is 30/p57600 (20)(9) (3.555 )/p58G/p5830 or(18.150, 30 ). The mean survival time estimated by (7.2.31 )is/afii9839/p24/p5866.667 days, and the 95%confidence interval for /afii9839computed from (7.2.33 )is 2(600) 31.526/p58/afii9839/p582(600) 8.231 or(38.064, 145.790 ). When data are progressively censored, Gehan (1970 )derives an estimate for Gand a modified MLE for the hazard rate /afii9838. Suppose that rout of then individuals in the study die before the end of the study and n/p57rindividuals are alive at the time of the last follow-up or termination. The nsurvival times are denoted by t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8,t/p62/p7/p80/p62/p16/p8,...,t/p62/p7/p76/p8 An estimate of Gobtained by G/p19/p58max/p1t/p7/p16/p8/p571 n/afii9838/p19,0/p2(7.2.40)176       and the variance of G/p19is Var(G/p19)/p581 (n/afii9838/p19)/p17/p11/p591 r/p571/p2(7.2.41) Whennis large,Gand Var (G/p19)can be estimated by G/p19/p60t/p7/p16/p8(7.2.42) and Var/p19(G/p19)/p601 (n/afii9838/p19)/p17(7.2.43) A modified MLE for /afii9838is /afii9838/p19/p58r/p571 /afii9814/p80/p71/p14/p16t/p7/p71/p8/p59/afii9814/p76/p71/p14/p80/p62/p16t/p62/p7/p71/p8/p57nt/p7/p16/p8(7.2.44 ) with variance Var(/afii9838/p19)/p58/afii9838/p17 r/p571(7.2.45) Any percentile of survival time t/p78may be estimated by equating S(t)t opand solvingfort/p19/p78; that is,t/p19/p78/p58/p57 (log/p67p)//afii9838/p19/p59G/p19. The followingexample illustrates the procedures. Example 7.8 Suppose that 19 patients with brain tumor are followed in a clinical trial for a year. Their survival times in weeks are 3, 4, 6, 8, 8, 10, 12,16,17,30,33,3 /p59,8/p59,13/p59,21/p59,26/p59,3 5/p59,44/p59,and 45 /p59.Inthis case n/p5819, r/p5811,t/p7/p16/p8/p583,/afii9814/p16/p16/p71/p14/p16t/p7/p71/p8/p58147, and /afii9814/p16/p24/p71/p14/p16/p17t/p62/p7/p71/p8/p58195. The hazard rate /afii9838per week and its variance may be estimated by (7.2.44 )and (7.2.45 )as /afii9838/p19/p5810 147/p59195/p5719(3)/p580.035 and Var/p19(/afii9838/p19)/p58(0.035 )/p17 10/p580.0001 The guarantee time Gand its variance may then be estimated by (7.2.40 )and (7.2.41 ):  177 G/p19/p58max /p13/p571 19/p590.035,0/p2/p581.496 and Var/p19(G/p19)/p581 (19/p590.035)/p17/p11/p591 10/p2/p582.487 Thus, after a guarantee time of approximately 1.5 weeks, the chance of death per week is 0.035. The estimated median survival time is t/p19/p15/p13/p20/p58/p57log0.5 0.035/p591.496 /p5821.3 weeks The probability of survivingat least six months (or 26 weeks )is estimated by S/p19(26)/p58exp[/p570.35(26/p571.496 )]/p580.424 7.3 WEIBULL DISTRIBUTION The Weibull distribution has the density and survivorship functions f(t)/p58/afii9828/afii9838/p65t/p65/p92/p16exp[/p57(/afii9838t)/p65] S(t)/p58e/p92/p7/p72/p82/p8/p65t/p460, /afii9828/p570, /afii9838/p570 (7.3.1 ) The MLE of the parameters /afii9838and /afii9828involves equations to be solved simulta- neously. Numerical methods such as the Newton —Raphson iterative procedure (7.1.13 )can be applied.We begin with the case whereno censoredobservations are presented. Lett/p16,t/p17,...,t/p76be the exact survival times of nindividuals under investiga- tion. If their survival times follow the Weibull distribution, the log-likelihoodfunction is l(/afii9838,/afii9828)/p58nlog/afii9828/p59n/afii9828log/afii9838/p59/p76/p26 /p71/p14/p16[(/afii9828/p571)logt/p71/p57/afii9838/p65t/p65/p71]( 7.3.2) The MLE of /afii9838and /afii9828in(7.3.1 )can be obtained by solvingthe followingtwo equations simultaneously: n/p57/afii9838/p19 /afii9828/p24/p76/p26 /p71/p14/p16t/p71/afii9828/p24/p580 (7.3.3 ) n /afii9828/p24/p59nlog/afii9838/p19/p59/p76/p26 /p71/p14/p16logt/p71/p57/afii9838/p19/afii9828/p24/p76/p26 /p71/p14/p16t/p71/afii9828/p24(log/afii9838/p19/p59logt/p71)/p580( 7.3.4)178       Next, let us consider a typical laboratory experiment in which subjects are entered at the same time and the experiment is terminated after rof then subjects have failed (or after a fixed period of time T). In both of these cases the data collected are singly censored. The ordered survival data are t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p58t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8 If the time to failure follows the Weibull distribution with the density function given in (7.3.1 ), the MLE of /afii9838and /afii9828may be obtained by solvingthe following two equations simultaneously: r/p57/afii9838/p19/afii9828/p24/p3/p80/p26 /p71/p14/p16t/p71/afii9828/p24/p59(n/p57r)t/afii9828/p24 /p7/p80/p8/p4/p580 (7.3.5 ) r /afii9828/p24/p59rlog/afii9838/p19/p59/p80/p26 /p71/p14/p16logt/p71 /p59/afii9838/p19/afii9828/p24/p3/p80/p26 /p71/p14/p16t/p71/afii9828/p24(log/afii9838/p19/p59logt/p71)/p59(n/p57r)t/afii9828/p24 /p7/p80/p8(log/afii9838/p19/p59logt/p7/p80/p8)/p4/p580 (7.3.6 ) When data are progressively censored, we have t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8,t/p62/p7/p80/p62/p16/p8,...,t/p62/p7/p76/p8 If the survival distribution is Weibull defined by (7.3.1 ), the log-likelihood function is l(/afii9838,/afii9828)/p58rlog/afii9828/p59r/afii9828log/afii9838/p59/p80/p26 /p71/p14/p16[(/afii9828/p571)logt/p7/p71/p8/p57/afii9838/p65t/p65/p7/p71/p8]/p57/p76/p26 /p71/p14/p80/p62/p16/afii9838/p65t/p62/p65/p7/p71/p8 (7.3.7) The MLE of /afii9838and /afii9828may be obtained by solvingthe followingtwo equations simultaneously: r/p57/afii9838/p19/p65/p24/p1/p80/p26 /p71/p14/p16t/p71/afii9828/p24/p59/p76/p26 /p71/p14/p80/p62/p16t/p62/p71/afii9828/p24/p2/p580( 7.3.8) r /afii9828/p24/p59rlog/afii9838/p19/p59/p80/p26 /p71/p14/p16logt/p71/p57/afii9838/p19/afii9828/p24/p80/p26 /p71/p14/p16t/p71/afii9828/p24(log/afii9838/p19/p59logt/p71) /p57/afii9838/p19/afii9828/p24/p76/p26 /p71/p14/p80/p62/p16t/p71/p59/afii9828/p24(log/afii9838/p19/p59logt/p62/p71)/p580( 7.3.9) The followingexample illustrates the use of available computer software to obtain the MLE of /afii9838and /afii9828.  179 Example 7.9 Referringto Example 7.5, for the observed survival data in the file ‘‘EXAMPLE.DAT’’, we can use either SAS or BMDP to obtain theestimated parameters of the Weibull distribution. The codes given in Example7.5 can be used except that d /p58exponential in the SAS code must be changed to d /p58weibull and accel /p58exponential in BMDP code be changed to ac- cel/p58weibull. If SAS is used, the estimated parameters of the Weibull distribu- tion are /afii9838/p19/p58exp(/p57INTERCEPT )and /afii9828/p24/p581 SCALE where INTERCEPT and SCALE are produced by SAS procedure LIFEREG. If BMDP is used, /afii9838/p19/p58exp(/p57INTERCEPT )and /afii9828/p24/p581 SCALE where CONSTANT and SCALE are given by procedure 2L. 7.4 LOGNORMAL DISTRIBUTION If the survival time Tfollows the lognormal distribution with density function f(t)/p581 t/afii9846/p402/afii9843exp/p3/p571 2/afii9846/p17(logt/p57/afii9839)/p17/p4(7.4.1) the mean and the variance are exp (/afii9839/p59/p16/p17/afii9846/p17)and [exp (/afii9846/p17)/p571]exp (2/afii9839/p59/afii9846/p17), respectively. Estimation of the two parameters /afii9839and /afii9846/p17has been investigated either by using (7.4.1 )directly or by usingthe fact that Y/p58logTfollows the normal distribution with mean /afii9839and variance /afii9846/p17. In the following, we discuss theestimation of /afii9839and /afii9846/p17forsampleswith andwithout censoredobservations. 7.4.1 Estimation of /afii9839and/afii98462for Data without Censored Observations Estimationsof /afii9839and /afii9846/p17for completesamples by maximumlikelihood methods havebeen studied by many authors:forexample, Cohen (1951 )and Harterand Moore (1966 ). But the simplest way to obtain estimates of /afii9839and /afii9846/p17with optimum properties is by consideringthe distribution of Y/p58logT. Lett/p16, t/p17,...,t/p76be the survival times of nsubjects. The MLE of /afii9839is the sample mean ofYgiven by /afii9839/p24/p581 n/p76/p26 /p71/p14/p16logt/p71(7.4.2) The MLE of /afii9846/p17is /afii9846/p24/p17/p581 n/p3/p76/p26 /p71/p14/p16(logt/p71)/p17/p57(/afii9814/p76/p71/p14/p16logt/p71)/p17 n /p4(7.4.3)180       The estimate /afii9839/p24is also unbiased but /afii9846/p24/p17is not. The best unbiased estimates of /afii9839and /afii9846/p17are /afii9839/p24and the sample variance s/p17/p58 /afii9846/p24/p17[n/(n/p571)]. Ifnis moderately large, the difference between s/p17and /afii9846/p24/p17is negligible. One of the properties of the MLE is that if /afii9835/p21/p19is the MLE of /afii9835/p21,g(/afii9835/p21/p19)is the MLE ofg(/afii9835/p21)ifg(/afii9835/p21)is a finite function. Therefore, the MLEs of the mean and variance ofTare, respectively, exp (/afii9839/p24/p59/p16/p17/afii9846/p24/p17)and exp[ (/afii9846/p24/p17/p571)] exp (2/afii9839/p24/p59/afii9846/p24/p17). It is known that /afii9839/p24/p58y/p21is normally distributed with mean /afii9839and variance /afii9846/p17/n. Hence, if /afii9846is known, a 100 (1/p57/afii9825)% confidence interval for /afii9839is /afii9839/p24/p60Z/p63/p30/p17/afii9846//p40n.I f /afii9846is unknown, we can use Student’s t-distribution. A 100(1/p57/afii9825)%confidence interval for /afii9839is/afii9839/p24/p60t/p63/p30/p17/p11/p7/p76/p92/p16/p8s//p40n/p571, wheret/p63/p30/p17/p11/p7/p76/p92/p16/p8is the 100 /afii9825/2 percentage point of Student’s t-distribution with n/p571 degrees of freedom (Table B-7 ). Confidence intervals for /afii9846/p17can be obtained by usingthe fact that n/afii9846/p24//afii9846/p17 has a chi-square distribution with n/p571 degrees of freedom. A 100 (1/p57/afii9825)% confidence interval for /afii9846/p17is n/afii9846/p24/p17 /afii9851/p17/p7/p76/p92/p16/p8/p11/p63/p30/p17/p58/afii9846/p17/p58n/afii9846/p24/p17 /afii9851/p17/p7/p76/p92/p16/p8/p11/p16/p92 /p63/p30/p17(7.4.4) The followinghypothetical example illustrates the procedures. Example 7.10 Fivemelanoma (resected )patientsreceivingimmunotherapy BCG are followed. The remissionduration in weeks are, in order of magnitude,8, 16, 23, 27, and 28. Suppose that the remission times follow a lognormaldistribution. In this case, parameters are estimated by (7.4.2 )and (7.4.3 )as follows: t logt (logt)/p17 82 .079 4 .322 16 2 .773 7 .690 23 3 .135 9 .828 27 3 .296 10 .864 28 3 .332 11 .102——— ———14.615 43 .806 /afii9839/p24/p5814.615 5/p582.923 /afii9846/p24/p17/p581 5/p343.806 /p571 5(14.615 )/p17/p4/p580.217 s/p17/p585/afii9846/p24/p17 5/p571/p580.271  181 The mean remission time is exp (2.923 /p590.217/2 ), or 20.728, weeks and the standard deviation of the remission time is /p43[exp (0.217 )/p571] exp(5.846 /p590.217 )/p44/p16/p30/p17, or 10.204, weeks. A 95% confidence interval for /afii9839is 2.923 /p572.776 /p10.521 /p404/p2/p58/afii9839/p582.923 /p592.776 /p10.521 /p404/p2 or(2.200,3.646 ). A 95% confidence interval for /afii9846/p17, following (7.4.4 ),i s 5(0.217 ) 11.1433/p58/afii9846/p17/p585(0.217 ) 0.4844 or(0.097, 2.240 ). 7.4.2 Estimation of /afii9839and/afii98462for Data with Censored Observations We first consider samples with singly censored observations. The data consist ofrexact survival times t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8andn/p57rright-censored survival times that are at least t/p7/p80/p8, denoted by t/p62/p7/p80/p8. Again, we use the fact that Y/p58logT has normal distribution with mean /afii9839and variance /afii9846/p17. Estimates of /afii9839and /afii9846/p17 can be obtained from the transformed data y/p71/p58logt/p71. Many authors have investigatedthe estimation of /afii9839and /afii9846/p17: for example, Gupta (1952 ), Sarhan and Greenberg (1956, 1957, 1958, 1962 ), Saw (1959 ), and Cohen (1959, 1961 ).W e shall discuss the methods of Sarhan and Greenbergand Cohen because of theavailable table that reduces computation time and efforts. The best linear estimates of /afii9839and /afii9846proposed by Sarhan and Greenbergare linear combinations of the logarithms of the rexact survival times: /afii9839/p24/p58/p80/p26 /p71/p14/p16a/p71logt/p7/p71/p8(7.4.5) and /afii9846/p24/p58/p80/p26 /p71/p14/p16b/p71logt/p7/p71/p8(7.4.6) where the coefficients a/p71andb/p71are calculated and tabulated by Saharan and Greenbergfor n/p4520 and are partially reproduced in Table B-8. The variance and covariance of /afii9839/p24and /afii9846/p24are tabulated in Table B-9. The followingexample illustrates the procedure. Example 7.11 Supposethatinastudyoftheefficacyofanewdrug,12mice with tumors are given the drug. The experimenter decides to terminate thestudy after 9 mice have died. The survival times are, in weeks, 5, 8, 9, 10, 12,15, 20, 21, 25, 25 /p59,2 5/p59, and 25 /p59. Assume that the times to death of these182       mice follow the lognormal distribution. In this case n/p5812,r/p589, and n/p57r/p583. Using (7.4.5 ),(7.4.6 ), and Table B-8, /afii9839/p24and /afii9846/p24can be calculated as /afii9839/p24/p580.036log5 /p590.0581log8 /p590.0682log9 /p590.0759log10 /p590.0827log12 /p590.0888log15 /p590.0948log20 /p590.1006log21 /p590.3950log25 /p582.811 /afii9846/p24/p58/p570.2545log5 /p570.1487log8 /p570.1007log9 /p570.0633log10 /p570.0308log12 /p570.0007log15 /p590.0286log20 /p590.0582log21 /p590.5119log25 /p580.747 The variance of /afii9839/p24and /afii9846/p24given in Table B-9 are, respectively, 0.0926 and 0.0723 and the covariance of /afii9839/p24and /afii9846/p24is 0.0152. Cohen’s (1959, 1961 )MLEs for the normal distribution can be used for n/p5720. Let y/p21/p581 r/p80/p26 /p71/p14/p16logt/p7/p71/p8(7.4.7) and s/p17/p581 r/p3/p26(logt/p7/p71/p8)/p17/p57(/afii9814logt/p7/p71/p8)/p17 r /p4(7.4.8) Then the MLEs of /afii9839and /afii9846/p17are /afii9839/p24/p58y/p21/p57/afii9838/p19(y/p21/p57logt/p7/p80/p8)( 7 .4.9) and /afii9846/p24/p17/p58s/p17/p59 /afii9838/p19(y/p21/p57logt/p7/p80/p8)/p17 (7.4.10) where the value of /afii9838/p19has been tabulated by Cohen (1961 )as a function of aand b. The proportion of censored observations, b, is calculated as b/p58n/p57r n and a/p581/p57Y(Y/p57c) (Y/p57c)/p17 whereY/p58[b/(1/p57b)]f(c)/F(c),f(c) andF(c) beingthe density and distribu-  183 tion functions, respectively, of the standard normal distribution, evaluated at c/p58(logt/p7/p80/p8/p57/afii9839)//afii9846. Table 7.1 gives values of /afii9838/p19forb/p580.01 to 0.90 and a/p580.00 to 1.00. For a censored sample, after computing a/p58s/p17/(y/p21/p57logt/p7/p80/p8)/p17, and b/p58(n/p57r)/n, enter Table 7.1 with these values of aandbto obtain /afii9838/p19. For values not tabulated, two-way linear interpolation can be used. The asymptotic variances and covariance are the following: Var(/afii9839/p24)/p58/afii9846/p17 nm/p16 Var(/afii9846/p24)/p58/afii9846/p17 nm/p17(7.4.11 ) Cov(/afii9839/p24,/afii9846/p24)/p58/afii9846/p17 nm/p18 wherem/p16,m/p17, andm/p18are also tabulated by Cohen (1961 ). The table is reproduced in Table 7.2. For any censored sample, compute c/p24/p58(logt/p7/p80/p8/p57/afii9839/p24)//afii9846/p24 and then enter the appropriate columns of Table 7.2 with y/p58/p57c/p24, and interpolateto obtain the required values of m/p71,i/p581, 2, 3, if the experiment was terminated after a predetermined time. If the experiment was terminated aftera given proportion of animals have died, enter Table 7.2 through the percentcensored column with percentage censored /p58100band interpolate to obtain the required value of m/p71. To illustrate the use of Tables 7.1 and 7.2 for the computation of /afii9839/p24,/afii9846/p24/p17, Var(/afii9839/p24), Var (/afii9846/p24), and Cov (/afii9839/p24,/afii9846/p24), consider Example 7.12, adapted from Cohen (1961 ). Example 7.12 Suppose that in a laboratory experiment 300 insects were followed until 119 died within seconds, y/p21/p581,304.832 seconds, s/p17/p5812,128.250, and logt/p7/p16/p16/p24/p8/p581,450.000 seconds. In this case n/p58300 andr/p58119. Accord- ingly, a/p24/p5812,128.25 (1,304.832/p571,450) /p17 /p580.575b/p58300/p57119 300/p580.603 From Table 7.1, /afii9838/p19is approximately 1.36. Using (7.4.9 )and (7.4.10 ), we obtain /afii9839/p24/p581,304.832 /p571.36(1,304.832 /p571,450 )/p581,502.26 seconds /afii9846/p24/p17/p5812,128.250 /p591.36(1,304.832 /p571,450 )/p17/p5840,788.55 and /afii9846/p24/p17/p58201.96 seconds. For the variance and covariance of /afii9839/p24and /afii9846/p24, we enter Table 7.2 with percentage censored 100 b/p5860.3 and interpolate linearly to obtain m/p16/p582.002,184       Table7.1 Estimate d Value s for /afii9839/p24and/afii98462 b a.01 .02 .03 .04 .05 .06 .07 .08 .09 .10 .15 .20 .25 .30 .35 .40 .45 .50 .55 .60 .65 .70 .80 .90 .00 .010100 .020400 .030902 .041583 .052507 .063627 .074953 .086488 .09824 .11020 .17342 .24268 .31862 .4021 .4941 .5961 .7096 0.8368 0.9808 1.145 1 .336 1.561 2.176 3.282 .05 .010551 .021294 .032225 .043350 .054670 .066189 .077909 .089834 .10197 .11431 .17935 .25033 .32793 .4130 .5066 .6101 .7252 0.8540 0.9994 1.166 1 .358 1.585 2.203 3.314 .10 .010950 .022082 .033398 .044902 .056596 .068483 .080568 .092852 .10534 .11804 .18479 .25741 .33662 .4233 .5184 .6234 .7400 0.8703 1.017 1.185 1. 379 1.608 2.229 3.345 .15 .011310 .022798 .034466 .046318 .058356 .070586 .038009 .095629 .10845 .12148 .18985 .26405 .34480 .4330 .5296 .6361 .7542 0.8860 1.035 1.204 1. 400 1.630 2.255 3.376 .20 .011642 .023459 .035453 .047629 .059990 .072539 .085280 .098216 .11135 .12469 .19460 .27031 .35255 .4422 .5403 .6483 .7678 0.9012 1.051 1.222 1. 419 1.651 2.280 3.405 .25 .011952 .024076 .036377 .048858 .061522 .074372 .087413 .10065 .11408 .12772 .19910 .27626 .35993 .4510 .5506 .6600 .7810 0.9158 1.067 1.240 1.4 39 1.672 2.305 3.435 .30 .012243 .024658 .037249 .050018 .062969 .076106 .089433 .10295 .11667 .13059 .20338 .28193 .36700 .4595 .5604 6.713 .7937 0.9300 1.083 1.257 1.4 57 1.693 2.329 3.464 .35 .012520 .025211 .038077 .051120 .064345 .077756 .091355 .10515 .11914 .13333 .20747 .28737 .37379 .4676 .5699 .6821 .8060 0.9437 1.098 1.274 1.4 76 1.713 2.353 3.492 .40 .012784 .025738 .038866 .052173 .065660 .079332 .093193 .10725 .12150 .13595 .21139 .29260 .38033 .4755 .5791 .6927 .8179 0.9570 1.113 1.290 1.4 94 1.732 2.376 3.520 .45 .013036 .026243 .039624 .053182 .066921 .080845 .094958 .10926 .12377 .13847 .21517 .29765 .38665 .4831 .5880 .7029 .8295 0.9700 1.127 1.306 1.5 11 1.751 2.399 3.547 .50 .013279 .026728 .040352 .054153 .068135 .082301 .096657 .11121 .12595 .14090 .21882 .30253 .39276 .4904 .5967 .7129 .8408 0.9826 1.141 1.321 1.5 28 1.770 2.421 3.575 .55 .013513 .027196 .041054 .055089 .069306 .083708 .098298 .11308 .12806 .14325 .22235 .30725 .39870 .4976 .6051 .7225 .8517 0.9950 1.155 1.337 1.5 45 1.788 2.443 3.601 .60 .013739 .027649 .041733 .055995 .070439 .085068 .099887 .11490 .13011 .14552 .22578 .31184 .40447 .5045 .6133 .7320 .8625 1.007 1.169 1.351 1.56 1 1.806 2.465 3.628 .65 .013958 .028087 .042391 .056874 .071538 .086388 .10143 .11666 .13209 .14773 .22910 .31630 .41008 .5114 .6213 .7412 .8729 1.019 1.182 1.366 1.577 1.824 2.486 3.654 .70 .014171 .028513 .043030 .057726 .072605 .087670 .10292 .11837 .13402 .14987 .23234 .32065 .41555 .5180 .6291 .7502 .8832 1.030 1.195 1.380 1.593 1.841 2.507 3.679 .75 .014378 .028927 .043652 .058556 .073643 .088917 .10438 .12004 .13590 .15196 .23550 .32489 .42090 .5245 .6367 .7590 .8932 1.042 1.207 1.394 1.608 1.858 2.528 3.705 .80 .014579 .029330 .044258 .059364 .074655 .090133 .10580 .12167 .13773 .15400 .23858 .32903 .42612 .5308 .6441 .7676 .9031 1.053 1.220 1.408 1.624 1.875 2.548 3.730 .85 .014755 .029723 .044848 .060153 .075642 .901319 .10719 .12325 .13952 .15599 .24158 .33307 .43122 .5370 .6515 .7761 .9127 1.064 1.232 1.422 1.639 1.892 2.568 3.754 .90 .014967 .030107 .045425 .060923 .076606 .092477 .10854 .12480 .14126 .15793 .24452 .33703 .43622 .5430 .6586 .7844 .9222 1.074 1.244 1.435 1.653 1.908 2.588 3.779 .95 .015154 .030483 .045989 .061676 .077549 .0093611 .10987 .12632 .14297 .15983 .24740 .34091 .44112 .5490 .6656 .7925 .9314 1.085 1.255 1.448 1.66 8 1.924 2.607 3.803 1.00 .015338 .030850 .046540 .062413 .078471 .094720 .11116 .12780 .14465 .16170 .25022 .34471 .44592 .5548 .6724 .8005 .9406 1.095 1.267 1.461 1.68 2 1.940 2.626 3.827 Source:Cohen (1961 ). /p63For all values 0 /p45a/p451,/afii9838/p580. 185 Table 7.2 Estimated Values of m1,m2. and m3for Var( /afii9839/p24), Var( /afii9846/p24), and Cov( /afii9839/p24,/afii9846/p24) Percentage ym/p16m/p17m/p18Censored /p574.0 1.00000 0.500030 0.000006 0.00 /p573.5 1.00001 0.500208 0.000052 0.02 /p573.0 1.00010 0.501180 0.000335 0.13 /p572.5 1.00056 0.505280 0.001712 0.62 /p572.4 1.00078 0.506935 0.002312 0.82 /p572.3 1.00107 0.509030 0.003099 1.07 /p572.2 1.00147 0.511658 0.004121 1.39 /p572.1 1.00200 0.514926 0.005438 1.79 /p572.0 1.00270 0.518960 0.007123 2.28 /p571.9 1.00363 0.523899 0.009266 2.87 /p571.8 1.00485 0.529899 0.011971 3.59 /p571.7 1.00645 0.537141 0.015368 4.46 /p571.6 1.00852 0.545827 0.019610 5.48 /p571.5 1.01120 0.556186 0.024884 6.68 /p571.4 1.01467 0.568417 0.031410 8.08 /p571.3 1.01914 0.582981 0.039460 9.68 /p571.2 1.02488 0.600046 0.049355 11.51 /p571.1 1.03224 0.620049 0.061491 13.57 /p571.0 1.04168 0.643438 0.076345 15.87 /p570.9 1.05376 0.670724 0.094501 18.41 /p570.8 1.06923 0.702513 0.116674 21.19 /p570.7 1.08904 0.739515 0.143744 24.20 /p570.6 1.11442 0.782574 0.176698 27.43 /p570.5 1.14696 0.832691 0.217183 30.85 /p570.4 1.18876 0.891077 0.266577 34.46 /p570.3 1.24252 0.959181 0.327080 38.21 /p570.2 1.31180 1.03877 0.401326 42.07 /p570.1 1.40127 1.13198 0.492641 46.02 0.0 1.51709 1.24145 0.605233 50.00 0.1 1.66743 1.37042 0.744459 53.980.2 1.86310 1.52288 0.917165 57.930.3 2.11857 1.70381 1.13214 61.790.4 2.45318 1.91942 1.40071 65.540.5 2.89293 2.17751 1.73757 69.15 0.6 3.47293 2.48793 2.16185 72.57 0.7 4.24075 2.86318 2.69858 75.800.8 5.2612 3.3192 3.3807 78.810.9 6.6229 3.8765 4.2517 81.591.0 8.4477 4.5614 5.3696 84.131.1 10.903 5.4082 6.8116 86.43 1.2 14.224 6.4616 8.6818 88.49 1.3 18.735 7.7804 11.121 90.321.4 24.892 9.4423 14.319 91.92186       Table7.2 Continued Percentage ym/p16m/p17m/p18Censored 1.5 33.339 11.550 18.539 93.32 1.6 44.986 14.243 24.139 94.521.7 61.132 17.706 31.616 95.54 1.8 83.638 22.193 41.664 96.41 1.9 115.19 28.046 55.252 97.13 2.0 159.66 35.740 63.750 97.722.1 222.74 45.930 99.100 98.212.2 312.73 59.526 134.08 98.612.3 441.92 77.810 182.68 98.932.4 628.58 102.59 250.68 99.18 2.5 899.99 136.44 346.53 99.38 Source:Cohen (1961 ). m/p17/p581.635, andm/p18/p581.051.Substitutingthese values and /afii9846/p24/p17/p5840,788.55 into (7.4.11 ), we obtain Var(/afii9839/p24)/p6040,788.55 (2.022 ) 300/p58274.91 var(/afii9846/p24)/p6040,788.55 (1.635 ) 300/p58222.30 Cov(/afii9839/p24,/afii9846/p24)/p6040,788.55 (1.051 ) 300/p58142.90 When the data are progressively censored, let t/p16,t/p17,...,t/p80be the uncensored andt/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76be the censored observations, the likelihood function, using (7.4.1 )and (7.1.1 ),i s l(/afii9839,/afii9846/p17)/p58/p57rlog(2/afii9843/afii9846/p17) 2/p57/p80/p26 /p71/p14/p16/p1logt/p71/p59(logt/p71/p57/afii9839)/p17 2/afii9846/p17 /p2 /p59/p76/p26 /p71/p14/p80/p62/p16log/p7/p16/p27 /p82/p62/p711 x/p402/afii9843/afii9846/p17exp/p3/p571 2/afii9846/p17(logx/p57/afii9839)/p17/p4dx/p8 and the MLE of /afii9839and /afii9846/p17can be obtained by solvingthe followingtwo  187 equations: /p80/p26 /p71/p14/p16logt/p71/p57/afii9839 /afii9846/p17/p59/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p62/p71logx/p57/afii9839 x/afii9846/p17/p402/afii9843/afii9846/p17exp/p3/p571 2/afii9846/p17(logx/p57/afii9839)/p17/p4dx /p16/p27 /p82/p62/p711 x/p402/afii9843/afii9846exp/p3/p571 2/afii9846/p17(logx/p57/afii9839)/p17/p4dx/p580 /p57n 2/afii9846/p17/p59/p80/p26 /p71/p14/p16(logt/p71/p57/afii9839)/p17 2/afii9846/p19 /p59/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p62/p71(logx/p57/afii9839)/p17 x2/afii9846/p19/p402/afii9843/afii9846/p17exp/p3/p571 2/afii9846/p17(logx/p57/afii9839)/p17/p4dx /p16/p27 /p82/p62/p711 x/p402/afii9843/afii9846/p17exp/p3/p571 2/afii9846/p17(logx/p57/afii9839)/p17/p4dx/p580 Again,this can be done by applyingthe Newton —Raphson iterative procedure. The followingexample illustrate the use of SAS and BMDP to obtain estimates of the lognormal parameters. Example 7.13 Referringto Example 7.5, for the observed survival data in the file ‘‘EXAMPLE.DAT’’, by changing d /p58exponential in SAS code to d/p58lnormal and accel /p58exponential in BMDP code to accel /p58lnormal we can obtain the estimated parameters of the lognormal distribution. If SAS isused, the estimated parameters of the lognormal distribution are /afii9839/p24/p58INTERCEPT and /afii9846/p24/p17/p58SCALE /p17 where INTERCEPT and SCALE are the names of output estimated par- ameters in SAS procedure LIFEREG. If BMDP is used, /afii9839/p24/p58CONSTANT and /afii9846/p24/p17/p58SCALE /p17 where CONSTANT and SCALE are given by procedure 2L. 7.5 STANDARD AND GENERALIZED GAMMA DISTRIBUTIONS The density function of the standard gamma distribution is f(t)/p58/afii9838 /afii9772(/afii9828)(/afii9838t)/p65/p92/p16exp(/p57/afii9838t)t/p460, /afii9838,/afii9828/p570 (7.5.1 )188       where /afii9772(/afii9828)/p58/p7/p25/p27/p15x/p65/p92/p16e/p92/p86dx (/afii9828/p571)! if /afii9828is an integer(7.5.2 ) In this section we discuss the MLE of /afii9838and /afii9828for data with and without censored observations. 7.5.1 Estimation of /afii9838and/afii9828for Data without Censored Observations Suppose that the npatients under study are followed to death and their exact survival times t/p16,t/p17,...,t/p76are known. The MLE of /afii9838and /afii9828can be obtained by solvingsimultaneously the two equations n/afii9828/p24 /afii9838/p19/p57/p76/p26 /p71/p14/p16t/p71/p580( 7 .5.3) and nlog/afii9838/p19/p57n/afii9772/p30(/afii9828/p24) /afii9772(/afii9828/p24)/p59/p76/p26 /p71/p14/p16logt/p71/p580( 7 .5.4) where /afii9772/p30(/afii9828)is the derivative of /afii9772(/afii9828), /afii9772/p30(/afii9828)/p58/p16/p27 /p15x/p65/p92/p16log(x)e/p92/p86dx (7.5.5) From (7.5.3 ), we have /afii9838/p19/p58n/afii9828/p24 /afii9814/p76/p71/p14/p16t/p71(7.5.6) On eliminating /afii9838, we substitute (7.5.6 )into (7.5.4 )and obtain /afii9772/p30(/afii9828/p24) /afii9772(/afii9828/p24)/p57log/afii9828/p24/p57log(/afii9811/p76/p71/p14/p16t/p71)/p16/p30/p76 /afii9814/p76/p71/p14/p16t/p71/n/p580( 7 .5.7) to solve for /afii9828/p24. This can be done by usingthe Newton —Raphson iterative procedure. Tables for the solution of (7.5.7 )for/afii9828/p24as a function of Rare given by Greenwood and Durand (1960 ), whereRis the ratio of the geometric mean to the arithmetic mean of the nobservations: R/p58(/afii9811/p76/p71/p14/p16t/p71)/p16/p30/p76 /afii9814/p76/p71/p14/p16t/p71/n(7.5.8)     189 Wilket al. (1962a )show thatthe relationshipbetween /afii9828/p24and 1/ (1/p57R) is linear. A table of /afii9828/p24values of as a function of 1/ (1/p57R)given in their paper is reproduced in Table B-10. Thus if Rand 1/ (1/p57R) are computed from the sample, a MLE of /afii9828can be found from Table B-10. For values not tabulated, linear interpolation can be used. Having /afii9828/p24so determined, /afii9838/p19can be obtained from (7.5.6 ). In the method of moments (Fisher, 1922 ), the estimators are obtained simply by equatingthe population mean and variance to the sample mean and variance. The moment estimators of /afii9828and /afii9838are /afii9838*/p58/afii9814/p76/p71/p14/p16t/p71/afii9814/p76/p71/p14/p16(t/p71/p57t/p16)/p17(7.5.9) and /afii9828*/p58(/afii9814/p76/p71/p14/p16t/p71)/p17 n/afii9814/p76/p71/p14/p16(t/p71/p57t/p16)/p17(7.5.10) Both types of estimators give biased estimates. The moment estimators are easy to calculate but are inefficient in the sense that their variances are largerthan the variance of the MLE. To reduce the bias, Lilliefors (1971 )suggests correction factors for these two types of estimators. The corrected MLE of /afii9828 and /afii9838are, respectively, /afii9828/p24/p65/p58/afii9828/p24 1/p593/n(7.5.11) and /afii9838/p19/p65/p58/afii9828/p24/p65t/p16/p11/p571 n/afii9828/p24/p65/p2(7.5.12) The corrected moments estimators of /afii9828and /afii9838are /afii9828*/p65/p58/afii9828* 1/p592/n/p573 n(7.5.13) and /afii9838*/p65/p58/afii9828*/p65t/p11/p571 n/afii9828*/p65/p2(7.5.14) Lilliefors shows by the Monte Carlo method that the corrected MLE and the method-of-moment estimates are approximately unbiased. In addition, aslongas /afii9828/p462, the corrected moments estimators have no more bias than the corrected MLE and for n/p5810 have considerably less bias. For n/p5810, 20 and /afii9828/p462, the variance is close to that of the MLE.190       Example 7.14 Ten patients with melanoma achieve remission after surgery and therapy. They are followed to relapse. The durations of remission inmonths are recordedas follows:5, 8, 10,11, 15, 20,21, 23, 30,and 40. Assumingthat the distribution of remission duration is standard gamma, we firstcalculate the MLE of /afii9828and /afii9838accordingto Wilk et al. (1962a ). To compute R, we obtain /afii9814/p76/p71/p14/p16t/p71/p58183 and (/afii9811/p76/p71/p14/p16t/p71)/p16/p30/p16/p15 /p5815.43. Therefore R/p580.84 and 1/(1/p57R)/p586.25. From Table B-10, /afii9828/p24/p582.89830 for 1/ (1/p57R)/p586.0 and /afii9828/p24/p583.14984 for 1/ (1/p57R)/p586.5. By linear interpolation, for 1/ (1/p57R)/p586.25, /afii9828/p24/p583.02407. From (7.5.6 ),/afii9838/p19/p580.16525. The corrected MLEs obtained from (7.5.11 )and (7.5.12 )are/afii9828/p24/p65/p582.326 and /afii9838/p19/p65/p580.122. The moment estimates of /afii9828 and /afii9838following (7.5.9 )and (7.5.10 )are /afii9838*/p580.173 and /afii9828*/p583.171. With the correction factors, /afii9838*/p65/p580.1225 and /afii9828*/p65/p582.3425, which are very close to the corrected MLE. 7.5.2 Estimation of /afii9828and/afii9838for Data with Censored Observations When data are singly censored, the survival times can be ordered as t/p7/p16/p8/p45t/p7/p17/p8/p45/p37/p45t/p7/p80/p8/p45t/p62/p7/p80/p62/p16/p8/p58/p37/p58t/p62/p7/p76/p8 whererpersons in the study have exact survival times recorded and n/p57r others have their lives terminated after the rth death occurs. In this case, the maximum likelihood procedure becomes much more complicated. Let /afii9834/p58/afii9838t/p7/p80/p8,P/p58[/afii9811/p80/p71/p14/p16t/p7/p71/p8]/p16/p30/p80/t/p7/p80/p8, andS/p58/afii9814/p80/p71/p14/p16t/p7/p71/p8/rt/p7/p80/p8. The MLE of /afii9834and /afii9828and can be obtained by solvingsimultaneously logP/p58n/afii9772/p30(/afii9828) r/afii9772(/afii9828) /p57n rlog/afii9834/p57/p1n r/p571/p2J/p30(/afii9828,/afii9834) J(/afii9828,/afii9834)(7.5.15) and S/p58/afii9828 /afii9834/p571 /afii9834/p1n r/p571/p2e/p92/p69 J(/afii9828,/afii9834)(7.5.16) where J(/afii9828,/afii9834)/p58/p16/p27 /p16t/p65/p92/p16e/p92/p69/p82dt (7.5.17) and J/p30(/afii9828,/afii9834)/p58/p42 /p42/afii9828J(/afii9828,/afii9834)/p58/p16/p27 /p16t/p65/p92/p16logte/p92/p69/p82dt (7.5.18) Wilk et al. (1962a )generate, for a grid of values of PandSandn/r, tables of values of /afii9828/p24and /afii9839/p24/p58/afii9828/p24//afii9834/p24based on the solutions of (7.5.15 )and (7.5.16 ). The tables are reproduced in Table B-11. Thus, to find /afii9828/p24and /afii9838/p19, one needs to     191 computePandSfirst. For specific values of P,S, andn/r,/afii9828/p24and /afii9839/p24may be looked up from Table B-11. Then /afii9838/p19can be obtained from /afii9838/p19/p58/afii9828/p24/[/afii9839/p24t/p7/p80/p8]. Interpolations may be needed when any of the values of P,S, andn/rare not tabulated. Example 7.15, adapted from Wilk et al. (1962a ), illustrates the procedure of calculating /afii9828/p24,/afii9839/p24, and /afii9838/p19when Table B-11 is used. This method can also be used in the case of a complete sample (no censored observations ); that is,r/p58n.I f it is obvious that some of the observations may be outliers (too large or too small ), it is reasonable not to use them in estimation. In this case, ris the number of observations used in the estimation procedure. Example 7.15 Consider an experiment with n/p5834 animals. The following data are the lifetimes t/p71in weeks of 34 animals: 3, 4, 5, 6, 6, 7, 8, 8, 9, 9, 9, 10, 10, 11, 11, 11, 13, 13, 13, 13, 13, 17, 17, 19, 19, 25, 29, 33, 42, 42, 52, 52 /p59,5 2/p59, and 52 /p59. The study is terminated when 31 animals have died and the other 3 are sacrificed. In our notation, n/p5834 andr/p5831. 1. Compute n/r,P, andS: n r/p5834 31/p581.10 S/p58/afii9814t/p71rt/p7/p80/p8/p58487 (31)(52)/p580.30 To compute P, it is easier first to compute log P: logP/p581 r/p26logt/p7/p71/p8/p57logt/p7/p80/p8/p581 31/p5933.90207 /p571.716/p58/p570.622385 HenceP/p580.24. 2. Consider the entries for n/r/p581.10 andP/p580.24 in Table B-11: S/p580.28: /afii9828/p24/p581.986 /afii9839/p24/p580.365 S/p580.32: /afii9828/p24/p581.449 /afii9839/p24/p580.410 Usinglinear interpolation, approximate estimates of /afii9828and /afii9839are/afii9828/p24/p581.72 and /afii9839/p24/p580.39. 3. Finally, /afii9838/p19/p581.72/ (0.39/p5952)/p580.085. For a more accurate two-way interpolation, the reader is referred to Wilk et al. (1962a ).192       When the data are progressively censored, let t/p16,t/p17,...,t/p80be the uncensored andt/p62/p80/p62/p16,t/p62/p80/p62/p17,...,t/p62/p76be the censored observations; the likelihood function is l(/afii9838,/afii9828)/p58logL(/afii9838,/afii9828)/p58n/afii9828log/afii9838/p57nlog/afii9772(/afii9828) /p59/p80/p26 /p71/p14/p16[(/afii9828/p571)logt/p71/p57/afii9838t/p71]/p59/p76/p26 /p71/p14/p80/p62/p16log/p1/p16/p27 /p82/p62/p71x/p65/p92/p16e/p92/p72/p86dx/p2 and the MLE of /afii9838and /afii9828can be obtained by solvingthe two equations n/afii9828 /afii9838/p57/p80/p26 /p71/p14/p16t/p71/p57/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p71/p62x/p65e/p92/p72/p86dx /p16/p27 /p82/p71/p62x/p65/p92/p16e/p92/p72/p86dx/p580 nlog/afii9838/p57n/afii9772/p30(/afii9828) /afii9772(/afii9828)/p59/p80/p26 /p71/p14/p16logt/p71/p59/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p71/p62x/p65/p92/p16e/p92/p72/p86log(x)dx /p16/p27 /p82/p71/p62x/p65/p92/p16e/p92/p72/p86dx/p580 usingthe Newton —Raphson iterative procedure. 7.5.3 Estimation of /afii9838,/afii9828, and /afii9825in the Extended Generalized Gamma Distribution for Data with or without Censored Observations The extended generalized gamma distribution has density function defined in (6.4.10 ), f(t)/p58/p34/afii9825/p34/afii9828/p65/afii9838/p63/p65t/p63/p65/p92/p16exp[/p57/afii9828(/afii9838t)/p63] /afii9772(/afii9828)t/p570, /afii9828/p570, /afii9838/p570( 7.5.19) Let us consider the case /afii9825/p570. Lett/p16,t/p17,...,t/p80be the uncensored and t/p62/p80/p62/p16,...,t/p62/p76the censored observations from npersons and the survival times follow the generalized gamma distribution. Then the likelihood function is l(/afii9838,/afii9828,/afii9825)/p58n/afii9825/afii9828log/afii9838/p59nlog/afii9825/p59n/afii9828log/afii9828/p57nlog/afii9772(/afii9828) /p59/p80/p26 /p71/p14/p16[(/afii9825/afii9828/p571)logt/p71/p57/afii9828(/afii9838t/p71)/p63] /p59/p76/p26 /p71/p14/p80/p62/p16log/p7/p16/p27 /p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x)/p63]dx/p8     193 and the MLE of /afii9838,/afii9828, and /afii9825can be obtained by solvingthe three equations n/afii9825/afii9828 /afii9838/p57/afii9828/afii9825/afii9838/p63/p92/p16/p80/p26 /p71/p14/p16t/p63/p71/p57/afii9828/afii9825/afii9838/p63/p92/p16/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p71/p62x/p63/p65/p92/p16/p62/p63exp(/p57/afii9828(/afii9838x)/p63)dx /p16/p27 /p82/p71/p62x/p63/p65/p92/p16exp(/p57/afii9828(/afii9838x)/p63)dx/p580( 7.5.20) n/afii9825log/afii9838/p59nlog/afii9828/p59n/p57n/afii9772/p30(/afii9828) /afii9772(/afii9828)/p59/p80/p26 /p71/p14/p16[/afii9825logt/p71/p57(/afii9838t/p71)/p63] /p59/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x)/p63][/afii9825log(x)/p57(/afii9838x)/p63]dx /p16/p27 /p82/p71/p62x/p63/p65/p92/p16exp(/p57/afii9828(/afii9838x)/p63)dx/p580 (7.5.21 ) n/afii9828log/afii9838/p59n /afii9825/p59/p80/p26 /p71/p14/p16[/afii9828logt/p71/p57/afii9828(/afii9838t/p71)/p63log(/afii9838t/p71)] /p59/p76/p26 /p71/p14/p80/p62/p16/p16/p27 /p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x)/p63][/afii9828log/p67(x)/p57/afii9828(/afii9838x)/p63log(/afii9838x)]dx /p16/p27 /p82/p71/p62x/p63/p65/p92/p16exp[/p57/afii9828(/afii9838x]/p63)dx/p580( 7.5.22) usingthe Newton —Raphson iterative procedure. If all the observed survival times are uncensored, the respective equations for the MLE of /afii9838,/afii9828, and /afii9825can be obtained simply by replacing rwithnin (7.5.20 )—(7.5.22 ). The SAS procedure LIFEREG can be used to obtain the MLE of /afii9838,/afii9828, and /afii9825in the extended generalized gamma distribution. Example 7.16 Referringto Example 7.5, for the observed survival data in the file ‘‘EXAMPLE.DAT‘, by changing d /p58exponential in the SAS code to d/p58gamma, one can obtain the MLE of the parameters of the extended generalized gamma distribution: /afii9838/p19/p58exp(/p57INTERCEPT ) /afii9825/p24/p58SHAPE1 SCALE/afii9828/p24/p581 SHAPE1 /p17 where INTERCEPT, SHAPE1, and SCALE are given by the SAS LIFEREG procedure.194       7.6 LOG-LOGISTIC DISTRIBUTION The log-logistic distribution has the density function f(t)/p58/afii9825/afii9828t/p65/p92/p16 (1/p59/afii9825t/p65)/p17(7.6.1) and survivorship function S(t)/p581 1/p59/afii9825t/p65(7.6.2) wheret/p460,/afii9825/p570,/afii9828/p570.Lett/p16,t/p17,...,t/p80be the uncensored and t/p62/p80/p62/p16, t/p62/p80/p62/p17,...,t/p62/p76the censored observations from npersons and the survival times follow the log-logistic distribution. Then the MLE of /afii9825and /afii9828can be obtained from solvingthe followingtwo simultaneous equations: r/p57/afii9825/p12/p80/p26 /p71/p14/p16t/p65/p711/p59/afii9825t/p65/p71/p59/p76/p26 /p71/p14/p80/p62/p16t/p62/p65/p711/p59/afii9825t/p62/p65/p71/p2/p580 (7.6.3 ) r /afii9828/p59/p80/p26 /p71/p14/p16log(t/p71)/p57/afii9825/p32/p80/p26 /p71/p14/p16t/p65/p71log(t/p71) 1/p59/afii9825t/p65/p71/p59/p76/p26 /p71/p14/p80/p62/p16t/p62/p65/p71log(t/p62/p71) 1/p59/afii9825t/p62/p65/p71/p4/p580 (7.6.4 ) usingthe Newton —Raphson iterative procedure. If all the survival times observed are uncensored, the respective equations for the MLE of /afii9825and /afii9828can be obtained simply by replacing rwithnin(7.6.3 )and (7.6.4 ). Example 7.17 Referringto Example 7.5, for the observed survival data in file ‘‘EXAMPLE.DAT‘, replacingd /p58exponential in the SAS code by d/p58llogistic and accel /p58exponential in the BMDP code by accel /p58llogistic, we can obtain the estimated parameters of the log-logistic distribution. If SASis used, the estimated parameters of the log-logistic distribution are /afii9825/p24/p58exp/p1/p57INTERCEPT SCALE /p2and /afii9828/p24/p581 SCALE where INTERCEPT and SCALE are given by the SAS LIFEREG procedure. If BMDP is used, /afii9825/p24/p58exp/p1/p57CONSTANT SCALE /p2and /afii9828/p24/p581 SCALE where CONSTANT and SCALE are produced by the BMDP procedure 2L.-  195 Example 7.18 Assume that the tumor-free time of the 30 rats in the low-fat diet group in Table 3.4 follows the log-logistic distribution. The estimates ofthe two parameters from either SAS or BMDP are /afii9825/p24/p580.000025484 and /afii9828/p24/p582.01866. Therefore, from Section 6.5, the median survival time for this group is 188.64 days, and the hazard function approach the peak at 190.37days. 7.7 OTHER PARAMETRIC SURVIVAL DISTRIBUTIONS The Gompertz distribution (Section 6.6 )has the followingsurvivorship and probability density functions: S(t)/p58exp /p3/p57e/p72 /afii9828(e/p65/p82/p571)/p4(7.7.1) f(t)/p58exp/p3(/afii9838/p59/afii9828t)/p571 /afii9828(e/p72/p62/p65/p82/p57e/p72)/p4(7.7.2) 7.7.1 Estimation of /afii9838and/afii9828for Data with or without Censored Observations Assumethat t/p16,t/p17,...,t/p76are the observedsurvivaltimes from nindividualsand the survival times follow the Gompertz distribution, without loss of generality,andassumethat t/p16,t/p17,...,t/p80areuncensoredand t/p62/p80,t/p62/p80/p62/p17,...,t/p62/p76right-censored. The MLE of /afii9838and /afii9828can be obtained by solvingthe equations r/p59e/p72 /afii9828/p7/p80/p26 /p71/p14/p16[1/p57exp(/afii9828t/p71)]/p59/p76/p26 /p71/p14/p80/p62/p16[1/p59exp(/afii9828t/p62/p71)]/p8/p580( 7.7.3) /p80/p26 /p71/p14/p16t/p71/p57e/p72 /afii9828/p17/p7/p80/p26 /p71/p14/p16[1/p59(/afii9828t/p71/p571)exp (/afii9828t/p71)]/p59/p76/p26 /p71/p14/p80/p62/p16[1/p59(/afii9828t/p62/p71/p571)exp (/afii9828t/p62/p71)]/p8/p580 (7.7.4 ) usingthe Newton —Raphson iterative procedure. If allt/p16,t/p17,...,t/p76are uncensored, the MLE of /afii9838and /afii9828can be obtained similarly by replacing rwithnin(7.7.3 )and (7.7.4 ). The MLE of the parameters of the other models in Section 6.6 can be obtained in a similarmanner. Bibliographical Remarks In addition to the papers cited in this chapter, Gross and Clark (1975 )have chapterson estimation and inferencein the exponentialdistribution and on the estimation of parameters of three distributions, includingthe Weibull and196       gamma. Mann et al. (1974 ), Lawless (1982 ), and Nelson (1982 )also provide a chapter on the estimation procedures for survival distributions, includingtheexponential, Weibull, gamma, and lognormal. A more recent book is by Kleinand Moeschberger (1997 ). Readers with a background in mathematical statis- tics and an interest in mathematical treatment of estimation procedures arereferred to these books. EXERCISES 7.1Consider the survival times given in Exercise 8.2. Assuming that they follow the one-parameter exponential distribution, obtain:(a)The MLE of /afii9838 (b)The MLE of /afii9839 (c)The 95% confidence intervals for /afii9838and /afii9839 7.2Assumingthatthe correctentriesbetweenerrors inExercise 8.3follow the two-parameter exponential distribution, obtain:(a)An estimate of G (b)The MLE of /afii9838 (c)The MLE of /afii9839 (d)The probability of 100 correct entries between two errors 7.3Consider the survival data in Exercise 8.5. Obtain the MLE of the parameter (s)and mean survival times, assuming: (a)A one-parameter exponential distribution (b)A Weibull distribution 7.4In a study of deep venous thrombosis, the followingblood clot lysis times in hours were recorded from 20 patients: 2, 3, 4, 5.5, 9, 13, 16.5, 17.5, 12.5,7, 6, 17.5, 11.5, 6, 14, 25, 49, 37.5, 49, and 28. Assume that the blood clotlysis times follow the lognormal distribution.(a)Obtain MLEs of the parameters /afii9839and /afii9846/p17. (b)Obtain 95% confidence intervals for /afii9839and /afii9846/p17. 7.5Consider the followingtumor-free times in days of 10 animals: 2, 3.5, 5, 7, 9, 10, 15, 20, 30, and 40. Assume that the tumor-free times follow thelog-logistic distribution. Estimate the parameters /afii9825and /afii9828. 197 CHAPTER 8 Graphical Methods for Survival Distribution Fitting The use of probabilitymodels for survival experience has play ed an increasing-lyimportant role in biomedical sciences. Survival models summarize thesurvival pattern, suggest further studies, and generate hypotheses. In thischapter we introduce three graphical methods for survival distribution fitting. In Section 8.1 we discuss the advantages of the graphical techniques. In Section 8.2 we discuss probabilityplotting, including how to make probabilityplots and how to estimate parameters from them. In Section 8.3 we discuss thetheoryand applications of hazard plotting for censored data. In Section 8.4 weintroduce the Cox —Snell residual method. 8.1 INTRODUCTION Graphical methods have long been used for displayand interpretation of data because theyare simple and effective. Often used in place of or in conjunctionwith numerical analysis, a plot of data serves a number of purposes simulta-neouslythat no numerical method can. The basic idea of the three graphicalmethods is to see if the survival time itself, or a function of it, has a linearrelationship with the distribution function and the cumulative hazard functionof a given parametric distribution, or a function of the distribution functionand the cumulative hazard function. If such a linear relationship exists, it canbe demonstrated graphicallyas a straight line. Thus, if one chooses theappropriate distribution and makes a probability, or hazard, plot, the resultwill be a straight line fit to the data. Parameters of the distribution chosen canbe estimated from the probabilityor hazard plots without tedious numericalcalculations. Such estimates maybe adequate and useful for preliminarypurposes. However, prior information is often not sufficient to choose asuitable distribution, and the plot maynot be a straight line. If the plot is nota straight line, there is no need to estimate the parameters and an alternative 198 Figure 8.1 Two curved normal probabilityplots. Figure 8.2 Two skewed densityfunctions.distributionmaybe selected.If the Cox —Snell residual plotting method is used, estimates of the parameters must be obtained first. A nonlinear plot can provide insight into the data. There are several pos- sible interpretations. First, the wrong theoretical distribution might havebeen used. Second, the sample might be from a mixture of populations. In thelatter case, it is necessaryto separate the data accordinglyand make a separateplot for each population. If one or two points are wayout of line, theymightbe the results of errors in collecting and recording the data or theymight notbe from the same population. Other reasons for peculiar looking plots andinterpretations of them are discussed byKing (1971 )and Hahn and Shapiro (1967 ). Consider the normal probabilityplots in Figures 8.1 aand b. The plot in Figure 8.1 ais convex, indicating that the data have a long tail to the left and could be from a distribution with a negativelyskewed densityfunction such asin Figure 8.2 a. On the contrary, the concave plot in Figure 8.1 bindicates that 199 the data have a long tail to the right and could be from a distribution with a positivelyskewed densityfunction such as in Figure 8.2 b. From the discussion in Chapter 6, we maytryto fit a lognormal or gamma distribution. The advantages of graphical methods can be summarized as follows: 1. Theyare fast and simple to use, in contrast with numerical methods, which maybe computationallytedious and require considerable analy ti-cal sophistication. The additional accuracyof numerical methods is usuallynot great enough in practice to warrant the effort involved. 2. Probabilityand hazard plots provide approximate estimates of the parameters of the distribution bysimple graphical means. 3. Theyallow one to assess whether a particular theoretical distribution provides an adequate fit to the data. 4. Peculiar appearance of a plot or points in a plot can provide insight into the data when the reasons for the peculiarities are determined. 5. A graph provides a visual representation of the data that is easyto grasp. This is useful not onlyfor oneself but also in presenting data to others,since a plot allows one to assess conclusions drawn from the data bygraphical or numerical means. 8.2 PROBABILITY PLOTTING The basic ideas in probabilityplotting are illustrated bythe following example. Example 8.1 Consider the white blood cell counts (WBCs )of 23 pediatric leukemia patients given in Table 8.1, ranging from 8000 to 120,000. A sample cumulative distribution is constructed byordering the data from smallest to largest,as shownin Table 8.1. A sample cumulativedistributioncurve can thenbemade byplottingeachWBCvalue versusthepercentage ofthe sampleequalto or less than that value. That is, the ith ordered data value in a sample of n values is plotted against the percentage 100 i/n. Note that for tied observations, we compute and plot the sample distribution onlyfor the one with the largestivalue. This gives a conservative estimate of the survivorship function. For example, the third value of WBC, 10, is plotted against a percentage of100/p593/23/p5813%. A plot of the cumulative distribution function for most large populations contains manycloselyspaced values and can be well approximated byasmooth curve drawn though the points. In contrast, a sample cumulativedistribution function has a relativelysmall number of points and thus some-what ragged appearance. To approximate the population cumulative distribu-tion function, one draws a smooth curve through the data points, obtaining abest fit byey e. Such a curve from the WBC data is given in Figure 8.3. It is an200       Table 8.1 Ordered WBCsData and Sample Cumulative Distribution for Example 8.1 Sample Distribution WBC Order, (10/p18) ii / 23 ( i/p570.5)/23 /afii9818/p92/p16(F)/p63 81 82 0 .087 0 .065 /p571.512 10 3 0 .130 0 .109 /p571.233 15 4 0 .174 0 .152 /p571.027 20 5 0 .217 0 .196 /p570.857 30 6 0 .261 0 .239 /p570.709 50 7 50 8 50 950 1050 11 0 .478 0 .457 /p570.109 60 1260 13 0 .565 0 .543 0 .109 75 14 75 15 0 .652 0 .630 0 .333 80 1680 17 0 .739 0 .717 0 .575 90 1890 19 90 20 0 .870 0 .848 1 .027 100 21 0 .913 0 .891 1 .233 110 22 0 .957 0 .935 1 .512 120 23 1 .000 0 .978 2 .019 /p63/afii9818/p92/p16(·) denotes the inverse of the standard normal distribu- tion function. estimate of the cumulative distribution function of the population and is used to obtain estimates and other information about the population. An estimate of the population median (50th percentile )is obtained by enteringtheplot onthepercentagescale at50%goinghorizontallyto thefittedline and then verticallydown to the data scale to read the estimate of themedian. For the WBC data, an estimate of the population median is 65,000.The median is a representative of nominal value for the population since halfof the population values are above it and half below. An estimate of anyotherpercentile can be obtained similarlybyentering the plot at the appropriatepoint on the percentage scale going horizontallyto the fitted line and thenverticallydown to the data scale where the estimate is read. For example, anestimate for the 25th percentile is 40,000.  201 Figure 8.3 Sample cumulative distribution curve of the WBC data. One can obtain an estimate of the proportion of the population that has a WBC below a specific value in a similar way. For example, to find theproportion of the population with a WBC of 10,000 or less, you enter the ploton the horizontal axis at the given value, 10, go verticallyup to the line fittedto the data, and then horizontallyto the probabilityscale, where the estimateof the population proportion is read, 8%. An estimate of the proportion of apopulation between two given values is obtained byfirst getting an estimate ofthe proportion below each value and then taking the difference. For example,the estimate of the population proportion with WBC between 10,000 and65,000 is 50 /p578/p5842%. As mentioned above, a smooth curve can be fitted byey e to a sample cumulative distribution function to obtain an estimate of the populationdistribution function. Also, one can fit data with a theoretical cumulativedistribution function byusing a probabilityplot and then use this plot to estimatethe parametersin thetheoreticalcumulativedistributionfunction.Thedistribution maybe the normal, lognormal, exponential, Weibull, gamma, orlog-logistic. To make a probabilityplot, one generallyuses (i/p570.5)/nor i/(n/p591) to estimate the sample cumulative distribution function at the ith ordered value of the nobservations in the sample. The (i/p570.5)/nfor the WBC data are given in Table 8.1. The probabilityplot is so constructed that if the theoretical distribution is adequate for the data, the graph of a function of t(used as the y-axis )versus a function of the sample cumulative distribution function (used as the x-axis ) will be close to a straight line. The parameters of the theoretical distributioncan then be estimated from a fitted line. This is carried out as follows. Step 1. A theoretical distribution for the survival time thas to be selected. Step 2.The sample cumulative distribution function is estimated byusing (i/p570.5)/nori/(n/p591),i/p581, 2, ...,n, for the ith ordered tvalue. For tied202       Figure 8.4 Normal probabilityplot of the WBC data in Example 8.1.observations have the same value, the sample cumulative distribution function is plotted against onlythe twith the largest ivalue. Step 3.Plot tor a function of it versus the estimated sample cumulative distribution or a function of it. Step 4.Fit a straight line through the points byey e. The position of the straight line should be chosen to provide a fit to the bulk of the data and mayignore outliers or data points of doubtful validity. Figure8.4 givesa normal probabilityplot of the WBCversus /afii9818/p92/p16(F), where /afii9818/p92/p16(·)is the inverse of the standard normal distribution function. The values of/afii9818/p92/p16(F/p19(WBC/p7/p71/p8))are shown in Table 8.1. The plot is reasonablylinear. The straight line fitted byey e in a probabilityplot can be used to estimatepercentiles and proportions within given limits in the same manner as for thesample cumulative distribution curve. In addition, a probabilityplot providesestimates of the parameters of the theoretical distribution chosen. The mean(or median )WBC estimated from the normal probabilityplot in Figure 8.4 is 56,000 [at /afii9818/p92/p16(F)/p580,F/p580.5 and WBC /p5856,000]. At /afii9818/p92/p16(F)/p581, WBC /p5891,000, which corresponds to the mean plus 1 standard deviation. Thus, the standard deviation is estimated as 35,000. We now discuss probabilityplots of the exponential, Weibull, lognormal, and log-logistic distributions.  203 Table 8.2 Probability Plotting for Example 8.2 Order, F, ti (i/p570.5)/21 log[1/ (1/p57F)] 11 1 2 0.071 0.074 23 2 4 0.167 0.1823 5 0.214 0.241464 7 0.310 0.37058 5 9 0.405 0.519 6 10 0.452 0.60281 18 12 0.548 0.7939 13 0.595 0.904 10 14 10 15 0.690 1.173 12 16 0.738 1.34014 17 0.786 1.54016 18 0.833 1.79220 19 0.881 2.12824 20 0.929 2.639 34 21 0.976 3.738Exponential Distribution The exponential cumulative distribution function is F(t)/p581/p57exp[/p57(/afii9838t)] t/p570( 8 .2.1) The probabilityplot for the exponential distribution is based on the relation- ship between tand F(t), from (8.2.1 ), t/p581 /afii9838log1 1/p57F(t)(8.2.2) This relationship is linear between tand the function log[1/ (1/p57F(t))]. Thus, an exponential probabilityplot is made byplotting the ith ordered observed survival time t/p7/p71/p8versus log[1/ (1/p57F/p19(t/p7/p71/p8))], where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8), for example, (i/p570.5)/n, for i/p581,...,n. From (8.2.2 ), at log /p431/[1/p57F(t)]/p44/p581,t/p581//afii9838. This fact can be used to estimate 1/ /afii9838and thus /afii9838from the fitted straight line. That is, the value t204       Figure 8.5 Exponential probabilityplot of the data in Example 8.2.corresponding to log /p431/[1/p57F(t)]/p44/p581 is an estimate of the mean 1/ /afii9838and its reciprocal is an estimate of the hazard rate /afii9838. Example 8.2 Suppose that 21 patients with acute leukemia have the following remission times in months: 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 8, 8, 9, 10, 10, 12,14, 16, 20, 24, and 34. We would like to know if the remission time follows theexponential distribution. The ordered remission times t/p7/p71/p8and the log /p431/ [1/p57F(t)]/p44are given in Table 8.2. The exponential probabilityplot is shown in Figure 8.5. A straight line is fitted to the points byey e,and the plot indicatesthat the exponential distribution fits the data verywell. At the point log[1/(1/p57F(t))]/p581.0, the corresponding t, approximately9.0 months, is an esti- mateof the mean1/ /afii9838andthus an estimateof thehazardrateis /afii9838/p19/p581/9/p580.111 per month. An alternative is to use (7.2.5 )to estimate /afii9838,/afii9838/p19/p5821/198 /p580.107, which is veryclose to the graphical estimate. Weibull Distribution The Weibull cumulative distribution function is F(t)/p581/p57exp[/p57(/afii9838t)/p65] t/p570,/afii9828/p570,/afii9838/p570( 8 .2.3) The probabilityplot for the Weibull distribution is based on the relationship logt/p58log1 /afii9838 /p591 /afii9828log/p3log1 1/p57F(t)/p4(8.2.4)  205 between tand the cumulative distributionfunction Foftobtained from (8.2.3 ). This relationship is linear between log tand the function log (log/p431/[1/p57F(t)]/p44). Thus, a Weibull probabilityplot is a graph of log( t/p7/p71/p8) and log (log/p431/ [1/p57F/p19(t/p7/p71/p8)]/p44), where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8), for example, (i/p570.5)/n, for i/p581,...,n. The shape parameter /afii9828is estimatedgraphicallyas the reciprocal of the slope of the straight line fitted to the graph. If the fitted line is appropriate, then atlog(log/p431/[1/p57F(t)]/p44)/p580, the corresponding log (t)is an estimate of log (1//afii9838) from (8.2.4 ). This fact can be used to estimate 1/ /afii9838and thus /afii9838graphicallyfrom a Weibull probabilityplot. At log (log/p431/[1/p57F(t)]/p44)/p580.5,(8.2.4 )reduces to logt/p58log(1//afii9838)/p590.5//afii9828. This equation can be used to estimate /afii9828. Estimatesoftheparameterscan alsobeobtainedfromthemethoddescribed in Chapter 7 if the Weibull distribution appears to be a good fit graphically.The following hypothetical example illustrates the use of the Weibull probabil-ityplot. The small number of observations used in the example is onlyforillustrative purposes. In practice, manymore observations are needed toidentifyan appropriate theoretical model for the data. Example 8.3 Six mice with brain tumors have survival times, in months of 3, 4, 5, 6, 8, and 10. Log( t/p7/p71/p8) plotted against log (log/p431/[1/p57(i/p570.5)/6]/p44) for i/p581,...,6 is shown in Figure 8.6. A straight line is fitted to the data point by eye. From the fitted line, at log (log/p431/[1/p57F(t)]/p44)/p580, the corresponding log(t)/p581.9, and thus an estimate of 1/ /afii9838is approximately6.69 [ /p58exp(1.9)] months and an estimate of /afii9838is 0.150. At log (log/p431/[1/p57F(t)]/p44)/p580.5, the corresponding log (t)/p582.09, and thus an estimate of /afii9828/p580.5/(2.09 —1.9)/p582.63. The maximum likelihood estimates of /afii9828and /afii9838obtained from the SAS procedure LIFEREG are 2.75 and 0.148, respectively. The graphical estimatesof/afii9828and/afii9838are close to the MLE. Lognormal Distribution If the survival time tfollows a lognormal distribution with parameters /afii9839and /afii9846/p17, log tfollows the normal distribution with mean /afii9839and variance /afii9846/p17. Consequently, (logt/p57/afii9839)//afii9846has the standard normal distribution. Thus, the lognormal distribution function can be written as F(t)/p58/afii9818 /p1logt/p57/afii9839 /afii9846/p2t/p570( 8 .2.5) where /afii9818(·)is the standard normal distribution function and /afii9839and /afii9846are, respectively, the mean and standard deviation of log t. A probabilityplot for the lognormal distribution is based on the following relationship obtained from (8.2.5 ): logt/p58/afii9839/p59/afii9846/afii9818/p92/p16(F(t)) (8 .2.6)206       Figure 8.6 Weibull probabilityplot of the data in Example 8.3. The function /afii9818/p92/p16(·)is the inverse of the standard normal distribution func- tion or its 100 Fpercentile. This relationship is linear between the value logtand the function /afii9818/p92/p16(F(t)). Thus, a log-normal probabilityplot is a graph of log( t/p7/p71/p8) versus /afii9818/p92/p16(F/p19(t/p7/p71/p8)), where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8). From (8.2.6 ),a t/afii9818/p92/p16(F(t))/p580, log t/p58/afii9839; and at, /afii9818/p92/p16(F(t))/p581,/afii9846/p58logt/p57/afii9839. These facts can be used to estimate /afii9839and/afii9846from a straight line fitted to the graph. Example 8.4 In a studyof a new insecticide, 20 insects are exposed. Survival times in seconds are 3, 5, 6, 7, 8, 9, 10, 10, 12, 15, 15, 18, 19, 20, 22,25, 28, 30, 40, and 60. Suppose that prior experience indicates that the survivaltime follows a lognormal distribution; that is, some insects might react to theinsecticide veryslowlyand not die for a long time. The log( t/p7/p71/p8) versus /afii9818/p92/p16[(i/p570.5)/20], i/p581,...,20, are plotted in Figure 8.7. The plot shows a reasonablystraight line. From the fitted line, at /afii9818/p92/p16(F(t))/p580, log tis an estimate of /afii9839, which is equal to 2.64, and at /afii9818/p92/p16(F(t))/p581, log t/p583.4 and thus /afii9846/p583.4/p572.64/p580.76. /afii9818/p92/p16(F(t)) can be obtained byapply ing Microsoft Excel function NORMSINV.  207 Figure 8.7 Lognormal probabilityplot of the data in Example 8.4. Log-Logistic Distribution The log-logistic distribution function is F(t)/p58/afii9825t/p65 1/p59/afii9825t/p65t/p570,/afii9828/p570,/afii9825/p570( 8 .2.7) A probabilityplot for the log-logistic distribution is based on the following relationship obtained from (8.2.7 ): logt/p581 /afii9828log/p31 1/p57F(t)/p571/p4/p571 /afii9828log/afii9825 (8.2.8 ) Thus, a log-logistic probabilityplot is a graph of log( t/p7/p71/p8) versus log (/p431/ [1/p57F/p19(t/p7/p71/p8)]/p44/p571), where F/p19(t/p7/p71/p8) is an estimate of F(t/p7/p71/p8), for example, (i/p570.5)/n, fori/p581,...,n. From (8.2.8 ), at log /p43[1/(1/p57F)]/p571/p44/p580, log t/p58/p57(1//afii9828) log/afii9825; and at log /p43[1/(1/p57F)]/p571/p44/p581, log t/p58(1//afii9828)(1/p57log/afii9825). These facts can be used to estimate /afii9828and/afii9825. The following example illustrates the log-logistic probabilityplot. Example 8.5 Consider the following survival times of 10 experimental rats in days: 8, 15, 25, 30, 50, 90, 95, 100, 150, and 300. Figure 8.8 plots log( t/p7/p71/p8)208       Figure 8.8 Log-logistic probabilityplot of the data in Example 8.5. against log (/p431/[1/p57(i/p570.5)/10]/p44/p571) for i/p581,...,10. To estimate /afii9828and/afii9825, from the fitted line, at log (/p431/[1/p57F(t)]/p44/p571)/p580, log t/p584.0; and at log (/p431/ [1/p57F(t)]/p44/p571)/p581, log t/p584.6. Thus, we have two equations: 4.0/p58/p571 /afii9828log/afii9825and 4.6 /p581 /afii9828(1/p57log/afii9825) From these two equations, /afii9828/p24/p581.667 and /afii9825/p24/p580.0013. 8.3 HAZARD PLOTTING Hazard plotting (Nelson 1972, 1982 )is analogous to probabilityplotting, the principal difference being that the survival time (or a function of it )is plotted against the cumulative hazard function (or a function of it )rather than the distribution function. Hazard plotting is designed to handle censored data.Similar to probabilityplotting, estimates of parameters in the distribution canbe determined from the hazard plot with little computational effort. To determine if a set of survival time with censored observation is from a given theoretical distribution, we construct a hazard plot byplotting thesurvival time (or a function of it )versus an estimation cumulative hazard (or  209 a function of it ). The cumulativehazard function can be estimated byfollowing the steps below. Step 1.Orderthe nobservationsinthe samplefromsmallesttolargest without regard to whether theyare censored. If some uncensored and censoredobservations have the same value, theyshould be listed in random order. Inthe list of ordered values, the censored data are each marked with a plus. Step 2.Number the ordered observations in reverse order, with nassigned to the smallest data value, n/p571 to the second smallest, and so on. The numbers so obtained are called K values orreverse-order numbers . For the uncensored observation, Kis the number of subjects still at risk at that time. Step 3.Obtain the corresponding hazard value for each uncensored observa- tion. Censored observations do not have a hazard value. The hazard value foran uncensored observationis 1 /K. This is the fraction of the Kindividuals who survived that length of time and then failed. It is an observed conditionalfailure probabilityfor an uncensored observation. Step 4.For each uncensored observation, calculate the cumulative hazard value. This is the sum of the hazard values of the uncensored observation andof all preceding uncensored observations. For tied uncensored observations,the cumulative hazard is evaluated onlyat the smallest Kamong the uncen- sored observations. The table in the following example illustrates the procedure.Example 8.6 Consider the remission data of the 21 leukemia patients receiving 6-MP in Example 3.3. Table 8.3 illustrates the procedure for estima-ting the cumulative hazard function. We now discuss the basic idea underlying hazard plotting for the exponen- tial, Weibull, lognormal, and log-logistic distributions. Exponential Distribution The exponential distribution has constant hazard function h(t)/p58/afii9838. Thus, the cumulative hazard function is H(t)/p58/afii9838t (8.3.1 ) From (8.3.1 ), the time can be written as a linear function of the cumulative hazard H, t/p581 /afii9838 H(t)( 8 .3.2) Thus, tplots as a straight-line function of H. The slope of the fitted line is the210       Table 8.3 Estimation of Cumulative Hazard Reversed Cumulative Order, Hazard, Hazard, tK 1/K H /p19(t) 62 10 .048 6/p59 20 61 90 .053 61 80 .056 0 .156 71 70 .059 0 .215 9/p59 16 10 15 0 .067 0 .281 10/p59 14 11/p59 13 13 12 0 .083 0 .365 16 11 0 .091 0 .456 17/p59 10 19/p59 9 20/p59 8 22 7 0 .143 0 .598 23 6 0 .167 0 .765 25/p59 5 32/p59 4 32/p59 3 34/p59 2 35/p59 1 mean survival time 1/ /afii9838of the distribution. More simply, 1/ /afii9838is the value of t when H(t)/p581. This fact is used to estimate 1/ /afii9838from an exponential hazard plot. Example 8.7 Using the estimated cumulative hazard values H/p19(t) in Table 8.3, we construct the exponential hazard plot in Figure 3.5 byplotting eachexact time tagainst its corresponding H/p19(t). The configuration appears to be reasonablylinear, suggesting that the exponential distribution provides areasonable fit. In Chapter 3 we see that the Weibull distribution gives a betterfit than the exponential. We use the data here just to demonstrate how theparameter can be estimated. To find an estimate for the mean remission time of the leukemia patients, we can use H(t)/p580.5 since the time for which H/p581 is out of the range of the horizontal axis. At H(t)/p580.5,t/p5816.9, from (8.3.2 ), an estimate of /afii9838is 0.5/16.9 /p580.0296. Thus, an estimate of the mean remission time is 34 weeks.  211 Figure 8.9 Cumulativehazard functionsofthe Weibulldistributionwith /afii9828/p580.5, 1, 2,4.Weibull Distribution The Weibull distribution has the hazard function h(t)/p58/afii9838/afii9828(/afii9838t)/p65/p92/p16 t/p570 The cumulative hazard function is H(t)/p58(/afii9838t)/p65 t/p570( 8 .3.3) and is plotted in Figure 8.9 for four different values of /afii9828: 0.5, 1, 2, and 4. From (8.3.3 ), the time tcan be written as a function of the cumulative hazard function, that is, t/p581 /afii9838[H(t)]/p16/p30/p65 (8.3.4) Taking the logarithm of (8.3.4 ), we obtain logt/p58log1 /afii9838/p591 /afii9828logH(t)( 8 .3.5) Since log tis a linear function of log H(t), a plot of log tagainst log H(t)i sa straight line. For log H(t)/p580o r H(t)/p581,(8.3.5 )reduces to log t/p58log(1//afii9838), and thus the corresponding time tequals 1/ /afii9838. This fact is used to estimate 1/ /afii9838 and consequently, /afii9838. The slope of the fitted straight line is 1/ /afii9828,o ra t logH(t)/p581,(8.3.5 )can be written as /afii9828/p581/(logt/p59log/afii9838). This equation can be used to estimate /afii9828.212       Figure 8.10 Weibull hazard plot of the data in Example 8.8. Example 8.8 Consider the following survival times in months of 14 patients: 15, 25, 38, 40 /p59, 50, 55, 65, 80 /p59, 90, 140, 150 /p59, 155, 250 /p59, 252. Figure 8.10 is the hazard plot with log tversus log H(t) of the data. From the fitted line, at log H(t)/p580, log t/p584.8.Thus, t/p58121.5 and the estimate of /afii9838is /afii9838/p19/p581/t/p580.0082. Similarly, at, log H(t)/p581, log t/p585.6, and thus /afii9828/p24/p581/ (5.6/p574.8)/p581.25. Lognormal Distribution The densityfunction of a lognormal distribution is f(t)/p581 t/afii9846/p402/afii9843exp/p3/p571 2/afii9846/p17(logt/p57/afii9839)/p17/p4 /p581 t/afii9846g/p1logt/p57/afii9839 /afii9846/p2t/p570 (8.3.6 ) where g(x) is the standard normal densityfunction. The lognormal cumulative distribution function is F(t)/p58/afii9818/p1logt/p57/afii9839 /afii9846/p2t/p570( 8 .3.7)  213 Figure 8.11 Cumulative hazard functions of the lognormal distribution with /afii9846/p580.1, 0.5, 1.0.where /afii9818(·)is the standard normal distribution function. Thus, by (2.10), the hazard function can be written as h(t)/p581 t/afii9846g/p1logt/p57/afii9839 /afii9846/p2 1/p57/afii9818/p1logt/p57/afii9839 /afii9846/p2(8.3.8) The cumulative hazard function, plotted in Figure 8.11 for three values of /afii9846,i s H(t)/p58/p57log/p31/p57/afii9818/p1logt/p57/afii9839 /afii9846/p2/p4(8.3.9) From (8.3.9 ), the logarithm of the survival time tas a function of the cumulative hazard His logt/p58/afii9839/p59/afii9846/afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]( 8 .3.10) where /afii9818/p92/p16(·)is the inverse of the standard normal distribution function. Thus, log tis a linear function of /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]. The log-normal hazard plot is a graph of log tversus /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]. From (8.3.10 ),a t /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p580, log t/p58/afii9839; and at /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p581, log t/p58/afii9839/p59/afii9846. These facts can be used to estimate /afii9839and/afii9846. Example 8.9 Consider the following remission times in months of 18 cancer patients: 4, 5, 6, 7, 8, 9 /p59, 12, 12 /p59, 13, 15, 18, 20, 25, 26 /p59,2 8/p59, 35, 35/p59, 56. Figure 8.12 gives the log-normal hazard plot. From the fitted line by eye, at /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p580, log t/p582.8; and at /afii9818/p92/p16[1/p57e/p92/p38/p7/p82/p8]/p581,214       Figure 8.12 Lognormal hazard plot of the data in Example 8.9.logt/p583.76. Thus, the estimate of /afii9839is 2.8 and the estimate of /afii9846is 3.76/p572.8/p580.96. Log-Logistic Distribution The cumulative hazard function of the log-logistic distribution is H(t)/p58log(1/p59/afii9825t/p65) This equation can be written as logt/p581 /afii9828log/p43exp[ H(t)]/p571/p44/p571 /afii9828log/afii9825 (8.3.11 ) Thus, log tis a linearfunction of log /p43exp[ H(t)]/p571/p44. A log-logistic hazardplot is a graph of log tversus log /p43exp[ H(t)]/p571/p44. From (8.3.11 ),a t log/p43exp[ H(t)]/p571/p44/p580, log t/p58/p57(1//afii9828) log/afii9825; and at log /p43exp[ H(t)]/p571/p44/p581, logt/p58(1//afii9828)/p57(1//afii9828) log/afii9825. These facts can be used to estimate /afii9828and/afii9825. 8.4 COX--SNELL RESIDUAL METHOD The Cox —Snell (1968 )residual method can be applied to anyparametric model. The Cox —Snell residual r/p71for the ith individual with observed survival time t/p71, uncensored or censored, is defined as r/p71/p58/p57logS/p19(t/p71) i/p581, 2, ...,n (8.4.1)—   215 where S/p19(t) is the estimated survival function based on the MLE of the parameters.If the observed t/p71is censored, the corresponding r/p71is also censored. Since the cumulative hazard function H(t)/p58/p57logS(t), the Cox —Snell residual r/p71is an estimated cumulated hazard value at t/p71. The important propertyof the Cox —Snell residual is that if the model selected fits the data, r/p71’s follow the unit exponential distribution with densityfunction f/p48(r)/p58e/p92/p80. Let S/p48(r) denote the survival function of the Cox —Snell residual r/p71. Then S/p48/p58/p25/p27/p80f/p48(x)dx/p58/p25/p27/p80e/p92/p86dx/p58e/p92/p80, and /p57logS/p48(r)/p58/p57log(e/p92/p80)/p58r (8.4.2 ) Let S/p19/p48(r) denote the Kaplan —Meier estimate of S/p48(r).It is clear from (8.4.2 ) that the plot of r/p71versus /p57logS/p19/p48(r/p71)should be a straight line with unit slope and zero intercept if the fitted survival distribution is appropriate, regardlessof the form of the distribution. The procedure for using Cox —Snell residuals can be summarized as follows. 1. Use the methods shown in Sections 7.1 to 7.7 to find the MLE of the parameters of the selected theoretical distribution. 2. Calculate Cox —Snell residuals r/p71/p58/p57logS/p19(t/p71),i/p581, 2, ...,n, where S/p19(t/p71) is the estimated survival function with the MLE of the parameters. 3. Applythe Kaplan —Meier method to estimate the survival function S/p48(r) of the Cox —Snell residuals r/p71’s obtained in step 2, then using the estimate S/p19/p48(r), calculate /p57logS/p19/p48(r/p71),i/p581, 2, ...,n. 4.Plot r/p71versus /p57logS/p19/p48(r/p71),i/p581, 2, ...,n. If the plot is closed to a straight line with unit slope and zero intercept, the fitted distribution is appropri-ate. From (8.4.1 ), if an individual survival time is right-censored, say, t/p62/p71and the fitted model is correct, the corresponding Cox —Snell residual /p57logS(t/p62/p71)/p58H(t/p62/p71) is smaller than the residual evaluated at an uncensored observationwiththesame value t/p71since H(t) is a monotone-increasingfunction oft. To take this into account, two modified Cox —Snell residuals have been proposed for censored observations (Crowleyand Hu, 1977 ). One is based on the mean, and the other is based on the median (/p58log2 /p580.693 )of the unit exponential distribution byassuming that difference between H(t/p71)and H(t/p62/p71) also follows the unit exponential distribution. For a censored observation t/p62/p71, the modified residual r/p62/p71is defined as r/p62/p71/p58r/p71/p591( 8 .4.3) or r/p62/p71/p58r/p71/p590.693 where r/p71/p58/p57logS/p19(t/p71)( 8 .4.4) Example 8.10 Consider the tumor-free time data observed from rats fed with saturated diets in Table 3.4. We select the lognormal distribution for this216       Figure 8.13 Cox —Snell residual plot for the fitted lognormal model on the tumor-free time data for rats fed with saturated diets.set of data for illustrative purposes. Using methods discussed in Chapter 7, the MLE of the parameters obtained are /afii9839/p584.76458 and /afii9846/p580.56053. We then calculate the Cox —Snell residuals r/p71/p58/p57logS(t/p71)/p58/p57log[1 /p57F(t/p71)], where F(t) is the distribution function of the lognormal distribution. An easywayto compute r/p71for the lognormal distributionis to use the relationshipbetween the normal and lognormal distributions, i.e., the distribution function of thelognormaldistribution, F(t), is equivalent to /afii9818[(logt/p57/afii9839)//afii9846], where /afii9818()is the distribution function of the standard normal distribution. We can use Micro-soft Excel function NORMSDIST to calculate /afii9818(t). Thus, for the lognormal distribution, S(t/p71)/p581/p57/afii9818([log( t/p71)/p574.76458] /0.56053) Using the specific notation of NORMSDIST, ln for log, r/p71/p58/p57ln(1/p57normsdist /p43[ln(t/p71)/p574.76458] /0.56053 /p44) The r/p71’s so obtained are given in Table 8.4. The next step is to obtain the Kaplan —Meier estimate of the survival function S(r/p71), and compute /p57logS(r/p71). These values are also given in Table 8.4. Figure 8.13 gives the graph of r/p71versus /p57logS/p19/p48(r/p71),i/p581,...,22. The graph is close to a straight line with unit slope and zero intercept. Therefore, a—   217 Table 8.4 Kaplan--Meier Estimate of Survivorship Function for the Cox--Snell Residuals from the Fitted Lognormal Model on Tumor-Free Time Data for Rats Fed with Saturated Diets tr /p63 S/p19/p48(r)/p64 /p57logS/p19/p48(r) 0.000 1 .000 0 .000 43 0 .037 0 .967 0 .034 46 0 .049 0 .933 0 .069 56 0 .098 0 .900 0 .105 58 0 .110 0 .867 0 .143 68 0 .181 0 .833 0 .182 75 0 .239 0 .800 0 .223 79 0 .275 0 .767 0 .266 81 0 .294 0 .733 0 .310 86 0 .342 0 .667 0 .405 86 0 .342 0 .667 0 .405 89 0 .373 0 .633 0 .457 96 0 .447 0 .600 0 .511 98 0 .469 0 .567 0 .568 105 0 .548 0 .533 0 .629 107 0 .571 0 .500 0 .693 110 0 .606 0 .467 0 .762 117 0 .690 0 .433 0 .836 124 0 .776 0 .400 0 .916 126 0 .800 0 .367 1 .003 133 0 .889 0 .333 1 .099 142 1 .004 0 .267 1 .322 142 1 .004 0 .267 1 .322 165 1 .305 0 .233 1 .455 170/p59 1.371/p59 200/p59 1.769/p59 200/p59 1.769/p59 200/p59 1.769/p59 200/p59 1.769/p59 200/p59 1.769/p59 200/p59 1.769/p59 /p63r, ordered Cox —Snell residuals from the fitted lognormal model. /p64S/p48(r), Kaplan —Meier estimate of survivorship function for the Cox —Snell residuals. lognormal model maybe appropriate for the tumor-free times observed. In Chapter 9 (Example 9.2 )we will see that the lognormal model was not rejected based on a goodness-of-fit test. Thus the result is consistent with thoseobtainedbyusing the analy ticalmethod.A weakness of theCox —Snell residual method is that the plot does not indicate the kind of departure the data havefrom the model selected if the configuration is not linear.218       Bibliographical Remarks Probabilityplotting has been widelyused since Daniel’s (1959 )classical work on the use of half-normal plot. A quite complete and excellent treatment ofprobabilityplotting is given byKing (1971 ). Although examples given are applications to industrial reliability, its interpretation of probability plots ofmanydistributions, such as the uniform, lognormal, Weibull, and gamma, areapplicable to biomedical research. Recent applications of probabilityplotting include Leitner et al. (1986 ), Horner (1987 ), Waters et al. (1991 ), and Tsumagari et al. (2000 ). Hazard plotting was developed byNelson (1972, 1982 ). Applications in- cluded Gore (1983 )and Wurpel et al. (1986 ). EXERCISES 8.1Show that the Cox —Snell residuals defined in (8.4.1 )follow the unit exponential distribution with densityfunction f(r)/p58exp(/p57r). 8.2Consider the following survival times of 16 patients in weeks: 4, 20, 22, 25, 38, 38, 40, 44, 56, 83, 89, 98, 110, 138, 145, and 27. (a)Does the exponential distribution provide a reasonable fit to the survival data? Use the probabilityplotting technique. (b)Estimate graphicallythe parameter /afii9838of the exponential distribution and consequently, the mean survival time. 8.3To computerize patients’ records, a data clerk is hired to transcribe medical data from the patients’ charts to computer coding forms. Thenumber of correct entries between errors is listed in chronological orderof occurrence over a period of five days as follows: 73, 12, 40, 65, 100,15, 70, 40, 110, 64, 200, 6, 90, 102, 20, 102, 90, 34. The assumption is thatthe data clerk, during the five days, would not change her error rateappreciably. Use the technique of probability plotting to evaluate theassumption above. What is your conclusion? 8.4Twenty-five rats were injected with a give tumor inoculum. Their times, in days, to the development of a tumor of a certain size are given below. 30 53 77 91 118 38 54 78 95 12045 58 81 101 12546 66 84 108 13450 69 85 115 135 Which of the distributions discussedin this chapter providea reasonable fit to the data? Estimate graphicallythe parameters of the distribution chosen. 219 8.5In a clinical study, 28 patients with cancer of the head and neck did not respondtochemotherapy.Theirsurvivaltimesinweeksaregivenbelow. 1.7 8.3 14.0 22.7 6.0 /p5913.1/p59 5.1 9.6 15.9 33.0 7.4 /p5913.4/p59 5.3 11.3 16.7 3.7 /p598.0/p5916.1/p59 6.0 12.1 17.0 5.0 /p598.3/p59 8.3 12.3 21.0 5.9 /p599.1/p59 (a)Make a hazard plot for each of the following distributions:exponen- tial, Weibull, lognormal, and log-logistic. (b)Which distribution provides a reasonable fit to the data? Estimate graphicallythe parameters of the distribution chosen. 8.6Thirty-one patients with advanced melanoma treated with combined chemotherapy, immunotherapy, and hormonal therapy have survivaltimes as given below. 26.3/p5916.1 24.0 4.3 31.3 /p59 94.0 49.6 77.9 97.6 /p5917.6/p59 9.1 27.3 16.6 /p597.3 16.3 34.6/p5961.9/p593.4 75.6 /p59 9.4 46.6 /p5910.9 14.3 25.7 22.4 /p5913.0 56.4 88.7 7.1 64.4 /p599.1 (a)Make a hazard plot for each of the following distributions:exponen- tial, Weibull, lognormal, and log-logistic. (b)Which distribution provides a reasonable fit to the data? Estimate the parameters of the distribution chosen. 8.7Consider the survival times of the hypernephroma patients in Exercise Table 3.1 (see Exercise 4.5 ). Make a hazard plot for the distribution you chose in Exercise 6.8. Did you make a good selection? If not, try twoother distributions. 8.8Consider the following survival times in weeks of 10 mice with injection of tumor cells: 5, 16, 18 /p59, 20, 22 /p59,2 4/p59, 25, 30 /p59, 35, 40 /p59. Make an exponential hazard plot. Does the exponential distribution provide areasonable fit? If not, is the lognormal distribution better? 8.9Consider the following survival times in months of 25 patients with cancer of the prostate. Use a graphical method to see if the survival timeof prostate cancer patients follows the exponential distribution with/afii9838/p580.01: 2, 19, 19, 25, 30, 35, 40, 45, 45, 48, 60, 62, 69, 89, 90, 110, 145, 160, 9 /p59,1 0/p59,2 0/p59,4 0/p59,5 0/p59, 110/p59, 130/p59. 8.10Make a log-logistic hazard plot of the following data and estimate the two parameters: 20, 30, 32 /p59, 40, 60, 100, 150, 200 /p59, 300.220       CHAPTER 9 Tests of Goodness of Fit and Distribution Selection In Chapter 8 we discuss three graphical methods for checking if a parametricdistribution fits the observed data. Parametric distributions can be groupedintofamilies.First, anygivendistributionwithdifferentparametervaluesformsa family. Second, if a distribution includes other distributions as its specialcases, this distribution is a nesting (larger )family of these distributions. For example, the distributions introduced in Chapter 6 belongto more than onenested family. First, the Weibull distribution reduces to the exponential when/afii9828/p581. Therefore, the exponential distribution is a special case of the Weibull and the two distributions are said to belongto one family, the Weibull family.Second, consider the standard gamma distribution; when /afii9828/p581, it reduces to the exponential, and when /afii9838/p58/p16/p17 and/afii9828/p58/p16/p17/afii9840, it becomes the chi-square distribution with /afii9840degrees of freedom. Thus, the gamma distribution includes the exponential and chi-square as a family. Now let us consider the generalizedgamma distribution. It reduces to the exponential if /afii9825/p58/afii9828/p581, the Weibull if /afii9828/p581, the lognormal if /afii9828/p59/p45, and the gamma if /afii9825/p581. Thus, the generalized gamma distribution includes these four distributions and represents a largefamily of distributions. The relationship of the generalized gamma distributionto the exponential, Weibull, lognormal, and gamma distributions allows us toevaluate the appropriateness of these distributions relative to each other andto a more general distribution. It is known that the generalized gammadistribution is a special case of the generalized F-distribution and therefore belongs to the generalized Ffamily (Kalbfleisch and Prentice, 1980 )Because of its complexity, we do not cover the generalized Ffamily. In this chapter we discuss several analytical procedures for comparing parametric distributions and assessingg oodness of fit. In Section 9.1 weintroduce several widely used statistics for testingthe appropriateness of adistribution. Readers who are not familiar with linear algebra or are notinterested in the mathematical details may skip this section without loss ofcontinuity.In Section 9.2 we discuss statistics for testingwhether a distribution 221 is appropriate by comparingit with other distributions in the same family or a more general family. Section 9.3 covers the selection of a distribution basedon Baysian information criteria. Section 9.4 covers the statistics for testingwhethera given distributionwith knownparameters is appropriate. Allthe teststatistics discussed in Sections 9.1 to 9.4 are based on asymptotic likelihoodinferences. In Section 9.5 we introduce the test statistic of Hollander andProschan (1979 )for testingwhether a distribution with g iven parameters is appropriate. Computer codes for BMDP or SAS that can be used to carry out the test procedures are provided. 9.1 GOODNESS-OF-FIT TEST STATISTICS BASED ON ASYMPTOTIC LIKELIHOOD INFERENCES We take the exponential distribution as an example to see how to construct statistics to test whether it is appropriate for the observed survival times. Asnoted in Chapter 6, the Weibull family with /afii9828/p581, the gamma family with /afii9828/p581, and the generalized gamma family with /afii9825/p58/afii9828/p581 reduce to the exponential distribution. Therefore, to test if the exponential distribution isappropriate for the observed survival time, we can first fit a Weibull distribu-tion and test if /afii9828/p581, or fit a gamma distribution, then test if /afii9828/p581, or fit a generalized gamma distribution, then test if /afii9825/p58/afii9828/p581. Similarly, to test whether the family of Weibull distributions, or the gamma distributions, or thelognormal distributions is appropriate for the survival data observed, we canfit a generalized gamma distribution (their nestingdistribution )and then test if/afii9828/p581, or/afii9825/p581, or with /afii9828/p59/p45,respectively.Thus,testingtheappropriateness of a family of distributions is equivalent to testingwhether a subset of theparameters in its nestingdistribution equal to some specific values. If the datacan be assumed to follow a certain distribution but the values of its parametersare uncertain, we need to test only that the parameters are equal to certainvalues. In the following, we separately introduce test statistics for testingwhether some of the parameters in a distribution are equal to certain valuesand whether all parameters in a distribution are equal to certain values.Readers who are interested in a detailed discussion of these statistics arereferred to Kalbfleisch and Prentice (1980 ). 9.1.1 Testing a Subset of Parameters in a Distribution Letb/p58(b/p16,b/p17)denote all the parameters in a parametric distribution, where b/p16andb/p17are subsets of parameters, and let the hypothesis be H/p15:b/p17/p58b/p15(9.1.1 ) whereb/p15is a vector of specific numbers. Let b/p19be the MLE of b,b/p19/p16(b/p15)the MLE of b/p16givenb/p17/p58b/p15, andV/p19/p17(b/p19)the submatrix of the covariance matrix in222         (7.1.5 ),V/p19(b/p19), correspondingto b/p17. Under H/p15and some mild assumptions, both of the followingtwo statistics have an asymptotic chi-square distribution withdegrees of freedom equal to the dimension of (or the number of parameters in ) b/p17. Log-likelihood ratio statistic: X/p42/p582[l(b/p19)/p57l(b/p19/p16(b/p15),b/p15)] (9.1.2 ) Wald statistic: X/p53/p58(b/p19/p17/p57b/p15)/p30V/p19/p92/p16/p17(b/p19)(b/p19/p17/p57b/p15)( 9.1.3 ) If the number of parameters in b/p17is equal to q, for a given significant level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p79/p11/p63when the likelihood ratio statistic is used; or if X/p53/p57/afii9851/p17/p79/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p79/p11/p16/p92/p63/p30/p17,(two-sided test )orX/p53/p57/afii9851/p17/p79/p11/p63(one-sided test ) when the Wald’s statistic is used, where /afii9851/p17/p79/p11/p63,/afii9851/p17/p79/p11/p63/p30/p17and/afii9851/p17/p79/p11/p16/p92/p63/p30/p17are the 100(1/p57/afii9825), 100 (1/p57/afii9825/2), and 100 /afii9825/2 percentile points of the chi-square dis- tribution with qdegrees of freedom; that is, P(/afii9851/p17/p79/p57/afii9851/p17/p79/p11/p63)/p58/afii9825andP(/afii9851/p17/p79/p57/afii9851/p17/p79/p11/p63/p30/p17)/p58P(/afii9851/p17/p79/p58/afii9851/p17/p79/p11/p16/p92/p63/p30/p17)/p58/afii9825 2 Example 9.1 Suppose that we wish to test whether the observed data are from an exponential distribution. We can use a Weibull distribution and testwhether its shape parameter, /afii9828, is equal to 1. The Weibull distribution has two parameters, /afii9838and/afii9828; thusb/p58(/afii9838,/afii9828)and the null and alternativehypotheses are: H/p15:/afii9828/p581(the underlyingdistribution is an exponential distribution ) (9.1.4 ) H/p16:/afii9828/p341(the underlyingdistribution is a Weibull distribution ) Letb/p19/p58(/afii9838/p19,/afii9828/p19)be the MLE of b,l/p53(b/p19)/p58l/p53(/afii9838/p19,/afii9828/p19) andl/p35(/afii9838/p19)be the log-likelihood of the Weibull and exponential distributions, respectively, l/p35(/afii9838/p19)/p89l/p53(/afii9838/p19(1),1), where /afii9838/p19(1)is the MLE of /afii9838in the Weibull distribution given /afii9828/p581. The log-likelihoodratio andWald statisticsdefined in (9.1.2 )and(9.1.3 )in this case become X/p42/p582[l/p53(/afii9838/p19,/afii9828/p19)/p57l/p53(/afii9838/p19(1),1)] (9.1.5 ) and X/p53/p58(/afii9828/p24/p571)V/p19/p92/p16/p17(/afii9838/p19,/afii9828/p24)(/afii9828/p24/p571) (9 .1.6) --   223 respectively, where V/p19/p17(/afii9838/p19,/afii9828/p24) is the second diagonal element of the covariance matrix V/p19(/afii9838/p19,/afii9828/p24)/p58/p57/p1/p42/p17l/p53(/afii9838/p19,/afii9828/p24) /p42/afii9838/p17/p42/p17l/p53(/afii9838/p19,/afii9828/p24) /p42/afii9838/p42/afii9828 /p42/p17l/p53(/afii9838/p19,/afii9828/p24) /p42/afii9828/p42/afii9838/p42/p17l/p53(/afii9838/p19,/afii9828/p24) /p42/afii9828/p17/p2/p92/p16 (9.1.7) and V/p19/p92/p16/p17(/afii9838/p19,/afii9828/p24)/p58/p57[/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9838/p17][/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9828/p17]/p57(/p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9838 /p42/afii9828)/p17 /p42/p17l/p53(/afii9838/p19,/afii9828/p24)//p42/afii9838/p17(9.1.8) For a given significant-level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63, when the likelihood ratio statistic is used; or if X/p53/p57/afii9851/p17/p16/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p16/p11/p16/p92/p63/p30/p17, when the Wald statistic is used. It must be pointed out that failure to reject H/p15in(9.1.4 )does not imply that anexponentialdistributionprovidesthe bestfit to thedata. On the other hand,rejection of H/p15does not indicate that a Weibull distribution is the choice either. Further testingof other distributions is needed. The details andexamples are given in Section 9.2. Since the gamma and generalized gamma distribution also include the exponential as a special case, similar test statistics can be constructed to testthe null hypothesisthat the data are from the exponentialdistributionby usingthe gamma, the generalized gamma, or the extended generalized gammadistribution. 9.1.2 Testing All Parameters in a Distribution To test whether all of the parameters in bequal a given set of known values b/p15, the null hypothesis is H/p15:b/p58b/p15(9.1.9 ) and the followingthree test statistics can be used.Log-likelihood ratio statistic: X/p42/p582[l(b/p19)/p57l(b/p15)] (9.1.10 ) Wald statistic: X/p53/p58/p57(b/p19/p57b/p15)/p30/p42/p17l(b/p15) /p42b/p42b/p30 (b/p19/p57b/p15) /p3or/p58/p57(b/p19/p57b/p15)/p30/p42/p17l(b/p19) /p42b/p42b/p30(b/p19/p57b/p15)/p4 (9.1.11 )224         Score statistic: X/p49/p58/p3/p42l(b/p15) /p42b/p4/p30/p3/p57/p42/p17l(b/p15) /p42b/p42b/p30/p4/p92/p16/p42l(b/p15) /p42b /p1or/p58/p3/p42l(b/p15) /p42b/p4/p30V/p19(b/p19)/p42l(b/p15) /p42b/p2 (9.1.12 ) whereV/p19(b/p19)is the estimated covariance matrix in (7.1.5 ). Under H/p15and the assumption that b/p19has approximately multinormal distribution, each of the three statistics has an asymptotic chi-square distribution with p(the dimension ofbor the number of parameters in b)degrees of freedom. For a given significant-level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p78/p11/p63, when the likelihood ratio statistic is used; or if X/p53/p57/afii9851/p17/p78/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p78/p11/p16/p92/p63/p30/p17, when the Wald statistic is used;or if X/p49/p57/afii9851/p17/p78/p11/p63/p30/p17orX/p49/p58/afii9851/p17/p78/p11/p16/p92/p63/p30/p17, when the score statistic is used. It must be pointed out that rejection of H/p15in(9.1.9 )means only that the given distribution with the known parameters b/p15, not the family of distribu- tions to which the given distribution belongs, is not appropriate for theobserved data. It is possible that a distribution with different b/p15in the family may be appropriate. 9.2 TESTS FOR APPROPRIATENESS OF A FAMILY OF DISTRIBUTIONS The usual method for testingwhether a distribution is appropriate for the observed data is to compare the distribution with a larger or more generalfamily that includes the distribution of interest as a special case (Hagar and Bain, 1970 ). Letl/p35(/afii9838),l/p53(/afii9838,/afii9828),l/p37(/afii9838,/afii9828),l/p42/p44(/afii9839,/afii9846/p17), andl/p37/p37(/afii9825,/afii9838,/afii9828) denote, respectively, the log-likelihoodfunction definedin (7.1.1 )based on the exponential,Weibull, gamma, lognormal, and extended generalized gamma distribution, and l/p35(/afii9838/p19), l/p53(/afii9838/p19,/afii9828/p24),l/p37(/afii9838/p19,/afii9828/p24),l/p42/p44(/afii9839/p24,/afii9846/p24/p17), andl/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)denote the respective log-likelihood values where /afii9838/p19,(/afii9838/p19,/afii9828/p24),(/afii9838/p19,/afii9828/p24),(/afii9839/p24,/afii9846/p24/p17), and (/afii9825/p24,/afii9838/p19,/afii9828/p24)are the MLE. For example, the log-likelihood of the exponential distribution can be obtained from l/p35(/afii9838/p19)/p58/p80/p26 /p71/p14/p16log(/afii9838/p19e /p57/afii9838/p19t/p71)/p59/p76/p26 /p71/p14/p80/p62/p16log(e/p57/afii9838t/p62/p71)/p58rlog/afii9838/p19/p57/afii9838/p19/p80/p26 /p71/p14/p16t/p71/p57/afii9838/p19/p76/p26 /p71/p14/p80/p62/p16t/p62/p71 for a set of observed survival times t/p16,...,t/p80,t/p62/p80/p62/p16,...,t/p62/p76. The log-likelihood value and the estimated covariance matrix in (7.1.5 )and parameters for each of the distributions discussed in Sections 7.2 to 7.6 can be obtained from SASorBMDP.Theresults canbe used toconstructthelog-likelihoodratiostatisticand the Wald statistic defined in (9.1.2 )and (9.1.3 ). In the following, we        225 introduceseveral testsfor the appropriatenessofa familyofdistributionsbased on the log-likelihoods. Construction of the respective Wald statistics is left tothe reader as exercises. 1.Testing the hypothesis that the underlying distribution is exponential . The null hypothesis is H/p15:The underlyingdistribution is an exponential distribution If the Weibull distribution is used, testingthe null hypothesis above is equivalent to testingthe followingnull and alternative hypotheses: H/p15:/afii9828/p581(the underlyingdistribution is an exponential distribution ) H/p16:/afii9828/p341(the underlyingdistribution is a Weibull distribution ) Let/afii9838/p19(1)be the MLE of /afii9838in the Weibull distribution given /afii9828/p581, the log-likelihood ratio statistic is X/p42/p582[l/p53(/afii9838/p19,/afii9828/p24)/p57l/p53(/afii9838/p19(1),1)] (9.2.1 ) which has an asymptotic chi-square distribution with 1 degree of freedom. For a given level of significance /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63. Note that l/p53(/afii9838/p19(1),1)/p89l/p35(/afii9838/p19). Similarly, a log-likelihood ratio statistic can be constructed by using the gamma or the extended generalized gamma distribution. These will be left tothe reader as exercises. 2.Testing the hypothesis that the underlying distribution is Weibull . The null hypothesis is H/p15:The underlyingdistribution is a Weibull distribution We can use the extended generalized gamma distribution and test whether its parameter /afii9828equals 1. Thus the null and alternativehypotheses can be stated as H/p15:/afii9828/p581(the underlyingdistribution is a Weibull distribution ) H/p16:/afii9828/p341(the underlyingdistribution is an extended g eneralized gamma distribution ) Let/afii9838/p19(1)and/afii9825/p24(1)be the MLE of /afii9838and/afii9825in the extended generalized gamma distribution given /afii9828/p581. Accordingto Section 6.4, an extended g eneralized226         gamma distribution with /afii9828/p581 is a Weibull distribution. The likelihood ratio statistic is X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p37/p37(/afii9825/p24(1),/afii9838/p19(1), 1)] (9 .2.2) which follows asymptotically the chi-square distribution with 1 degree of freedom. H/p15is rejected at a significance level of /afii9825ifX/p42/p57/afii9851/p17/p16/p11/p63. Note that l/p37/p37(/afii9825/p24(1),/afii9838/p24(1), 1) /p89l/p53(/afii9838/p19,/afii9828/p24). 3.Testing the hypothesis that the underlying distribution is standard gamma . The null hypothesis is H/p15:The underlyingdistribution is a g amma distribution Followingthe same log ic in Section 6.4, the null hypothesis above is equivalent to the following if the extended generalized gamma distribution is used. H/p15:/afii9825/p581(the underlyingdistribution is a standard g amma distribution ) H/p16:/afii9825/p341(the underlying distribution is a generalized gamma distribution ). The likelihood test statistic is X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p37/p37(1,/afii9838/p19(1),/afii9828/p24(1))] (9 .2.3) where /afii9828/p24(1)and/afii9838/p19(1)are the MLE of /afii9828and/afii9838given /afii9825/p581, which has an asymptotic chi-square distribution with 1 degree of freedom under H/p15. The rejection rule is the same as that for the exponential or Weibull distribution.Note that l/p37/p37(1,/afii9838/p19(1),/afii9828/p24(1))/p89l/p37(/afii9838/p19,/afii9828/p24). 4. Testing the hypothesis that the underlying distribution is lognormal . The null hypothesis is H/p15:the underlyingdistribution is a log normal distribution The log-likelihood test statistic is X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p42/p44(/afii9839/p24,/afii9846/p24/p17)] which has an asymptotic chi-square distribution with 1 degree of freedom underH/p15. The rejection rule is the same as that for the exponential or Weibull distribution. For the log-logistic and extended generalized gamma distributions, it can be shown that a generalized F-distribution (Kalbfleisch and Prentice, 1980 ) includes the exponential, Weibull, lognormal, gamma, generalized gamma,        227 Table 9.1 Summary of Goodness-of-Fit Tests for Testing Whether a Family of Models Is Appropriate forthe Observed Data /p63 Hypothesized Model LL X/p42df Generalized gamma l/p37/p37Lognormal l/p42/p442(l/p37/p37/p57l/p42/p44)1 Gamma l/p372(l/p37/p37/p57l/p37)1 Weibull l/p532(l/p37/p37/p57l/p53)1 Exponential l/p352(l/p37/p37/p57l/p35)2 Exponential l/p352(l/p37/p57l/p35)1 Exponential l/p352(l/p53/p57l/p35)1 /p63LL, log-likelihood; X/p42, likelihood ratio chi-square statistic; df, degrees of freedom. extended generalized gamma, and log-logistic distributions as special cases. Therefore, one can follow the same logic to construct either the log-likelihoodratio or the Wald statistic to test the appropriateness of a family of generalizedgamma or log-logistic distributions. However, methods for testing the appro-priateness of a generalized F-distribution remain unknown. Unless we can find a more general distribution that includes the generalized F-distribution as a special case, there is no formal way to check whether the generalized F- distribution is appropriate. However, the generalized gamma distribution is arich family and includes a considerable number of distributions. It should besufficient for most applications. All the tests introduced in this section aresummarized in Table 9.1. As pointed out in Section 9.1, when usingany of the testingprocedures above, failure to reject H/p15does not imply that the hypothesized distribution provides a perfect fit to the data. On the other hand, rejection of H/p15does not mean that the distribution under the alternative hypothesis is the best choiceeither. In practice, with the help of available computer software, it is easy to fitseveral distributions simultaneously and then select the most appropriate one,usually the simplest one, as the final choice for the data. The followingexamples illustrate the procedure. Example 9.2 Consider the tumor-free times of the 30 rats that are fed with a saturated diet in Table 3.4. UsingSAS, we obtainthe MLE of the parametersand the log-likelihoods for the exponential, Weibull, lognormal, and generaliz-ed gamma distributions. The results are given in Table 9.2. For example, theMLE of /afii9838in the exponential distribution is 5.054 and the corresponding log-likelihood is /p5735.359, and the MLE of the two parameters in the Weibull distribution are /afii9838/p19/p585.002 and /afii9828/p24/p580.500 and the correspondinglog -likelihood228         Table 9.2 Goodness-of-Fit Tests Based on Asymptotic Likelihood Inference for the Tumor-Free Time Data from Rats F e dw i t hS a t ur a t e dD i e ti nT a b l e3 . 4 /p63 Estimated Parameters Model A /p64 B/p65 C/p66 LL X/p42pValue BIC AIC Exponential 5.054 — — /p5735.359 11.922 /p67 /p580.001 — — Exponential 5.054 — — /p5735.359 19.762 /p68 /p580.001 /p5737.060 /p5737.359 Weibull 5.002 0.500 — /p5729.398 7.840 /p680.005 /p5732.800 /p5733.398 Lognormal 4.765 0.561 — /p5726.641 2.326 /p680.127 /p5730.042 /p5730.641 Extended generalized 4.495 0.527 /p571.088 /p5725.478 — — /p5730.580 /p5731.478 gamma Log-logistic 4.739 0.332 — /p5726.867 — — /p5730.268 /p5730.867 /p63LL, log-likelihood; X/p42,l i k e l i h o o dr a t i os t a t i s t i c ; pvalue,P(X/p17/p57X/p42). /p64A/p58/p57log/afii9838for the exponential and the extended generalized gamma, /p58/p57 (1//afii9828)log/afii9838for the Weibull, /p58/afii9839for the lognormal, and /p58/p57 (1//afii9828)log/afii9825for the log-logistic distribution. /p65B/p581//afii9828for the Weibull and log-logistic, /p58/afii9846for the lognormal, and /p581/(/afii9825/afii9828/p15/p13/p20)for the extended generalized gamma distribution. /p66C/p581//afii9828for the extended generalized gamma distribution. /p67Relative to the Weibull. /p68Relativew to the extended generalized gamma. 229 is/p5729.398. To test the null hypothesis that the underlyingdistribution is an exponential distribution versus the alternative hypothesis that the underly-ingdistribution is Weibull (or extended generalized gamma ), the likelihood ratio test statistic X/p42/p582(35.359 /p5729.398 )/p5811.922 [or 2 (35.359 /p5725.478 )/p58 19.762]. The probability of observingsuch a chi-square value is /p580.001; therefore, the exponential distribution is rejected and the Weibull or thegeneralized gamma is preferred. However, the Weibull distribution is also rejected at the 0.001 level relative to the extended generalized gamma distribution (X/p42/p587.840,p/p580.001). This implies that the extended generalized gamma distribution may be better.However, the extended generalized gamma distribution is not significantlybetter than the lognormal distribution (X/p42/p582.326,p/p580.127 ). Thus, among these distributions, the lognormal and extended generalized gamma distribu-tions are our choices. Because of its simplicity, we may select the lognormaldistribution as the choice for this set of data. Example 9.3 Table 9.3 contains a set of remission times from 137 cancer patients. These remission times are a subset of the data from a bladder cancerstudy and are used here only for illustrative purposes. The results of goodnessof fit tests based on asymptotic likelihood inferences are shown in Table 9.4.From this table, we see that the exponential distribution is not rejected relativeto the Weibull distribution based on the statistic defined in (9.2.1 )(X/p42/p580.638, p/p580.425 ). The hypothesis that the underlyingdistribution is exponential versus the alternative hypothesis that the distribution is the extended general-ized gamma is rejected (X/p42/p586.772,p/p580.034 ). Furthermore, the Weibull and lognormal distributions are also rejected in favor of the extended generalizedgamma (X/p42/p586.135, 8.120,p/p580.013 and 0.004, respectively ). This implies that the exponential distribution may not be an appropriate distribution since theWeibull distribution (its nestingdistribution )is rejected. Therefore, we may accept the extended generalized gamma as our final choice of distribution forthe data. 9.3 SELECTION OF A DISTRIBUTION USING BIC OR AIC PROCEDURES The test procedures discussed in Section 9.2 require knowledge of the distribu- tion family to which the distribution of interest belongs. In this section weintroduce a simpler selection procedure called the Baysian information criterion (BIC; Schwarz, 1978 ). This criterion is based on the log-likelihood l(b/p19), the number of parameters in the distribution (p), and the total number of observations (n). For each candidate distribution, compute r/p58l(b/p19)/p57p 2 logn (9.3.1)230         Table 9.3 Remission Times (Months) of 137 Cancer Patients tt t t 4.50 32.15 3.88 13.80 19.13 4.87 3.02 /p59 5.85 14.24 5.71 19.36 /p59 7.09 7.87 7.59 20.28 5.32 5.49 3.02 46.12 4.33 /p59 2.02 4.51 5.17 2.839.22 1.05 0.20 8.373.82 9.47 36.66 14.77 26.31 79.05 10.06 8.53 4.65/p59 2.02 4.98 11.98 2.62 4.26 5.06 1.76 0.90 11.25 16.62 4.40 21.73 10.34 12.07 34.26 0.87/p59 10.66 6.97 2.07 0.51 12.03 0.08 17.123.36 2.64 1.40 12.63 43.01 14.76 2.75 7.66 0.81 1.19 7.32 4.183.36 8.66 1.26 13.291.46 14.83 6.76 23.63 24.80 /p59 5.62 8.60 /p59 3.25 10.86 /p59 18.10 7.62 7.63 17.14 25.74 3.52 2.87 15.96 17.36 9.74 3.31 7.28 1.35 0.40 2.264.33 9.02 5.41 2.69 22.69 6.94 2.54 11.79 2.46 7.26 2.69 5.34 3.48 4.70 /p59 8.26 6.93 4.23 3.70 0.50 10.756.54 3.64 5.32 13.118.65 3.57 5.09 7.395.41 11.64 2.092.23 6.25 7.93 4.34 25.82 12.02 whereb/p19denotes the MLE of all the parameters in the distribution. The candidate distribution with the largest rvalue is the distribution that fits the data the best. It has been shown that for some distribution familiesand under mild assumptions, for sufficiently large n, the distribution select- ed by the BIC procedure approaches the true underlyingdistribution, ifit exists.      231 Table 9.4 Goodness-of-Fit Tests Based on Asymptotic Likelihood Inference for Data in Table 9.3 /p63 Estimated Parameters Model A /p64 B/p65 C/p66 LL X/p42pValue BIC AIC Exponential 2.303 — — /p57198.234 0.638 /p670.425 — — Exponential 2.303 — — /p57198.234 6.772 /p680.034 /p57200.694 /p57200.234 Weibull 2.322 0.949 — /p57197.915 6.134 /p680.013 /p57202.835 /p57201.915 Lognormal 1.821 1.090 — /p57198.908 8.120 /p680.004 /p57203.828 /p57202.908 Extended generalized 2.087 0.999 0.520 /p57194.848 — — /p57202.228 /p57200.848 gamma Log-logistic 1.866 0.591 — /p57195.344 — — /p57200.264 /p57199.344 /p63LL, log-likelihood; X/p42,l i k e l i h o o dr a t i os t a t i s t i c ; pvalue,P(X/p17/p57X/p42). /p64/p11/p65/p11/p66See the footnotes in Table 9.2. /p67Relative to the Weibull. /p68Relative to the extended generalized gamma. 232 In general,thelargerthenumberofparameters pinadistribution,thelarger thelog-likelihood l(b/p19)in(9.3.1 ).Thus thefirst term representsthe gain by using a distribution with more parameters. But the larger the p, the larger the second term in (9.3.1 )is, which represents a penalty by havingmore parameters in the distribution. Therefore, the BIC provides a balance between the gain and thepenalty. Another widely used criterion is called an information criterion (AIC; Akaike, 1969 ), in which ris defined as r/p58l(b/p19)/p572p (9.3.2) Example 9.4 The values of the BIC and AIC for the various distributions considered in Examples 9.2 and 9.3 are listed in the last two columns in Tables9.2 and 9.4. Based on Table 9.2, the lognormal distribution would be selectedby either the BIC or AIC procedure, which is consistent with the resultsobtained in Example 9.2. The results in Table 9.4 show that the log-logisticdistribution, rather than the extended generalized gamma distribution, shouldbe selected based on either the BIC or AIC procedure. 9.4 TESTS FOR A SPECIFIC DISTRIBUTION WITH KNOWN PARAMETERS In this section we introduce the likelihood ratio statistic for testingif the survival data observed follow a given distribution with known parameters. Weuse the same notations as in Section 9.2. In addition to the exponential,Weibull,lognormal,gamma,generalizedgamma distributions,wealso considerthe log-logistic distribution. Let l/p42/p42(/afii9825,/afii9828) andl/p42/p42(/afii9825/p24,/afii9828/p24) denote its log-likelihood function and the log-likelihood with (/afii9825/p24,/afii9828/p24), the MLE of (/afii9825,/afii9828). 1.Testing the hypothesis that the underlying distribution is exponential with known parameter /afii9838/p15. The null hypothesis is H/p15:the underlyingdistribution is the exponential distribution with /afii9838/p58/afii9838/p15 The likelihood ratio test statistic based on (9.1.10 )is X/p42/p582[l/p35(/afii9838/p19)/p57l/p35(/afii9838/p15)] (9 .4.1) X/p42has an asymptotic chi-square distribution with 1 degree of freedom under H/p15.H/p15is rejected if X/p42/p57X/p17/p16/p11/p63, where /afii9825is the significance level. Similarly, the Wald test statistic and the score statistic can be derived by following (9.1.11 ) and (9.1.12 ). This is left to the reader as exercises. Example 9.5 Consider the followingsurvival times in weeks of 10 mice with a given tumor: 1, 3, 5, 8, 10 /p59, 15, 18, 19, 22, 25 /p59. We test the following        233 null hypothesis: H/p15:the underlyingdistribution of the observed data is exponential with /afii9838/p580.06 In this case, n/p5810,r/p588,/afii9814/p80/p71/p14/p16t/p71/p5891, and /afii9814/p76/p71/p14/p80/p62/p16t/p62/p71/p5835. The MLE of /afii9838 based on (7.2.16 )is /afii9838/p19/p588 91/p5935/p580.0635 andl/p35(/afii9838/p19)/p588(log0.0635 )/p570.0635 (91)/p570.0635 (35)/p58/p5730.055. Under H/p15, l/p35(/afii9838/p15)/p588(log0.06 )/p570.06(91)/p570.06(35)/p58/p5730.067. Thus, following (9.4.1 ), X/p42/p582[/p5730.055/p57(/p5730.067)] /p580.024.X/p17/p16/p11/p15/p13/p15/p20/p583.84; therefore, we cannot reject the null hypothesis that the data are from the exponential distributionwith/afii9838/p580.06. 2.Testing the hypothesis that the underlying distribution is Weibull with known parameters /afii9838/p15and/afii9828/p15. The null hypothesis is H/p15: the underlyingdistribution is Weibull with known parameters /afii9838/p58/afii9838/p15and/afii9828/p58/afii9828/p15 Based on (9.1.10 ), the likelihood ratio test statistic is X/p42/p582[l/p53(/afii9838/p19,/afii9828/p24)/p57l/p53(/afii9838/p15,/afii9828/p15)] (9 .4.2) UnderH/p15,X/p42has an asymptotic chi-square distribution with 2 degrees of freedom. H/p15is rejected if X/p42/p57X/p17/p17/p11/p63where /afii9825is the significance level. 3.Testing the hypothesis that the underlying distribution is lognormal with known parameters /afii9839/p15and/afii9846/p17/p15. Similar to the procedures above, the likelihood ratio test statistic is X/p42/p582[l/p42/p44(/afii9839/p24,/afii9846/p24/p17)/p57l/p42/p44(/afii9839/p15,/afii9846/p17/p15)] (9 .4.3) UnderH/p15,X/p42has an asymptotic chi-square distribution with 2 degrees of freedom. 4.Testing the hypothesis that the underlying distribution is standard gamma with known parameters /afii9838/p15and/afii9828/p15. The likelihood ratio test statistic is X/p42/p582[l/p37(/afii9838/p19,/afii9828/p24)/p57l/p37(/afii9838/p15,/afii9828/p15)] (9 .4.4) Under H/p15,X/p42is asymptotically chi-square distributed with 2 degrees of freedom.234         5.Testing the hypothesis that the underlying distribution is generalized gamma with known parameters /afii9825/p15,/afii9838/p15,and/afii9828/p15. The likelihood ratio test statistic is X/p42/p582[l/p37/p37(/afii9825/p24,/afii9838/p19,/afii9828/p24)/p57l/p37/p37(/afii9825/p15,/afii9838/p15,/afii9828/p15)] (9 .4.5) Under H/p15,X/p42is asymptotically chi-square distributed with 3 degrees of freedom. 6.Testing the hypothesis that the underlying distribution is log-logisticwith known parameters /afii9825/p15and/afii9828/p15. The likelihood ratio test statistic is X/p42/p582[l/p42/p42(/afii9825/p24,/afii9828/p24)/p57l/p42/p42(/afii9825/p15,/afii9828/p15)] (9 .4.6) UnderH/p15,X/p42has an asymptotic chi-square distribution with 2 degrees of freedom. Note that the respective Wald and score statistics can be constructed for these tests by following (9.1.11 )and (9.1.12 ). These are left to the reader as exercises. As noted in Section 9.3, the log-likelihood and estimated covariancematrix and parameters in (9.1.10 )—(9.1.12 )for each of the distributions dis- cussed in Sections 7.2 to 7.6 can be obtained from SAS or BMDP. The otherterms in these test statistics can also be obtained by usingSAS or BMDP. Thefollowingexample illustrates the use of SAS and BMDP. Example 9.6 To use the likelihood ratio statistic (9.4.1 )to test the null hypothesis in Example 9.5, H/p15:/afii9838/p58/afii9838/p15/p580.06, we need to calculate the log-likelihood l/p35(/afii9838/p19) andl/p35(/afii9838/p15).l/p35(/afii9838/p19) can be obtained by applyingeither the SASorBMDPcodes inExample7.5.We nowshowhowto useSAS or BMDPto calculate l/p35(/afii9838/p15). Suppose that the survival data of the 10 mice in Example 9.5 are saved in the file ‘‘C: /p33EXAMPLE.DAT’’. If SAS is used, we specify that the distribution is exponential by using D/p58EXPONENTIAL in the ‘‘model’’ statement and lettingINTERCEPT /p582.813 [ /p58/p57log/afii9838/p15/p58 /p57log(0.06)]. If BMDP is used, we specify the distribution by lettingAC- CEL /p58EXPONENTIAL and CONSTANT /p582.813. The followingSAS or BMDP codes can be used to obtain the l/p35(/afii9838/p15)and the terms needed for the Wald and score statistics in (9.1.11 )and (9.1.12 ). SAS code: data w1; infile ‘c: /p33example.dat’ missover; input t cens; run; proc lifereg; model t*cens (0)/p58/maxit /p580 covb itprint d /p58exponential intercept /p582.813; run;        235 BMDP code: /input file /p58‘c:/p33example.dat’ . variables /p582. format /p58free. /print level /p58brief. cova. iterations. /variable names /p58t, cens. /form time /p58t. status /p58cens. response /p581. /regress iteration /p580. accel /p58exponential. constant /p582.813. /end Similarly, to obtain the log-likelihood ratio statistic, the Wald and the score statistics in (9.1.10 )—(9.1.12 )for testingnull hypotheses about the parameters of other distributions, we can follow the same procedure but change the D /p58 and ACCEL /p58statements to reflect the distribution under the null hypothesis. We also need to provide values for the input variables INTERCEPT andSCALE, for Weibull, lognormal, and log-logistic distributions, if SAS is used.For the extended generalized gamma distribution, we need to provide a valuefor SHAPE1. BMDP does not have a procedure for the gamma distribution.For the Weibull, lognormal, and log-logistic distributions, we need to providevalues for CONSTANT and SCALE. All of these input variables are based onthe distribution and their relationship to the parameters under the nullhypothesis (see notes at the end of each of the SAS or BMDP codes in Section 7.2 to 7.6 ). 9.5 HOLLANDER AND PROSCHAN’S TEST FOR APPROPRIATENESS OF A GIVEN DISTRIBUTION WITH KNOWNPARAMETERS Another test for the appropriateness of a parametric distribution with known parameters was proposed by Hollander and Proschan (1979 ). Let 0/p58t/p7/p15/p8/p58t/p7/p16/p8/p58t/p7/p17/p8/p58/p37/p58t/p7/p76/p8be a set of distinct ordered survival times and some of the t/p7/p71/p8’s may be censored. If censored observations are tied with uncensoredobservations,treat the censoredobservationsof tieas beingg reaterthan the uncensored of the tie. Let S(t) be the underlyingsurvivorship function andS/p15(t) the survivorship function of the specific distribution. The null hypothesis is H/p15:S(t)/p58S/p15(t)236         Usingthe Kaplan —Meier product-limit method, S(t) is estimated as S/p19(t)/p58/p7/p73/p92/p16/p147 /p72/p14/p16/p1n/p57j n/p57j/p591/p2/p66/p7/p72/p8t/p7/p73/p92/p16/p8/p58t/p45t/p7/p73/p8,k/p581,...,n 0 t/p57t/p7/p76/p8(9.5.1) where /afii9829/p7/p72/p8/p581i ft/p7/p72/p8is uncensored and /afii9829/p7/p72/p8/p580i ft/p7/p72/p8is censored. Hollander and Proschan’s test statistic for the null hypothesis that the data are from adistribution with survivorship function S(t)i s C/p58 /p26 /p0/p12/p12 /p21/p14/p2/p4/p14/p19/p15/p18/p4/p3/p15/p1/p19/p4/p18/p22/p0/p20/p9/p15/p14/p19S/p15(t/p7/p71/p8)f/p19(t/p7/p71/p8)( 9 .5.2) where f/p19(t/p7/p71/p8) is the jump of the Kaplan —Meier estimates at consecutive uncensored observation and at the largest observation, uncensored or not, f/p19(t/p7/p71/p8)/p581 n/p71/p92/p16/p147 /p72/p14/p16/p1n/p57j/p591 n/p57j/p2/p16/p92/p66/p7/p72/p8. (9.5.3 ) Under the null hypothesis, C*/p58/p40n(C/p570.5) /afii9846/p24(9.5.4) follows approximatelythe standardnormal distribution, where /afii9846/p24is an estimate of the standard deviation of Cand /afii9846/p24/p17/p581 16/p76/p26 /p71/p14/p16n n/p57i/p591[S/p19/p15(t/p7/p71/p92/p16/p8)/p57S/p19/p15(t/p7/p71/p8)] (9 .5.5) To test H/p15:S/p58S/p15versusH/p16:S/p57S/p15, we reject H/p15ifC*/p58/p57Z/p63; to test H/p15versusH/p16:S/p58S/p15, we reject H/p15ifC*/p57Z/p63; and to test H/p15versusH/p16:S/p34S/p15, we reject H/p15ifC*/p58/p57Z/p63/p30/p17orC*/p57Z/p63/p30/p17, where Z/p63is the upper /afii9825percentile point of the standard normal distribution. The procedure for the calculation of C*can be summarized as follows. 1. Compute the Kaplan —Meier estimate S/p19(t) for each uncensored observa- tion. 2. Compute the jump of the Kaplan —Meier distribution at each t/p7/p71/p8uncen- sored, that is, f/p19(t/p7/p71/p8), which is the difference of F/p19(t)/p581/p57S/p19(t) at two consecutive uncensored observations. 3. Compute S/p15(t/p7/p71/p8) for each observation.   ’  237 4. Multiply S/p15(t/p7/p71/p8)b yf/p19(t/p7/p71/p8) and sum over all uncensored t/p7/p71/p8’s to obtain C. 5. Compute /afii9846/p24/p17accordingto (9.5.5 )and consequently, C*accordingto (9.5.4 ). Example 9.7 Consider the survival times in weeks of 10 mice in Example 9.5: 8, 5, 10 /p59, 1, 3, 18, 22, 15, 25 /p59, and 19. We wish to test that the survival time followsan exponentialdistributionwith /afii9838/p580.06.The null and alternative hypotheses are H/p15:S(t)/p58S/p15(t) H/p16:S(t)/p34S/p15(t) whereS/p15(t)/p58exp(/p570.06t). Followingtheprocedureoutlinedabove,wefirstarrangetheobservationsin ascendingorderand computetheKaplan —Meierestimatesasshownincolumn (d)of Table 9.5. The jumps are given in column (e). For example, the first jump is between S/p19(0) and S/p19(1) or 1 /p570.9/p580.1. Column (f)gives the survival function under the null hypothesis, for example, S/p15(3)/p58exp(/p570.06/p593)/p58 0.835. Following (9.5.2 ), column (g)gives the value of C/p580.4808. The last three columns are for calculation of the estimated variance of C. Thus, /afii9846/p24/p17/p581 16(1.3166 )/p580.0823 and C*/p58/p4010(0.4808 /p570.5) /p400.0823/p58/p570.2116 For/afii9825/p580.05,Z/p63/p30/p17/p581.96,C*does not fall in the rejection region. From Table B-1 we obtain that the pvalue correspondingto C*/p58/p570.2116 is approxi- mately 0.84. Therefore, we conclude that there is insufficient evidence to saythat the data are not from an exponential distribution with /afii9838/p580.06. Figure 9.1, which plots the Kaplan —Meier estimates and the hypothesized theoretical distribution S/p15(t)/p58exp(/p570.06t), demonstrates a close agreement between the two. The result is consistent with that obtained in Example 9.5, where thelikelihood ratio test is used. Bibliographical Remarks Readers with a background in mathematical statistics and an interest in mathematical details about asymptotic likelihood theory, likelihood ratio, Wald’s, and score statistics are referred to Cox (1961, 1962 a), Atkinson (1970 ), Hagar and Bain (1970 ), Cox and Hinkley (1974 ), and Kalbfleisch and Prentice238         Table 9.5 Calculation of Test Statistic C*for Data in Example 9.7 (a)( b)( c)( d)( e)( f)( g)( h)( i)( j) t/p7/p71/p8in/p57i n/p57i/p591S/p19(t)f/p19(t/p7/p71/p8)S/p15(t/p7/p71/p8) (e)/p59(f)S/p19/p15(t/p7/p71/p8)S/p19/p15(t/p7/p71/p92/p16/p8)/p57S/p19/p15(t/p7/p71/p8)n n/p57i/p591/p59(i) 1 1 0.900 0.900 0.100 0.941 0.0941 0.7841 0.2159 /p63 0.2159 3 2 0.889 0.800 0.100 0.835 0.0835 0.4861 0.2980 0.3311 5 3 0.875 0.700 0.100 0.741 0.0741 0.3015 0.1846 0.2308 8 4 0.857 0.600 0.100 0.619 0.0619 0.1468 0.1547 0.2210 10/p59 5 — — 0 0.549 0 0.0908 0.0560 0.0933 15 6 0.800 0.480 0.120 0.407 0.0560 0.0274 0.0634 0.126818 7 0.750 0.360 0.120 0.340 0.0408 0.0134 0.0140 0.035019 8 0.667 0.240 0.120 0.320 0.0384 0.0105 0.0029 0.0097 22 9 0.500 0.120 0.120 0.267 0.0320 0.0051 0.0054 0.0270 25/p5910 — — 0 0.223 0 0.0025 0.0026 0.0260——— ———0.4808 1.3166 /p63S/p19/p15(t/p7/p15/p8)/p58S/p19/p15(0)/p581. 239 Figure 9.1 Kaplan—Meier estimator S/p19(t) and the hypothesized survival function S/p15(t)/p58exp(/p570.06t). (1980 ). There have been many papers about the AIC and BIC criteria in the literature since the introduction of AIC by Akaike (1969 ). The asymptotic properties of the two criteria and their relationships with other criteria werediscussed by Akaike (1974 ), Parzen (1974 ), Schwarz (1978 ), Hannan (1979 ), Shibata (1980 ),Wang (1984,1989 ), Rissanen (1986 ), and Wei (1992 ). Interested readers are referred to these papers for details. When there are no censored observations, the chi-square goodness of fit test introduced by Karl Pearson in 1900 can be used to test any distributionalassumption. In addition, tests for the exponential and lognormal (Shapiro and Wilk, 1965a, b )are available. EXERCISES 9.1Derive the likelihood ratio and Wald test statistics following (9.1.2 )and (9.1.3 )for the followingnull hypothesis: H/p15: The underlyingdistribution is exponential versus the alternatives (a)H/p16: The underlyingdistribution is g amma (b)H/p16: The underlying distribution is generalized gamma240         9.2Derive the Wald test statisticsfollowing (9.1.3 )for the followingnull and alternative hypotheses H/p15: The underlyingdistribution is Weibull H/p16: The underlying distribution is generalized gamma 9.3Derive the respective Wald test statistics by following (9.1.3 )for the null hypothesis that the distribution is the standard gamma versus thealternative hypothesis that the distribution is the generalized gamma. 9.4Derive the Wald and score statistics for testingthe null hypotheses in Section 9.4 by following (9.1.11 )and (9.1.12 ). 9.5Consider the survival time of 28 cancer patients in Exercise 8.5. (a)Obtain the log-likelihoods for the exponential, Weibull, lognormal, and generalized gamma distributions. Perform the likelihood ratiotest and select the best distribution amongthese four distributions. (b)Use the BIC and AIC procedures to select the best distribution amongthefourdistributionsinpart (a)plusthelog-logisticdistribu- tion. (c)Comparetheresults obtainedin parts (a)and(b)and those obtained in Exercise 8.5. 9.6Consider the survival time of 31 patients with advanced melanoma in Exercise 8.6.(a)Select the best distribution usingthe likelihood ratio, Wald, and score statistics among the exponential, Weibull, gamma, lognormal,and generalized gamma distributions. (b)Use the BIC and AIC procedures to do the same as in part (a)with the addition of the log-logistic distribution. (c)Comparetheresults obtainedin parts (a)and(b)and those obtained in Exercise 8.6. (d)Compare the MLE of the parameters with those estimates obtained by usingthe g raphical methods in Exercise 8.6. 9.7Do the same as in Exercise 9.5 for the data in Exercise Table 3.1. 9.8Consider the followingsurvival time in weeks of 10 mice with injection of tumor cells: 5, 16, 18 /p59, 20, 22 /p59,2 4/p59, 25, 30 /p59, 35, 40 /p59. Do the data follow the exponential distribution with /afii9838/p580.02? (a)Use the likelihood ratio test. (b)Use Hollander and Proschan’s test statistic. (c)Plot the Kaplan —Meier estimator of S(t) and the hypothesized distribution. 241 9.9Consider the followingsurvival time in months of 25 patients with cancer of the prostate. Test the hypothesis that the survival time ofprostate cancer patients follows the exponential distribution with/afii9838/p580.01: 2, 19, 19, 25, 30, 35, 40, 45, 45, 48, 60, 62, 69, 89, 90, 110, 145, 160, 9 /p59,1 0/p59,2 0/p59,4 0/p59,5 0/p59, 110 /p59, 130 /p59. 9.10The Gompertz distribution belongs to the Gompertz —Makeham dis- tribution family and the Gompertz —Makeham distribution (Makeham, 1860 )has the followinghazard function: h(t)/p58/afii9825/p59exp(/afii9838/p59/afii9828t)/afii9825/p460 Construct alikelihood ratiostatistic and Wald’s statisticto test theappro- priateness of a Gompertz distribution.242         CHAPTER 10 Parametric Methods for Comparing Two Survival Distributions In Chapter 5 we discussed several nonparametric tests for comparing twosurvival distributions. If the distributions follow a known model, parametrictests are more powerful than nonparametric tests, but their computation ismoretedious.Inthischapterwe firstdiscussthelikelihoodratio testingeneralfor comparing two survival distributions in Section 10.1. Readers who are notfamiliarwith linear algebramay skip this section without loss of continuity. InSections 10.2 to 10.4 we present either the likelihood ratio test or other testsfor the comparison of two survival patterns that follow the exponential,Weibull, and gamma distributions. 10.1 LIKELIHOOD RATIO TEST FOR COMPARING TWO SURVIVAL DISTRIBUTIONS Letx/p16,...,x/p76/p129, and y/p16,...,y/p76/p130betheobservedexactorcensoredsurvivaltimes ofn/p16and n/p17subjectsfromtwogroups.Assumethatthesurvivaltimesfromthe two groups follow the same distribution with different parameters. We use thegeneral notation b/p58(b/p16,b/p17,...,b/p78)to denote the set of parameters of the distribution, p/p461.Let l/p71(b/p71),i/p581,2,denotethelog-likelihoodfunctionforthe observed survival times from each group, where b/p71/p58(b/p71/p16,...,b/p71/p73, b/p71/p73/p62/p16,...,b/p71/p78)/p58(b/p71/p16,b/p71/p17), andb/p71/p16/p58(b/p71/p16,...,b/p71/p73)andb/p71/p17/p58(b/p71/p73/p62/p16,...,b/p71/p78)are twosubsetsofthe pparameters, i/p581,2.Thenthejointlog-likelihoodfunction for the two groups is l(b/p16,b/p17)/p58l/p16(b/p16)/p59l/p17(b/p17). Letb/p19/p71denote the MLE of b/p71, andb/p19/p71/p17(b/p15)denote the MLE of b/p71/p17givenb/p71/p16/p58b/p15, whereb/p15is known. For example, if the survival time of the two groups follows the Weibull distribution with a scale parameter /afii9838and a shape parameter /afii9828. Then p/p582, b/p16/p58(/afii9838/p16,/afii9828/p16), andb/p17/p58(/afii9838/p17,/afii9828/p17), where /afii9838/p16,/afii9828/p16and/afii9838/p17,/afii9828/p17are the respective parameters in the two Weibull distributions. Let l/p16(/afii9838/p16,/afii9828/p16) and l/p17(/afii9838/p17,/afii9828/p17) denote the log-likelihood functions of the observed survival times from two groups; 243 then the joint log-likelihood function for the two groups is l(/afii9838/p16,/afii9828/p16,/afii9838/p17,/afii9828/p17)/p58l/p16(/afii9838/p16,/afii9828/p16)/p59l/p17(/afii9838/p17,/afii9828/p17) In this case, b/p71/p16may be a singleton /afii9838/p71, and similarly, b/p71/p17may be /afii9828/p71. The followingtests are widelyused in comparingtwosurvival distributions. Case 1. All parameters are unknown. When b/p71,i/p581, 2, are unknown, we test the hypothesis H/p15:b/p16/p58b/p17/p58b (10.1.1 ) that is, that the two groups have the same survival distribution with equal but unknown parameters b. The log-likelihood ratio test statistic X/p42/p58/p572[l(b/p19,b/p19)/p57l(b/p19/p16,b/p19/p17)] /p582[l/p16(b/p19/p16)/p59l/p17(b/p19/p17)/p57l(b/p19,b/p19)] (10.1.2 ) has an asymptotic chi-square distribution with pdegrees of freedom. For a given significance level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p78/p11/p63(10.1.3) or equivalently, if P(/afii9851/p17/p78/p57X/p42)/p58/afii9825 (10.1.4 ) where /afii9851/p17/p78denotes the chi-square random variable with pdegrees of freedom, and/afii9851/p17/p78/p11/p63is its 100 (1/p57/afii9825)percentile points, P(/afii9851/p17/p78/p57/afii9851/p17/p78/p11/p63)/p58/afii9825. In the case of comparing two Weibull distributions, it reduces to H/p15:/afii9838/p16/p58/afii9838/p17/p58/afii9838and/afii9828/p16/p58/afii9828/p17/p58/afii9828 where /afii9838and/afii9828are unknown, X/p42/p582[l/p16(/afii9838/p19/p16,/afii9828/p24/p16)/p59l/p17(/afii9838/p19/p17,/afii9828/p24/p17)/p57l(/afii9838/p19,/afii9828/p24,/afii9838/p19,/afii9828/p24)] and H/p15is rejected if X/p42/p57/afii9851/p17/p17/p11/p63, or equivalently, if P(/afii9851/p17/p17/p57X/p42)/p58/afii9825. Case 2. A subset of the parameters of the two survival distributions are known and equal, say, b/p16/p16/p58b/p17/p16/p58b/p15, where the values of b/p15are known. The null hypothesis is the equality of the remaining parameters, or H/p15:b/p16/p17/p58b/p17/p17/p58b/p28/p17(10.1.5 )244        whereb/p28/p17is unknown. The log-likelihood ratio statistic is X/p42/p582[l/p16(b/p15,b/p19/p16/p17(b/p15))/p59l/p17(b/p15,b/p19/p17/p17(b/p15))/p57l((b/p15,b/p19/p28/p17(b/p15)),(b/p15,b/p19/p28/p17(b/p15))] (10.1.6 ) whereb/p19/p16/p17(b/p15)is the MLE of b/p16/p17givenb/p16/p16/p58b/p15, and so are the others. X/p42has an asymptotic chi-square distribution with degrees of freedom equal to thenumber of parameters in b/p16/p17(orb/p17/p17). In the case of comparing two Weibull distributions, we may assume that /afii9838/p16/p58/afii9838/p17/p58/afii9838/p15(or/afii9828/p16/p58/afii9828/p17/p58/afii9828/p15), where the value of /afii9838/p15(or/afii9828/p15)is known,and test the null hypothesis H/p15:/afii9828/p16/p58/afii9828/p17/p58/afii9828(or/afii9838/p16/p58/afii9838/p17/p58/afii9838) Then X/p42/p582[l/p16(/afii9838/p15,/afii9828/p24/p16(/afii9838/p15))/p59l/p17(/afii9838/p15,/afii9828/p24/p17(/afii9838/p15))/p57l(/afii9838/p15,/afii9828/p24(/afii9838/p15),/afii9838/p15,/afii9828/p24(/afii9838/p15))] (orX/p42/p582[l/p16(/afii9838/p19/p16(/afii9828/p15),/afii9828/p15)/p59l/p17(/afii9838/p19/p17(/afii9828/p15),/afii9828/p15)/p57l(/afii9838/p19(/afii9828/p15),/afii9828/p15,/afii9838/p19(/afii9828/p15),/afii9828/p15)] and H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63, equivalently, if P(/afii9851/p17/p16/p57X/p42)/p58/afii9825. Case 3. A subset of the parameters of the two survival distributions are equal but unknown, say, if b/p16/p16/p58b/p17/p16/p58b/p28/p16and the values of b/p28/p16are unknown. The null hypothesis is the equality of the remaining parameters, or H/p15:b/p16/p17/p58b/p17/p17/p58b/p28/p17(10.1.7 ) whereb/p28/p17is unknown and needs to be estimated.In addition, b/p28/p16also needs to be estimated. The log-likelihood ratio statistic, X/p42/p582[l((b/p19/p28/p16,b/p19/p16/p17),(b/p19/p28/p16,b/p19/p17/p17))/p57l(b/p19,b/p19)] (10.1.8 ) has an asymptoticchi-squaredistributionwith degrees of freedomequal to the number of parameters in b/p16/p17(orb/p17/p17). For the case of comparing two Weibull distributions, the derivation of X/p42in(10.1.8 )is left to the reader as an exercise. Case 4. A subset of the parameters of the two survival distributions are known but not equal, say, if b/p16/p16/p58b/p16/p15,b/p17/p16/p58b/p17/p15, andb/p16/p15andb/p17/p15are known        245 butb/p16/p15/p34b/p17/p15. The null hypothesis is the equality of the remaining parameters, or H/p15:b/p16/p17/p58b/p17/p17/p58b/p28/p17(10.1.9 ) The log-likelihood ratio statistic X/p42/p582[l/p16(b/p16/p15,b/p19/p16/p17(b/p16/p15))/p59l/p17(b/p17/p15,b/p19/p17/p17(b/p17/p15)) /p57l((b/p16/p15,b/p19/p28/p17(b/p16/p15,b/p17/p15)),(b/p17/p15,b/p19/p28/p17(b/p16/p15,b/p17/p15)))] (10.1.10 ) has an asymptoticchi-squaredistributionwith degrees of freedomequal to the number of parameters in b/p16/p17orb/p17/p17. For the case of comparing two Weibull distributions, the derivation of X/p42in(10.1.10 )is left to the reader as an exercise. 10.2 COMPARISON OF TWO EXPONENTIAL DISTRIBUTIONS Suppose that two survival distributions follow the exponential model with parameters /afii9838/p16and/afii9838/p17, respectively. Two tests can compare the distributions: thelikelihoodratiotestand an F-testsuggestedby Cox (1953 ).Thesetwotests can test the hypothesis that the two exponential distributions are equalwhether or not the samples include censored observations. 10.2.1 Likelihood Ratio Test Suppose that there are n/p16and n/p17individuals in groups 1 and 2, respectively, x/p16,...,x/p80/p129uncensored and x/p62/p80/p129/p62/p16,..., x/p62/p76/p129censored in group 1, and y/p16,...,y/p80/p130uncensored and y/p62/p80/p130/p62/p16,..., y/p62/p76/p130censored in group 2. Thus, in group 1, there are r/p16uncensored and n/p16/p57r/p16censored observations. In group 2, there are r/p17uncensored and n/p17/p57r/p17censored observations. If it is known that the survival times of the two groups follow the exponential distribution with densityfunction f/p71(t)/p58/afii9838/p71e/p92/p72 /p71/p82,i/p581,2,testingtheequalityoftwoexponentialdistribu- tionsisequivalenttotestingthehypothesis H/p15:/afii9838/p16/p58/afii9838/p17.Thisisbecausethetwo exponential distributions are characterized by the two parameters /afii9838/p16and/afii9838/p17. Thus, the null hypothesis is H/p15:/afii9838/p16/p58/afii9838/p17/p58/afii9838and the alternative hypothesis is H/p16:/afii9838/p16/p34/afii9838/p17. According to (10.1.2 ), the test statistic for the likelihood ratio test is X/p42/p58/p572logL(/afii9838/p19,/afii9838/p19) L(/afii9838/p19/p16,/afii9838/p19/p17)(10.2.1)246        wherethedenominatoristhelikelihoodfunctionforthetwogroupscombined, L(/afii9838/p19/p16,/afii9838/p19/p17)/p58/afii9838/p19/p80/p16/p16/afii9838/p19/p80/p17/p17exp/p3/p57/afii9838/p19/p16/p1/p80/p16/p26 /p71/p14/p16x/p71/p59/p76/p16/p26 /p71/p14/p80/p16/p62/p16x/p62/p71/p2 /p57/afii9838/p19/p17/p1/p80/p17/p26 /p71/p14/p16y/p71/p59/p76/p17/p26 /p71/p14/p80/p17/p62/p16y/p62/p71/p2/p4(10.2.2) and/afii9838/p19/p16and/afii9838/p19/p17are the MLE of /afii9838/p16and/afii9838/p17, respectively, from groups 1 and 2. From Section 7.2, /afii9838/p19/p16/p58r/p16/p26/p80/p16/p71/p14/p16x/p71/p59/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71/afii9838/p19/p17/p58r/p17/p26/p80/p17/p71/p14/p16y/p71/p59/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71(10.2.3) The numerator in (10.2.1 )is the likelihood function for the combined sample under the null hypothesis, that is, /afii9838/p16/p58/afii9838/p17/p58/afii9838, L(/afii9838/p19,/afii9838/p19)/p58/afii9838/p19/p80/p16/p62/p80/p17exp/p3/p57/afii9838/p19/p1/p80/p16/p26 /p71/p14/p16x/p71/p59/p76/p16/p26 /p71/p14/p80/p16/p62/p16x/p62/p71/p59/p80/p17/p26 /p71/p14/p16y/p71/p59/p76/p17/p26 /p71/p14/p80/p17/p62/p16/p2/p4 (10.2.4) where /afii9838/p19is the MLE of /afii9838obtained from the combined sample, /afii9838/p19/p58r/p16/p59r/p17/p26/p80/p16/p71/p14/p16x/p71/p59/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71/p59/p26/p80/p17/p71/p14/p16y/p71/p59/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71(10.2.5) From Section 10.1, X/p42has an approximate chi-square distribution with 1 degree of freedom for samples of at least 25 ( n/p16/p59n/p17/p4625) under the null hypothesis. For a given significance level /afii9825,H/p15is rejected if X/p42/p57/afii9851/p17/p16/p11/p63,o r equivalently, if P(/afii9851/p17/p16/p57X/p42)/p58/afii9825. The test procedure can be summarized as follows: 1. Compute /afii9838/p19/p16and/afii9838/p19/p17following (10.2.3 ). 2. Compute L(/afii9838/p19/p16,/afii9838/p19/p17)in(10.2.2 )usingthegivendataand /afii9838/p19/p16and/afii9838/p19/p17obtained in step 1. 3. Compute /afii9838/p19following (10.2.5 ). 4. Compute L(/afii9838/p19,/afii9838/p19)i n(10.2.4 ). 5. Compute X/p42in(10.2.1 ).I fX/p42/p57/afii9851/p17/p16/p11/p63(Table B-2 ), reject H/p15and conclude that the two exponential survival distributions are not equal. Otherwise,the data do not provide enough evidence to reject the null hypothesis. If there are no censored observations in the data, (10.2.1 )—(10.2.5 )are also applicable simply be letting n/p16/p58r/p16,n/p17/p58r/p17and omitting the terms involving     247 x/p62/p71and y/p62/p71. The likelihood ratio test is primarily for two-sided tests and is difficulttoapplyto a one-sidedtest.It isapproximateandshouldbe used withcaution when the sample size is small. The power of the test, similar to that ofother likelihood tests, is not high. That is, if the likelihood ratio test is usedregularly, one is more likely not to reject the null hypothesis when the twosurvival distributions are not equal. Example 10.1 Consider the remission data of the two treatment groups given in Example 5.1. The remission times in months are as follows: CMF: 23, 16 /p59,1 8/p59,2 0/p59,2 4/p59 Control: 15, 18, 19, 19, 20 Assume that the two distributions are exponential with parameters /afii9838/p16and/afii9838/p17, respectively. Using the likelihood ratio test, we test the following null hypoth-esis H/p15:/afii9838/p16/p58/afii9838/p17/p58/afii9838(the two treatments are equally effective ) against H/p16:/afii9838/p16/p34/afii9838/p17(the two treatments are not equally effective ) Following the above, we proceed as follows: 1. Compute /afii9838/p19/p16and/afii9838/p19/p17in(10.2.3 ): In this case, n/p16/p58n/p17/p585,r/p16/p581,r/p17/p585, /p26/p80 /p16/p71/p14/p16x/p71/p5823,/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71/p5878,/p26/p80/p17/p71/p14/p16y/p71/p5891,/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71/p580, /afii9838/p19/p16/p581 23/p5978/p581 101/p580.0099 /afii9838/p19/p17/p585 91/p580.0549 2. Compute L(/afii9838/p19/p16,/afii9838/p19/p17)in(10.2.2 ): L(/afii9838/p19/p16,/afii9838/p19/p17)/p58(0.0099)(0 .0549) /p20exp[/p570.0099 (101)/p570.0549 (91)] /p581.2290 (10)/p92/p16/p16 3. Compute /afii9838/p19in(10.2.5 ): /afii9838/p19/p581/p595 23/p5978/p5991/p586 192/p580.0313248        4. Compute L(/afii9838/p19/p16,/afii9838/p19/p17)i n(10.2.4 ): L(/afii9838/p19/p16,/afii9838/p19/p17)/p58(0.0313) /p21exp[/p570.0313 (192)]/p582.3085 (10)/p92/p16/p17 5. Compute X/p42in(10.2.1 ): X/p42/p58/p572log2.3085 (10)/p92/p16/p17 1.2290 (10)/p92/p16/p16/p58/p572log (0.1878 )/p583.344 From Table B-2 we obtain /afii9851/p17/p16/p11/p15/p13/p15/p20/p583.84. Thus we cannot reject H/p15at the 0.05 level. Recall that in Chapter 5 the null hypothesis was rejected at the 0.05level by using the four nonparametric tests. 10.2.2 Cox’s F-Test for Exponential Distributions If the times to failure can be assumed to follow the exponential distribution in both treatment groups, an F-test suggested by Cox (1953 )can be used to test for treatment differences whether or not censored observations are present.Suppose that we wish to test the hypothesis H/p15:/afii9838/p16/p58/afii9838/p17against either the one-sided alternative H/p16:/afii9838/p16/p58/afii9838/p17(orH/p17:/afii9838/p16/p57/afii9838/p17)or the two-sided alternative H/p18:/afii9838/p16/p34/afii9838/p17. An efficient test is to take t/p16/p16/t/p16/p17as having an F-distribution with (2r/p16,2r/p17)degrees of freedom, where t/p16/p16/p58/p26/p80 /p16/p71/p14/p16x/p71/p59/p26/p76/p16/p71/p14/p80/p16/p62/p16x/p62/p71r/p16(10.2.6) t/p16/p17/p58/p26/p80/p17/p71/p14/p16y/p71/p59/p26/p76/p17/p71/p14/p80/p17/p62/p16y/p62/p71r/p17(10.2.7) The test procedures are (1)forH/p16, reject H/p15ift/p16/p16/t/p16/p17/p57F/p17/p80/p16/p11/p17/p80/p17/p11/p63;(2)forH/p17, reject H/p15ift/p16/p16/t/p16/p17/p58F/p17/p80/p16/p11/p17/p80/p17/p11/p16/p92/p63; and (3)forH/p18, reject H/p15ift/p16/p16/t/p16/p17/p57F/p17/p80/p16/p11/p17/p80/p17/p11/p63/p30/p17or t/p16/p16/t/p16/p17/p58F/p17/p80/p16/p11/p17/p80/p17/p11/p16/p92/p63/p30/p17, where /afii9825is the significance level and F/p17/p80/p16/p11/p17/p80/p17/p11/p63is the upper 100/afii9825percentage point of the F-distribution with (2 r/p16,2r/p17) degrees of freedom. Similarly, the hypothesis that /afii9838/p16//afii9838/p17/p58kcan be tested by referring kt/p16/p16/t/p16/p17to the table of the F-distribution. When there are no censored observations, that is, n/p16/p58r/p16,n/p17/p58r/p17, the second terms of the numerators in (10.2.6 )and (10.2.7 )are zero. Then the test statistic t/p16/p16/t/p16/p17has an F-distribution with (2 n/p16,2n/p17)degrees of freedom. Confidence intervals for the ratio /afii9838/p16//afii9838/p17can be obtained from the fact that /afii9838/p16t/p16/p16//afii9838/p17t/p16/p17has the F-distribution with (2 n/p16,2n/p17) degrees of freedom. It follows thata100 (1/p57/afii9825)%confidenceintervalfortheratiooftwohazardrates /afii9838/p16//afii9838/p17is t/p16/p17t/p16/p16F/p17/p80/p16/p11/p17/p80/p17/p11/p16/p92/p63/p30/p17/p58/afii9838/p16/afii9838/p17/p58t/p16/p17t/p16/p16F/p17/p80/p16/p11/p17/p80/p17/p11/p63/p30/p17(10.2.8)     249 Example 10.2 Thirty-six patients with glioblastoma multiforme were divided into two groups; the experimental group contained 21 patients whohad surgery and chemotherapy, and the control group contained 15 patientswhohadsurgeryonly.Thesurvivaltimesinweeksareavailableaboutoneyearafter the start of the study (Burdette and Gerhan, 1970 ): Experimental: 1, 2, 2, 2, 6, 8, 8, 9, 13, 16, 17, 29, 34, 2 /p59,9/p59,1 3/p59,2 2/p59, 25/p59,3 6/p59,4 3/p59,4 5/p59 Control: 0, 2, 5, 7, 12, 42, 46, 54, 7 /p59,1 1/p59,1 9/p59,2 2/p59,3 0/p59,3 5/p59, 39/p59 The hypotheses are H/p15:/afii9838/p16/p58/afii9838/p17(no difference in survival between experimental and control groups ) H/p16:/afii9838/p16/p58/afii9838/p17(difference in survival favoring experimental group ) In this case, n/p16/p5821, n/p17/p5815, r/p16/p5813, r/p17/p588,/afii9814x/p71/p58147,/afii9814x/p62/p71/p58195, /afii9814y/p71/p58168, and /afii9814y/p62/p71/p58163. Hence t/p16/p16/p58147/p59195 13 /p5826.308 t/p16/p17/p58168/p59163 8/p5841.375 and t/p16/p16/t/p16/p17/p580.636 with (26,16 )degrees of freedom. For /afii9825/p580.05, F/p17/p21/p11/p16/p21/p11/p15/p13/p15/p20is approximately 2.23; hence the hypothesis H/p15is not rejected and the data do not provide enough evidence that the survival time is longer in the experimen-tal group. A 95%confidence interval for the ratio /afii9838/p16//afii9838/p17is 41.375 26.308(0.419 )/p58/afii9838/p16/afii9838/p17/p5841.375 26.308(2.625 ) or(0.659, 4.128 ). The estimate of /afii9838/p19/p16//afii9838/p19/p17according to (10.2.3 )is 1.58. Hence the data show that the death rate per week of the experimental group is close tothat of the control group. In Example 10.2, the guarantee time in both groups is zero. In the case where a group has a nonzero guarantee time, it can be subtracted from everyobservation in the group and the test then applied. Monte Carlo studies(Gehan and Thomas, 1969; Lee et al., 1975 )show that when samples are from exponential distributions, with or without censoring, the F-test is the most powerful test among the parametric or nonparametric tests discussed in thischapter and Chapter 5.250        10.3 COMPARISON OF TWO WEIBULL DISTRIBUTIONS It is well known that if the survival time Thas a Weibull distribution with shape parameter /afii9828, then T/p65has an exponential distribution. Thus, if /afii9828/p16and/afii9828/p17for the two groups are known, the most powerful Cox’s F-test described in Section 10.2 can be applied to the transformed observations. However, inpractice, /afii9828/p16and/afii9828/p17are probably unknown, and so are the scale parameters, /afii9838/p16and/afii9838/p17. In this case, the likelihood ratio tests described in Section 10.1 can be applied to test whether the observed survival times from the two groups havethe same Weibull distribution. To test the equality of two Weibull distribu-tions,itsufficestotest /afii9828/p16/p58/afii9828/p17and/afii9838/p16/p58/afii9838/p17.Ifthehypothesis /afii9828/p16/p58/afii9828/p17isrejected, we need not test the hypothesis /afii9838/p16/p58/afii9838/p17. If the hypothesis /afii9828/p16/p58/afii9828/p17is not rejected, we do need to test /afii9838/p16/p58/afii9838/p17. In the following, we introduce an additional two-sample test proposed by Thoman and Bain (1969 )for uncen- sored samples. Assume that independent random samples of equal size ( n/p16/p58n/p17/p58n) are obtained from Weibull distributions f/p16(t) and f/p17(t), where f/p71(t)/p58/afii9838/p71/afii9828/p71(/afii9838/p71t)/p65 /p71/p92/p16exp[/p57(/afii9838/p71t)/p65/p71] i/p581, 2 (10 .3.1) To test /afii9828/p16/p58/afii9828/p17, we use the property of the maximum likelihood estimator /afii9828/p24 (Thoman et al., 1969; Thoman and Bain, 1969 ). To test the null hypothesis H/p15:/afii9828/p16/p58/afii9828/p17against H/p16:/afii9828/p16/p57/afii9828/p17, we use the fact that (/afii9828/p24/p16//afii9828/p16)/(/afii9828/p24/p17//afii9828/p17)/p58/afii9828/p24/p16//afii9828/p24/p17under H/p15. Thepercentagepoints of /afii9828/p24/p16//afii9828/p24/p17are givenin TableB-12.We compute theMLEof /afii9828/p16and/afii9828/p17,thatis, /afii9828/p24/p16and/afii9828/p24/p17,andcompare /afii9828/p24/p16//afii9828/p17withthepercentage points for a given /afii9825in Table B-12. Reject H/p15at/afii9825level if /afii9828/p24/p16//afii9828/p24/p17/p57l/p63. For example,if n/p16/p58n/p17/p58n/p5810,acomputed /afii9828/p24/p16//afii9828/p24/p17/p571.897wouldleadtorejection ofH/p15at a significance level of 0.05. For /afii9825/p460.50, percentage points l/p63can be calculated by using the relationship l/p63/p581/l/p16/p92/p63. The procedure described above can be generalized to test H/p15:/afii9828/p16/p58k/afii9828/p17against H/p16:/afii9828/p16/p58k/p30/afii9828/p17. For the case when k/p58k/p30, the rejection region becomes /afii9828/p24/p16//afii9828/p24/p17/p57kl/p63, where /afii9825is the significance level. If the hypothesis H/p15:/afii9828/p16/p58/afii9828/p17is rejected, the two Weibull distributions are not the same. However, if the hypothesis is not rejected, we need to test theequality of the two scale parameters /afii9838/p16and/afii9838/p17. A test of H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p58/afii9838/p17suggested by Thoman and Bain (1969 )rejects H/p15if G/p58/p16/p17(/afii9828/p24/p16/p59/afii9828/p24/p17)(log/afii9838/p19/p17/p57log/afii9838/p19/p16)/p57z/p63(10.3.2) where z/p63issuchthat P(G/p58z/p63/p34H/p15)/p581/p57/afii9825and/afii9828/p24/p16,/afii9828/p24/p17,/afii9838/p19/p16,and/afii9838/p19/p17aretheMLEs of/afii9828/p16,/afii9828/p17,/afii9838/p16, and /afii9838/p17respectively. The percentage points z/p63are given in Table B-13.Forexample,ifthe commonsamplesizeis 10,thehypothesis H/p15:/afii9838/p16/p58/afii9838/p17is rejected if G/p460.918 at significance level 0.05. A test of H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p57/afii9838/p17can be constructed in a similar fashion. The critical points z/p63can be obtained from Table B-13 by using the fact that z/p63/p58/p57 z/p16/p92/p63.     251 Table 10.1 Survival Times of Patients in Two Treatment Groups Treatment 1 Treatment 2 5, 10, 17, 32, 32, 33, 34, 36, 43, 44, 44, 48, 48, 61, 64, 65, 65, 66, 67, 68, 82, 85,90, 92, 92, 102, 103, 106, 107, 114, 114,116,117,124,139,142,143,151,158,19520.9, 32.2, 33.2, 39.4, 40.0, 46.8, 57.3, 58.0,59.7,61.1,61.4,54.3,66.0,66.3,67.4,68.5, 69.9, 72.4, 73.0, 73.2, 88.7, 89.3,91.6, 93.1 94.2, 97.7, 101.6, 101.9, 107.6, 108.0, 109.7, 110.8, 114.1, 117.5, 119.2, 120.3, 133.0, 133.8, 163.3, 165.1 Example 10.3 illustrates the test procedures. The data are adapted and modified from Harter and Moore (1965 ). Forty observations are generated from a Weibull distribution with /afii9838/p16/p580.01 and /afii9828/p16/p582 and another 40 from a Weibull distribution with /afii9838/p17/p580.01 and /afii9828/p17/p583. The resulting data are shown in Table 10.1. For illustrative purposes, we consider the two samples as twotreatment groups. Example 10.3 Consider the survival times of the patients in the two treatmentgroupsinTable10.1.Thenullhypothesisisthatthetwopopulationshave the same shape parameter; that is, H/p15:/afii9828/p16/p58/afii9828/p17against H/p16:/afii9828/p16/p58/afii9828/p17. The MLEs are /afii9828/p24/p16/p581.945, /afii9828/p24/p17/p582.715, and hence /afii9828/p24/p17//afii9828/p24/p16/p581.396, which is significant at the 0.05 level ( l/p15/p13/p15/p20/p581.342 for n/p5840) but not significant at the 0.02 level ( l/p15/p13/p15/p17/p581.453 for n/p5840). If we choose /afii9825/p580.05 and reject H/p15, the decision is correct. An error of not rejecting H/p15would be committedif an /afii9825of 0.02 or 0.01 is chosen. This is because the two shape parameters are very close(/afii9828/p16/p582,/afii9828/p17/p583). To illustrate the procedure of testing the equality of the scale parameters, letusassumethatthehypothesis H/p15:/afii9828/p16/p58/afii9828/p17isnotrejected.Totest H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p57/afii9838/p17,weneedtheMLEsof /afii9838/p16and/afii9838/p17. HarterandMoore obtain /afii9838/p19/p16/p580.010776 and /afii9838/p19/p17/p580.010471. From (10.3.2 )we obtain G/p58/p16/p17 (1.945 /p592.715 )(4.559 /p574.530 )/p580.068 From Table B-13, the critical region for n/p5840 is G/p570.404. Hence we do not reject H/p15. This decision is correct since /afii9838/p16/p58/afii9838/p17/p580.01. Note that the MLEs of /afii9828/p16,/afii9828/p17,/afii9838/p16, and /afii9838/p17are very close to their real values. 10.4 COMPARISON OF TWO GAMMA DISTRIBUTIONS Suppose that x/p16,...,x/p76, and y/p16,...,y/p76are the survival times of patients receivingtwo different treatmentsand that they followthe gamma distribution with the density function given in (6.4.1 ). Let /afii9838/p16and/afii9828/p16be the parameters of252        Table 10.2 Survival Times of 40 Patients Receiving Two Different Treatments Treatment 1 (x) Treatment 2( y) 17, 28, 49, 98, 119 26, 34, 47, 59, 101, 133, 145, 146, 158, 160, 112, 114, 136, 154, 154,174, 211, 220, 231, 252, 161, 186, 197, 226, 226,256, 267, 322, 323, 327 243, 253, 269, 308, 465thexpopulation and /afii9838/p17and/afii9828/p17be those of the ypopulation. The likelihood ratio tests introduced in Section 10.1 can be used to test whether the survivaltimes observed from the xpopulation and the ypopulation have different gamma distributions. The estimation of the parameters is quite complicatedbut can be obtained using commercially available computer programs. In thefollowing we introduce an F-test for testing the null hypothesis H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p34/afii9838/p17, under the assumptions that the x/p71’s and y/p71’s are exact (uncensored )survival times, and that /afii9828/p16and/afii9828/p17are known (usually assumed equal ). Letx/p21and y/p21be the sample mean survival times of the two groups. The test is based on the fact that x/p21/y/p21has the F-distribution with 2 n/afii9828/p16and 2 n/afii9828/p17degrees of freedom (Rao, 1952 ). Thus the test procedure is to reject H/p15at the /afii9825level if x/p21/y/p21exceeds F/p17/p76/p65 /p16/p11/p17/p76/p65/p17/p11/p63/p30/p17, the 100 (/afii9825/2)percentage point of the F-distribution with (2n/afii9828/p16,2n/afii9828/p17)degrees of freedom. Since the F-table gives percentage points for integer degrees of freedom only, interpolations (linear or bilinear )are necessary when either 2 n/afii9828/p16or 2n/afii9828/p17is not an integer. The following example illustrates the test procedure. The data are adapted andmodifiedfromHarterandMoore (1965 ).Theysimulated40survivaltimes from the gamma distribution with parameters /afii9828/p16/p58/afii9828/p17/p58/afii9828/p582,/afii9838/p580.01. The 40 individuals are divided randomly into two groups for illustrative purposes. Example 10.4 Consider the survival time of the two treatment groups in Table 10.2. The two populations follow the gamma distributions with acommon shape parameter /afii9828/p582. To test the hypothesis H/p15:/afii9838/p16/p58/afii9838/p17against H/p16:/afii9838/p16/p34/afii9838/p17, we compute x/p21/p58181.80,y/p21/p58173.55, and x/p21/y/p21/p581.048. Under the nullhypothesis, x/p21/y/p21hasthe F-distributionwith (80,80 )degreesoffreedom.Use /afii9825/p580.05, F/p23/p15/p11/p23/p15/p11/p15/p13/p15/p17/p20/p601.45. Hence, we do not reject H/p15at the 0.05 level of significance. The result is what we would expect since the two samples aresimulated from the same overall sample of 40 with /afii9838/p580.01. To test the equality of two lognormal distributions, we use the fact that the logarithmic transformation of the observed survival times follows the normaldistributions, and thus we can use the standard tests based on the normaldistribution. In general, for other distributions, such as log-logistic and thegeneralized gamma, the log-likelihood ratio statistics defined in Section 10.1     253 can be applied to test whether the survival times observed from two groups have the same distribution. Readers can follow Example 10.2.1 in Section 10.2and use the respective likelihood functions derived in Chapter 7 to constructthe needed tests. Bibliographical Remarks In addition to the papers cited in this chapter, readers are referred to Mann et al.(1974 ), Gross and Clark (1975 ), Lawless (1982 ), and Nelson (1982 ). EXERCISES 10.1Derive the likelihood ratio tests in (10.1.8 )and (10.1.10 )for testing the equality of two Weibull distributions. 10.2Derive the likelihood ratio test in (10.1.2 )for testing the equality of two log-logistic distributions with unknown parameters. 10.3Consider the remission data of the leukemia patients in Example 3.3. Assume that the remission times of the two treatment groups follow theexponentialdistribution.Testthehypothesisthatthetwotreatmentsareequally effective using: (a)The likelihood ratio test (b)Cox’s F-test Obtain a 95%confidence interval for the ratio of the two hazard rates. 10.4For the same data in Exercise 10.3, test the hypothesis that /afii9838/p17/p585/afii9838/p16. 10.5Suppose that the survival time of two groups of lung cancer patients follows the Weibull distribution. A sample of 30 patients (15 from each group )was studied. Maximun likelihood estimates obtained from the two groups are, respectively, /afii9828/p24/p16/p583,/afii9838/p19/p16/p581.2 and /afii9828/p24/p17/p582,/afii9838/p19/p17/p580.5. Test the hypothesis that the two groups are from the same Weibull distribu-tion. 10.6Divide the lifetimes of 100 strips (delete the last one )of aluminum coupon in Table 6.4 randomly into two equal groups. This can be donebyassigningtheobservationsalternatelytothetwogroups.Assumethatthe two groups follow a gamma distribution with shape parameter/afii9828/p5812. Test the hypothesis that the two scale parameters are equal. 10.7Twelve brain tumor patients are randomized to receive radiation ther- apy or radiation therapy plus chemotherapy (BCNU )in a one-year clinical trial. The following survival times in weeks are recorded:254        1. Radiation /p59BCNU: 24, 30, 42, 15 /p59,4 0/p59,4 2/p59 2. Radiation: 10, 26, 28, 30, 41, 12 /p59 Assumingthatthesurvivaltimefollowsthe exponentialdistribution,use Cox’s F-test for exponential distributions to test the null hypothesis H/p15:/afii9838/p16/p58/afii9838/p17versus the alternative H/p16:/afii9838/p16/p58/afii9838/p17. 10.8Use one of the nonparametric tests discussed in Chapter 5 to test the equalityof survivaldistributionsofthe experimentaland controlgroupsin Example 10.2. Compare your result with that obtained in Example10.2. 255 CHAPTER11 ParametricMethodsforRegression ModelFittingandIdentificationofPrognosticFactors Prognosis,thepredictionofthefutureofanindividualpatientwithrespecttoduration,course,andoutcomeofadiseaseplaysanimportantroleinmedicalpractice.Beforeaphysiciancanmakeaprognosisanddecideonthetreatment,amedicalhistoryaswellaspathologic,clinical,andlaboratorydataareoftenneeded.Therefore,manymedicalchartscontainalargenumberofpatient (or individual )characteristics (also calledconcomitantvariables ,independentvari- ables,covariates,prognosticfactors ,o rriskfactors ), and it is often difficult to sort outwhich onesare mostcloselyrelatedto prognosis.The physiciancanusuallydecidewhichcharacteristicsare irrelevant,buta statisticalanalysisisusuallyneededtoprepareacompactsummaryofthedatathatcanrevealtheirrelationship. One way to achieve this purpose is to search for a theoreticalmodel (or distribution ), that fits the observed data and identify the most important factors. These models, usually regression models, extend themethods discussed in previous chapters to include covariates. In this chapterwe focus on parametric regression models (i.e., we assume that the survival time follows a theoretical distribution ). If an appropriate model can be assumed, the probability of surviving a given time when covariates areincorporatedcanbeestimated. InSection11.1wediscussbrieflypossibletypesofresponseandprognostic variables and things that can be done in a preliminary screening before aformal regression analysis. This section applies to methods discussed in thenext four chapters. In Section 11.2 we introduce the general structure of acommonly used parametric regression model, the accelerated failure time(AFT )model.Sections11.3to11.7coverseveralspecialcasesofAFTmodels. Fittingthesemodelsofteninvolvescomplicatedandtediouscomputationsandrequirescomputersoftware.Fortunately,mostoftheproceduresareavailableinsoftwarepackagessuchasSASandBMDP.TheSASandBMDPcodethat 256 canbeusedtofitthemodelsaregivenattheendoftheexamples.Readersmay findthesecodeshelpful.Section11.8introducestwoothermodels.InSection11.9wediscussthemodelselectionmethodsandgoodnessoffittests. 11.1 PRELIMINARY EXAMINATION OF DATA Informationconcerningpossibleprognosticfactorscanbeobtainedeitherfrom clinical studiesdesignedmainly to identifythem, sometimescalled prognostic studies,orfromongoingclinicaltrialsthatcomparetreatmentsasasubsidiary aspect. The dependent variable (also called the response variable ), or the outcome of prediction, may be dichotomous, polychotomous, or continuous.Examples of dichotomous dependent variables are response or nonresponse,life or death, and presence or absence of a given disease. Polychotomousdependentvariablesincludedifferentgradesofsymptoms (e.g.,noevidenceof disease,minorsymptom,majorsymptom )and scoresof psychiatricreactions (e.g.,feelingwell,tolerable,depressed,orverydepressed ).Continuousdepend- ent variables may be length of survival from start of treatment or length ofremission,bothmeasuredonanumericalscalebyacontinuousrangeofvalues.Of these dependent variables, response to a given treatment (yes or no ), developmentof a specific disease (yes or no ), lengthof remission,and length ofsurvivalare particularlycommonin practice.In this chapterwe focusourattention on continuous dependent variables such as survival time and re-missionduration.Dichotomousandmultiple-responsedependentvariablesarediscussedinChapter14. Aprognostic variable (or independent variable )or factor may be either numerical or nonnumerical. Numerical prognostic variables may be discrete,suchasthenumberofpreviousstrokesornumberoflymphnodemetastases,or continuous, such as age or blood pressure. Continuous variables can bemadediscretebygroupingpatientsintosubcategories (e.g.,fouragesubgroups: /p5820, 20—39, 40—59, and /p4660). Nonnumerical prognostic variables may be unordered (e.g.,raceordiagnosis )orordered (e.g.,severityofdiseasemaybe primary,local,ormetastatic ).Theycanalsobedichotomous (e.g.,alivereither is or is not enlarged ).Usually, the collectionof prognosticvariables includes someofeachtype. Before a statistical calculation is done, the data have to be examined carefully. If some of the variables are significantly correlated, one of thecorrelated variables is likely to be a predictor as good as all of them.Correlation coefficients between variables can be computed to detect signifi-cantly correlated variables. In deleting any highly correlated variables, infor-mationfromotherstudieshastobeincorporated.Ifotherstudiesshowthatagivenvariablehasprognosticvalue,itshouldberetained. Inthenexteight sectionswediscussmultivariateorregressiontechniques, which are useful in identifying prognostic factors. The regression techniquesinvolveafunctionoftheindependentvariablesorpossibleprognosticfactors.    257 Thevariablesmust bequantitative,with particularnumericalvaluesforeach patient. This raises no problem when the prognostic variables are naturallyquantitative (e.g.,age )andcanbeusedintheequationdirectly.However,ifa particular prognostic variable is qualitative (e.g., a histological classification into one of three cell types A, B, or C ), something needs to be done. This situation can be covered by the use of two dummy variables, e.g., x/p16, taking thevalue1forcelltypeAand0otherwise,and x/p17,takingthevalue1forcell typeBand0otherwise.Clearly,ifthereareonlytwocategories (e.g.,sex ),only onedummyvariableisneeded: x/p16is1foramale,0forafemale.Also,abetter descriptionofthedatamightbeobtainedby usingtransformedvaluesof theprognosticvariables (e.g.,squaresorlogarithms )orbyincludingproductssuch asx/p16x/p17(representing an interaction between x/p16andx/p17). Transforming the dependent variable (e.g., taking the logarithm of a response time )can also improvethefit. Inpractice,thereareusuallyalargernumberofpossibleprognosticfactors associatedwiththeoutcomes.Onewaytoreducethenumberoffactorsbeforeamultivariateanalysisisattemptedistoexaminetherelationshipbetweeneachindividual factor and the dependent variable (e.g., survival time ). From the univariate analysis, factors that have little or no effect on the dependentvariablecanbeexcludedfromthemultivariateanalysis.However,itwouldbedesirabletoincludefactorsthathavebeenreportedtohaveprognosticvaluesbyotherinvestigatorsandfactorsthatareconsideredimportantfrombiomedi-calviewpoints.Itisoftenusefultoconsidermodelselectionmethodstochoosethosesignificantfactorsamongallpossiblefactorsanddetermineanadequatemodel with as few variables as possible. Very often, a variable of significantprognosticvalueinonestudyisunimportantinanother.Therefore,confirma-tioninalaterstudyisveryimportantinidentifyingprognosticfactors. Another frequent problem in regression analysis is missing data. Three distinctionsaboutmissingdatacanbemade: (1)dependentversusindependent variables, (2)manyversusfewmissingdata,and (3)randomversusnonrandom loss of data. If the value of the dependent variable (e.g., survival time )is unknown,thereislittletodobutdropthatindividualfromanalysisandreducethe sample size. The problem of missing data is of different magnitudedependingonhowlargeaproportionofdata,eitherforthedependentvariableor for the independent variables, is missing. This problem is obviously lesscritical if 1%of data for one independent variable is missing than if 40%ofdataforseveralindependentvariablesismissing.Whenasubstantialpropor-tion of subjects has missing data for a variable, we may simply opt to dropthemandperformtheanalysisontheremainderofthesample.Itisdifficulttospecify‘‘howlarge’’and‘‘howsmall,’’butdropping10or15casesoutofseveralhundredwould raise no serious practicalobjection.However,if missing dataoccurinalargeproportionofpersonsandthesamplesizeisnotcomfortablylarge,aquestionofrandomnessmayberaised.Ifpeoplewithmissingdatadonotshowsignificantdifferencesinthedependentvariable,the problemis notserious.Ifthedataarenotmissingrandomly,resultsobtainedfromdropping258         subjects will be misleading. Thus, dropping cases is not always an adequate solutiontothemissingdataproblem. If the independentvariableis measuredon anominal or categoricalscale, analternativemethodistotreatindividualsinagroupwithmissinginforma-tion as another group. For quantitatively measured variables (e.g., age ), the meanofthevaluesavailablecanbeusedforamissingvalue.Thisprinciplecanalso be applied to nominal data. It does not mean that the mean is a goodestimateforthemissingvalue,butitdoesprovideconvenienceforanalysis. A more detailed discussion on missing data can be found in Cohen and Cohen (1975,Chap.7 ),LittleandRubin (1987 ),Efron (1994 ),Crawfordetal. (1995 ),Heitjan (1997 ),andSchafer (1999 ). 11.2 GENERAL STRUCTURE OF PARAMETRIC REGRESSION MODELS AND THEIR ASYMPTOTIC LIKELIHOOD INFERENCE When covariates are considered, we assume that the survival time, or a function of it, has an explicit relationship with the covariates. Furthermore,whena parametricmodelisconsidered,weassumethat thesurvivaltime (or afunctionofit )followsagiventheoreticaldistribution (ormodel )andhasan explicit relationship with the covariates. As an example, let us consider theWeibulldistributioninSection6.2.Let x/p58(x/p16,...,x/p78)denotethepcovariates considered. If the parameter /afii9838in the Weibull distribution is related to xas follows: /afii9838/p58e /p57(a/p15/p59/afii9814/p78/p71/p14/p16a/p71x/p71)/p58exp[/p57(a/p15/p59a/p30x)] wherea/p58(a/p16,...,a/p78)denotethecoefficientsfor x,thenthehazardfunctionof theWeibulldistributionin (6.2.4 )canbeextendedtoincludethecovariatesas follows: h(t,x)/p58/afii9838/p65/afii9828t/p65/p92/p16 /p58 /afii9828t/p65/p92/p16e/p57(a/p15/p59/afii9814/p78/p71/p14/p16a/p71x/p71)/afii9828/p58/afii9828t/p65/p92/p16exp[/p57(a/p15/p59a/p30x)/afii9828] (11.2.1 ) Thesurvivorshipfunctionin (6.2.3 )becomes S(t,x)/p58(e/p92/p82/p65)exp(/p57/afii9828(a/p15/p59a/p30x))(11.2.2 ) or log[/p57logS(t,x)]/p58/p57/afii9828(a/p15/p59a/p30x)/p59/afii9828logt(11.2.3) whichpresentsalinearrelationshipbetweenlog[ /p57logS(t,x)]andlogtandthe covariates. In Sections 11.2 to 11.7 we introduce a special model called the acceleratedfailuretimemodel. Analogous to conventional regression methods, survival time can also be analyzedbyusingthe acceleratedfailuretime (AFT )model.TheAFTmodel      259 forsurvivaltimeassumesthattherelationshipoflogarithmofsurvivaltime T andthecovariatesislinearandcanbewrittenas logT/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p59/afii9846/afii9830 (11.2.4 ) wherex/p72,j/p581,...,p, are the covariates, a/p72,j/p580, 1,...,pthe coefficients, /afii9846 (/afii9846/p570)is an unknown scale parameter, and /afii9830, the error term, is a random variablewithknownformsofdensityfunction g(/afii9830,d)andsurvivorshipfunction G(/afii9830,d)butunknownparameters d. Thismeansthat thesurvivalisdependent onboththecovariateandanunderlyingdistribution g. Considerasimplecasewherethereisonlyonecovariate xwithvalues0and 1.Then (11.2.4 )becomes logT/p58a/p15/p59a/p16x/p59/afii9846/afii9830 LetT/p15andT/p16denote the survival times for two individuals with x/p580 and x/p581, respectively. Then, T/p15/p58exp(a/p15/p59/afii9846/afii9830), andT/p16/p58exp(a/p15/p59a/p16/p59/afii9846/afii9830)/p58 T/p15exp(a/p16).Thus,T/p16/p57T/p15ifa/p16/p570andT/p16/p58T/p15ifa/p16/p580.Thismeansthatthe covariatexeither ‘‘accelerates’’ or ‘‘decelerates’’ the survival time or time to failure—thusthename acceleratedfailuretimemodels forthisfamilyofmodels. In the following we discuss the general form of the likelihood function of AFTmodels,theestimationproceduresoftheregressionparameters (a/p15,a,/afii9846, andd)in(11.2.4 )andtestsofsignificanceofthecovariatesonthesurvivaltime. The calculations of these procedures can be carried out using availablesoftwarepackagessuchasSASandBMDP.ReaderswhoarenotinterestedinthemathematicaldetailsmayskiptheremainingpartofthissectionandmoveontoSection11.3withoutlossofcontinuity. Lett/p16,t/p17,...,t/p76betheobservedsurvivaltimesfrom nindividuals,including exact, left-, right-, and interval-censored observations. Assume that the logsurvival time can be modeled by (11.2.4 )and let a/p30/p58(a/p16,a/p17,...,a/p78), and b/p30/p58(a/p30,d/p30,a/p15,/afii9846).Similarto (7.1.1 ),thelog-likelihoodfunctionintermsofthe densityfunction g(/afii9830) andsurvivorshipfunction G(/afii9830)o f/afii9830is l(b)/p58logL(b)/p58/p26log[g(/afii9830/p71)]/p59/p26log[G(/afii9830/p71)] /p26log[1 /p57G(/afii9830/p71)]/p59/p26log[G(/afii9834/p71)/p57G(/afii9830/p71)] (11.2.5 ) where /afii9830/p71/p58logt/p71/p57a/p15/p57/p26/p78/p72/p14/p16a/p72x/p72/p71/afii9846 (11.2.6) /afii9834/p71/p58log/afii9840/p71/p57a/p15/p57/p26/p78/p72/p14/p16a/p72x/p72/p71/afii9846(11.2.7)260         The first term in the log-likelihood function sums over uncensored observa- tions, the second term over right-censored observations, and the third termover left-censored observations, and the last term over interval-censoredobservationswith /afii9840/p71asthelowerendofacensoringinterval.Notethatthelast two summations in (11.2.5 )do not exist if there are no left- and interval- censoreddata. Alternatively,let /afii9839/p71/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71i/p581,2,...,n (11.2.8 ) Then (11.2.4 )becomes logT/p58/afii9839/p59/afii9846/afii9830 (11.2.9) The respective alternative log-likelihood function in terms of the density functionf(t,b)andsurvivorshipfunction S(t,b)ofTis l(b)/p58logL(b)/p58/p26log[f(t/p71,b)]/p59/p26log[S(t/p71,b)] /p59/p26log[1 /p57S(t/p71,b)]/p59/p26log[S(/afii9840/p71,b)/p57S(t/p71,b)] (11.2.10 ) wheref(t,b)canbederivedfrom (11.2.4 )throughthedensityfunction g(/afii9830)b y applyingthedensitytransformationrule f(t,b)/p58g((logt/p57/afii9839)//afii9846) /afii9846t (11.2.11) andS(t,b)isthecorrespondingsurvivorshipfunction.Thevector bin(11.2.10 ) and (11.2.11 )includes the regression coefficients and other parameters of the underlyingdistribution. Either (11.2.5 )or(11.2.10 )can be used to derive the maximum likelihood estimates (MLEs )of parameters in the model. For a given log-likelihood functionl(b),theMLE b/p19isasolutionofthefollowingsimultaneousequations: /p42(l(b)) /p42b/p71/p580 foralli (11.2.12) Usually, there is no closed solution for the MLE b/p19from (11.2.12 )and the Newton—RaphsoniterativeprocedureinSection7.1mustbeappliedtoobtain b/p19. By replacing the parameters bwith its MLE b/p19inS(t/p71,b), we have an estimated survivorship function S(t,b/p19), which takes into consideration the covariates. All of the hypothesis tests and the ways to construct confidence intervals showninSection7.1canbeappliedhere.Inaddition,wecanusethefollowingteststotestlinearrelationshipsamongtheregressioncoefficients a/p16,a/p17,...,a/p78.      261 To test a linear relationship among x/p16,...,x/p78is equivalent to testing the nullhypothesisthatthereisalinearrelationshipamong a/p16,a/p17,...,a/p78.H/p15can bewritteningeneralas H/p15:La/p58c (11.2.13 ) whereLisamatrixorvectorofconstantsforthelinearhypothesisand cisa knowncolumnvectorofconstants.ThefollowingWald’sstatisticscanbeused: X/p53/p58(La/p24/p57c)/p30[LV/p19/p63(a/p24)L/p30]/p92/p16(La/p24/p57c)( 11.2.14 ) whereV/p19/p63(a/p24)isthesubmatrixofthecovariancematrix V/p19(b/p19)correspondingto a. UndertheH/p15andsomemildassumptions, X/p53hasanasymptoticchi-square distributionwith /afii9840degrees of freedom,where /afii9840is the rank of L. For a given significancelevel /afii9825,H/p15isrejectedifX/p53/p57/afii9851/p17/p74/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p74,1/p57/afii9825/2. Forexample,if p/p583andwewishtotestif x/p16andx/p17haveequaleffectson thesurvivaltime,thenullhypothesisis H/p15:a/p16/p58a/p17(ora/p16/p57a/p17/p580).Itiseasy toseethatforthishypothesis,thecorresponding L/p58(1,/p571,0)and c/p580since La/p58(1,/p571, 0)(a/p16,a/p17,a/p18)/p30/p58a/p16/p57a/p17 Letthe (i,j)elementofV/p19/p63(a/p24)be/afii9840/p71/p72;thentheX/p53definedin (11.2.14 )becomes X/p53/p58(a/p24/p16/p57a/p24/p17)/p3(1,/p571, 0)/p1/afii9840/p16/p16/afii9840/p16/p17/afii9840/p16/p18 /afii9840/p17/p16/afii9840/p17/p17/afii9840/p17/p18 /afii9840/p18/p16/afii9840/p18/p17/afii9840/p18/p18/p2/p11 /p571 0/p2/p4/p92/p16 (a/p24/p16/p57a/p24/p17) /p58(a/p24/p16/p57a/p24/p17)/p17 /afii9840/p16/p16/p59/afii9840/p17/p17/p572/afii9840/p16/p17 X/p53has an asymptotic chi-square distribution with 1 degree of freedom (the rankofLis1). Ingeneral,totestifanytwocovariateshavethesameeffectson T,thenull hypothesiscanbewrittenas H/p15:a/p71/p58a/p72(ora/p71/p57a/p72/p580) (11 .2.15) The corresponding L/p58(0,...,0,1,0,...,0, /p571, 0,...,0 )andc/p580, and the X/p53in(11.2.14 )becomes X/p53/p58(a/p24/p71/p57a/p24/p72)/p17 /afii9840/p71/p71/p59/afii9840/p72/p72/p572/afii9840/p71/p72(11.2.16)262         whichhasanasymptoticchi-squaredistributionwith1degreeoffreedom. H/p15isrejectedifX/p53/p57/afii9851/p17/p16/p11/p63/p30/p17orX/p53/p58/afii9851/p17/p16/p11/p16/p92/p63/p30/p17. To test that noneof the covariatesis relatedto the survivaltime,the null hypothesisis H/p15:a/p580 (11.2.17 ) The respective test statistics for this overall null hypothesis are shown in Section9.1.Forexample,thelog-likelihoodratiostatisticstherebecomes X/p42/p58/p572[l(0,d/p19(0),a/p24/p15(0),/afii9846/p24(0))/p57l(b/p19)] (11.2.18 ) which has an asymptotic chi-square distribution with pdegrees of freedom underH/p15, wherepis the number of covariates; d/p19(0),a/p24/p15(0), and /afii9846/p24(0)are the MLEof d,a/p15,and /afii9846givena/p580. 11.3 EXPONENTIAL REGRESSION MODEL Toincorporatecovariatesintotheexponentialdistribution,weuse (11.2.4 )for thelogsurvivaltimeandlet /afii9846/p581: logT/p71/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71/p59/afii9830/p71/p58/afii9839/p71/p59/afii9830/p71, (11.3.1 ) where /afii9839/p71/p58a/p15/p59/afii9814/p78/p72/p14/p16a/p72x/p72/p71,/afii9830/p71’sareindependentlyidenticallydistributed (i.i.d.) random variables with a double exponential or extreme value distributionwhichhasthefollowingdensityfunction g(/afii9830) andsurvivorshipfunction G(/afii9830): g(/afii9830)/p58exp[/afii9830/p57exp(/afii9830)] (11.3.2 ) G(/afii9830)/p58exp[/p57exp(/afii9830)] (11.3.3 ) This model is the exponential regression model. Thas the exponential distributionwiththefollowinghazard,density,andsurvivorshipfunctions. h(t,/afii9838/p71)/p58/afii9838/p71/p58exp/p3/p57/p1a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71/p2/p4/p58exp(/p57/afii9839/p71)(11.3.4 ) f(t,/afii9838/p71)/p58/afii9838/p71exp(/p57/afii9838/p71t) (11.3.5 ) S(t,/afii9838/p71)/p58exp(/p57/afii9838/p71t)( 11.3.6 ) where /afii9838/p71isgivenin (11.3.4 ).Thus,theexponentialregressionmodelassumesa linear relationship between the covariates and the logarithm of hazard. Let   263 h/p71(t,/afii9838/p71) andh/p72(t,/afii9838/p72)be thehazards ofindividuals iandj; the hazardratio of thesetwoindividualsis h/p71(t,/afii9838/p71) h/p72(t,/afii9838/p72)/p58/afii9838/p71/afii9838/p72/p58exp[/p57(/afii9839/p71/p57/afii9839/p72)]/p58exp/p3/p57/p78/p26 /p73/p14/p16a/p73(x/p73/p71/p57x/p73/p72)/p4(11.3.7) This ratio is dependent only on the differences of the covariates of the two individualsandthe coefficients.It doesnotdependonthe time t.InChapter 12we introduce aclass of modelscalled proportionalhazardmodels inwhich the hazard ratio of any two individualsis assumed to be a time-independentconstant. The exponential regression model is therefore a special case of theproportionalhazardmodels. The MLE of b/p58(a/p15,a/p16,...,a/p78)is a solution of (11.2.12 ), using (11.2.10 ), wheref(t,/afii9838) andS(t,/afii9838) aregivenin (11.3.5 )and (11.3.6 ).Computerprograms inSASorBMDPcanbeusedtocarryoutthecomputation. In the following we introduce a practical exponential regression model. Suppose that there are n/p58n/p16/p59n/p17/p59/p37/p59n/p73individuals in ktreatment groups.Lett/p71/p72bethesurvivaltimeand x/p16/p71/p72,x/p17/p71/p72,...,x/p78/p71/p72thecovariatesof the jthindividualinthe ithgroup,where pisthenumberofcovariatesconsidered, i/p581,...,k, andj/p581,...,n/p71. Define the survivorship function for the jth individualinthe ithgroupas S/p71/p72(t)/p58exp(/p57/afii9838/p71/p72t) (11 .3.8) where /afii9838/p71/p72/p58exp(/p57/afii9839/p71/p72)and /afii9839/p71/p72/p58/p57 /p1a/p15/p71/p59/p78/p26 /p74/p14/p16a/p74x/p74/p71/p72/p2(11.3.9) This model was proposed by Glasser (1967 )and was later investigated by Prentice (1973 )and Breslow (1974 ). The term exp (/p57a/p15/p71)represents the underlyinghazardofthe ithgroupwhencovariatesareignored.Itisclearthat /afii9839/p71/p72definedin (11.3.9 )isaspecialcaseof (11.3.4 )byaddinganewindexforthe treatment groups. To construct the likelihood function, we use the followingindicatorvariablestodistinguishcensoredobservationsfromtheuncensored: /afii9829/p71/p72/p58/p71i ft/p71/p72uncensored 0i ft/p71/p72censored According to (11.2.10 )and (11.3.8 ), the likelihood function for the data can thenbewrittenas L(/afii9838/p71/p72)/p58/p73/p26 /p71/p14/p16/p76/p71/p147 /p72/p14/p16(/afii9838/p71/p72)/p66/p71/p72exp(/p57/afii9838/p71/p72t/p71/p72)264         Substituting (11.3.9 )in the logarithm of the function above, we obtain the log-likelihoodfunctionof a/p15/p58(a/p15/p16,a/p15/p17,...,a/p15/p73) anda/p58(a/p16,a/p17,...,a/p78): l(a/p15,a)/p58/p73/p26 /p71/p14/p16/p76/p71/p26 /p72/p14/p16/p3/afii9829/p71/p72/p1a/p15/p71/p59/p78/p26 /p74/p14/p16a/p74x/p74/p71/p72/p2/p57t/p71/p72exp/p1a/p15/p71/p59/p78/p26 /p74/p14/p16a/p74x/p74/p71/p72/p2/p4 /p58/p73/p26 /p71/p14/p16/p3a/p15/p71r/p71/p59/p78/p26 /p74/p14/p16a/p74s/p71/p74/p57exp(a/p15/p71)/p76/p71/p26 /p72/p14/p16t/p71/p72exp/p1/p78/p26 /p74/p14/p16a/p74x/p74/p71/p72/p2/p4(11.3.10) where s/p71/p74/p58/p76/p71/p26 /p72/p14/p16/afii9829/p71/p72x/p74/p71/p72 isthesumofthe lthcovariatecorrespondingtotheuncensoredsurvivaltimes intheithgroupandr/p71isthenumberofuncensoredtimesinthatgroup. Maximumlikelihoodestimatesof a/p15/p71’sanda/p74’scanbeobtainedbysolving thefollowingk/p59pequationssimultaneously.Theseequationsareobtainedby takingthederivativeof l(a/p15,a)in(11.3.10 )withrespecttothe ka/p15/p71’sandpa/p74’s: r/p71/p57exp(a/p24/p15/p71)/p76/p71/p26 /p72/p14/p16t/p71/p72exp/p1/p78/p26 /p74/p14/p16a/p24/p74x/p74/p71/p72/p2/p580i/p581,...,k(11.3.11) /p73/p26 /p71/p14/p16/p3s/p71/p74/p57exp(a/p24/p71)/p76/p71/p26 /p72/p14/p16t/p71/p72x/p74/p71/p72exp/p1/p78/p26 /p74/p14/p16a/p24/p74x/p74/p71/p72/p2/p4/p580l/p581,...,p(11.3.12) ThiscanbedonebyusingtheNewton —RaphsoniterativeprocedureinSection 7.1.ThestatisticalinferencesfortheMLEandthemodelarethesameasthosestated in Section 7.1. Let a/p24/p15anda/p24be the MLE of a/p15andain(11.3.10 ), and a/p24/p15(0)be the MLE of a/p15givena/p580. According to (11.2.18 ), the difference betweenl(a/p24/p15,a/p24)andl(a/p24/p15(0),0)can be used to test the overall null hypothesis (11.2.17 )that none of the covariates is related to the survival time by considering X/p42/p58/p572(l(a/p24/p15(0),0)/p57l(a/p24/p15,a/p24)) ( 11.3.13 ) aschi-squaredistributedwith pdegreesoffreedom.A X/p42greaterthanthe100 /afii9825 percentage point of the chi-square distribution with pdegrees of freedom indicates significant covariates. Thus, fitting the model with subsets of thecovariatesx/p16,x/p17,...,x/p78allowsselectionofsignificantcovariatesofprognostic variables.Forexample,if p/p582,totestthesignificanceof x/p17afteradjustingfor x/p16,thatis,H/p15:a/p17/p580,wecompute X/p42/p58/p572[l(a/p24/p15(0),a/p24/p16(0),0) /p57l(a/p24/p15,a/p24/p16,a/p24/p17)]   265 Table 11.1 Summary Statistics for the Five Regimens Additive Therapy Geometric Median 6-MP MTX Numberof Numberin Mean /p63of Mean Remission Regimen Cycle Cycle Patients Remission WBC Age (yr)Duration 1 A-D NM 46 20 9,000 4.61 510 2 A-D A-D 52 18 12,308 5.25 409 3 NM NM 64 18 15,014 5.70 3074 NM A-D 54 14 9,124 4.30 416 5 None None 52 17 13,421 5.02 420 1,2,4 — — 152 52 10,067 4.74 4353,5 — — 116 35 14,280 5.40 340 All — — 268 87 11.711 5.02 412 Source:Breslow (1974 ).ReproducedwithpermissionoftheBiometricSociety. /p63Thegeometricmean of x/p16,x/p17,...,x/p76isdefinedas (/afii9811/p76/p71/p14/p16x/p71)/p16/p30/p76. Itgivesalessbiasedmeasureof centraltendencythanthearithmeticmeanwhensomeobservationsareextremelylarge. wherea/p24/p15(0)anda/p24/p16(0) are,respectively,theMLEof a/p15anda/p16givena/p17/p580.X/p42followsthechi-squaredistributionwith1degreeoffreedom.Asignificant X/p42value indicates the importance of x/p17. This can be done automatically by a stepwiseprocedure.Inaddition,ifoneormoreofthecovariatesaretreatments,the equality of survival in specified treatment groups can be tested bycomparingtheresultingmaximumlog-likelihoodvalues.Havingestimatedthecoefficientsa/p15/p71anda/p74,asurvivorshipfunctionadjustedforcovariatescanthen beestimatedfrom (11.3.9 )and (11.3.8 ). The following example, adapted from Breslow (1974 ), illustrates howthis modelcanidentifyimportantprognosticfactors. Example 11.1 Twohundredandsixty-eightchildrenwithnewlydiagnosed and previously untreated ALL were entered into a chemotherapy trial. Aftersuccessful completion of an induction course of chemotherapy designed toinduce remission, the patients were randomized onto five maintenance regi-mens designed to maintain the remission as long as possible. Maintenancechemotherapyconsistedofalternatingeight-weekcyclesof6-MPandmethot-rexate (MTX )towhichactinomycin-D (A-D)ornitrogenmustard (NM)was added.TheregimensaregiveninTable11.1.Regimen5isthecontrol.Manyinvestigators had a prior feeling that actinomycin-D was the active additivedrug; therefore, pooled regimens 1, 2, and 4 (with actinomycin-D )were comparedtoregimens3and5 (withoutactinomycin-D ).Covariatesconsidered were initial WBC and age at diagnosis. Analysis of variance showed thatdifferences between the regimens with respect to these variables were notsignificant. Table 11.1 shows that the regimen with lowest (highest )WBC geometricmeanhasthelongest (shortest )estimatedremissionduration.Figure266         Figure 11.1 Remission curves of all patients by WBC at diagnosis. (From Breslow, 1974.ReproducedwithpermissionoftheBiometricSociety. ) 11.1 gives three remission curves by WBC; differences in duration were significant.It is well known that the initial WBC is an important prognosticfactorfor patientsfollowedfromdiagnosis;however,itisinterestingtoknowif this variable will continue to be important after the patient has achievedremission. To identify important prognostic variables, model (11.3.9 )was used to analyzetheeffectofWBCandageatdiagnosis.Previousstudies (Pierceetal., 1969; George et al., 1973 )showed that survival is longest for children in the middleagerange (6—8years ),suggestingthatbothlinearandquadraticterms in age be included. The WBC was transformed by taking the commonlogarithm.Thus,thenumberofcovariatesis p/p583.Letx/p16,x/p17,andx/p18denote log/p16/p15(WBC ), age, and age squared, and a/p16,a/p17, anda/p18be the respective coefficients.Insteadofusingastepwisefittingprocedure,themodelwasfittedfivetimesusingdifferentnumbersofcovariates.Table11.2givestheresults. Theestimatedregressioncoefficientswereobtainedbysolving (11.3.11 )and (11.3.12 ). Maximum log-likelihood values were calculated by substituting the regression coefficients with the estimates in (11.3.10 ). TheX/p42values were computedfollowing (11.3.13 ),whichshowtheeffectofthecovariatesincluded. Thefirst fit did not includeany covariates.The log-likelihoodso obtainedistheunadjustedvalue l(a/p24/p15(0),0)in(11.3.13 ).Thesecondfitincludedonly x/p16or log/p16/p15(WBC ), which yields a larger log-likelihood value than the first fit. Following (11.3.13 ),weobtain X/p42/p58/p572(l(a/p24/p15(0),0)/p57l(a/p24/p15,a/p24/p16))/p58/p572(/p571332.925 /p591316.399 )/p5833.05   267 Table 11.2 Regression Coefficients and Maximum Log-Likelihood Values for Five Fits RegressionCoefficient Covariates Maximum Fit Included Log-Likelihood b/p16b/p17b/p18/afii9851/p17df 1 None /p571332.925 2x/p16(log/p16/p15WBC ) /p571316.399 0.72 33.05 1 3x/p16,x/p17(age) /p571316.111 0.73 0.02 33.63 2 4x/p17,x/p18(agesquared ) /p571327.920 /p570.24 0.018 10.01 2 5x/p16,x/p17,x/p18/p571314.065 0.67 /p570.14 0.011 37.72 3 Source:Breslow (1974 ).ReproducedwithpermissionoftheBiometricSociety. with1degreeoffreedom.Thehighlysignificant (p/p580.001 )X/p42valueindicates theimportanceofWBC.Whenageandagesquaredareincluded (fit4)inthe model,theX/p42value,10.01,islessthanthatoffit2.ThisindicatesthatWBC isabetterpredictorthanageastheonlycovariate.TotestthesignificanceofageeffectsafteradjustingforWBC,wesubtractthelog-likelihoodvalueoffit2fromthatoffit5andobtain X/p42/p58/p572(/p571316.399/p591314.065)/p584.668 with3 /p571/p582degreesoffreedom.Thesignificanceofthis X/p42valueismarginal (p/p580.10).Comparingthemaximumlog-likelihoodvalueoffit2tothatoffit 5,wefindthatlogWBCaccountsforthemajorportionofthetotalcovariateeffect. Thus, log (WBC )was identified as the most important prognostic variable. In addition, subtracting the maximum log-likelihood value of fit 5fromthatoffit3yields X/p42/p58/p572(/p571316.111/p591314.065)/p584.092 with 1 degree of freedom. This significant (p/p580.05)value indicates that the agerelationshipisindeedaquadraticone,withchildren6to8yearsoldhavingthe most favorable prognosis. For a complete analysis of the data, theinterestedreaderisreferredtoBreslow (1974 ). TouseSAStoperformtheanalysis,letTbetheremissionduration,TGan indicatorvariable (TG/p581ifinregimengroups1,2,and4;0otherwise ),CENS asecondindicatorvariable (CENS /p580 whentis censored;1 otherwise ),and x1,x2,andx3belog/p16/p15(WBC ),age,andagesquared,respectively.Assumethat thedataaresavedin‘‘C: /p33RDT.DAT’’asatextfile,whichcontainssixcolumns, and that each row (consisting of six space-separated numbers )gives the observedT,CENS,TG,x1,x2,andx3fromachild.Forinstance,afirstrow268         inRDT.DATmaybe 500 1 0 4.079 5.2 27.04 whichrepresentsthata5.2-year-oldchildwithinitiallog/p16/p15(WBC )/p584.079who received regimen 3 or 5 relapsed after 500 days [i.e., t/p58500, CENS /p581, TG/p580,x1 /p584.079,x2 /p585.2,andx3 (agesquared )/p5827.04]. Forthisdataset,thefollowingSAScodecanbeusedtoperformfits1to5 inTable11.2byusingprocedureLIFEREG. dataw1; infile‘c: /p33rdt.dat’missover; inputtcenstgx1x2x3; run;proclifereg; model1:modelt*cens (0)/p58tg/d /p58exponential; model2:modelt*cens (0)/p58tgx1/d /p58exponential; model3:modelt*cens (0)/p58tgx1x2/d /p58exponential; model4:modelt*cens (0)/p58tgx2x3/d /p58exponential; model5:modelt*cens (0)/p58tgx1x2x3/d /p58exponential; run; ForBMDPprocedure2Lthefollowingcodecanbeusedforfit5. /input file /p58‘c:/p33rdt.dat’. variables /p586. format /p58free. /print level /p58brief. /variable names /p58t,cens,tg,x1,x2,x3. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58tg,x1,x2,x3. accel /p58exponential. /end 11.4 WEIBULL REGRESSION MODEL To consider the effects of covariates, we use the model (11.2.4 ); that is, the log-survival-timeofindividual iis logT/p71/p58a/p15/p59/p78/p26 /p73/p14/p16a/p73x/p73/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71(11.4.1 ) where /afii9839/p71/p58a/p15/p59/afii9814/p78/p73/p14/p16a/p73x/p73/p71and/afii9830has the distribution defined in (11.3.2 )and   269 (11.3.3 ). This model is the Weibull regression model. Thas the Weibull distributionwith /afii9838/p71/p58exp/p1/p57/afii9839/p71/afii9846/p2and /afii9828/p581 /afii9846(11.4.2) andthefollowinghazard,density,andsurvivorshipfunctionsthatarerelated withcovariatesvia /afii9838/p71in(11.4.2 ): h(t,/afii9838/p71,/afii9828)/p58/afii9838/p71/afii9828t/p65/p92/p16 (11.4.3) f(t,/afii9838/p71,/afii9828)/p58/afii9838/p71/afii9828t/p65/p92/p16exp(/p57/afii9838/p71t/p65) (11 .4.4) S(t,/afii9838/p71,/afii9828)/p58exp(/p57/afii9838/p71t/p65) (11 .4.5) Thehazardratioofanytwoindividuals iandj,basedon (11.4.3 )and (11.4.2 ), is h/p71h/p72/p58exp/p1/p57/afii9839/p71/p57/afii9839/p72/afii9846/p2/p58exp/p1/p571 /afii9846/p78/p26 /p73/p14/p16a/p73(x/p73/p71/p57x/p73/p72)/p2 which is not time-dependent. Therefore, similar to the exponentialregression model,theWeibullregressionmodelisalsoaspecialcaseoftheproportionalhazardmodels. The following example illustrates the use of the Weibull regression model andofcomputersoftwarepackages. Example 11.2 Considerthetumor-freetimeinTable3.4.Supposethatwe wishtoknowifthreedietshavethesameeffectonthetumor-freetime.LetTbe the tumor-free time; CENS be an index (or dummy )variable with CENS /p580ifTiscensoredand1otherwise;andLOW,SATU,andUNSAbe index variables indicating that a rat was fed a low-fat, saturated fat, orunsaturated fat diet, respectively (e.g., LOW /p581 if fed a low-fat diet; 0 otherwise ).Thedatafromthe90ratsinTable3.4canbepresentedusingthese fivevariables.Forexample,thethreeobservationsinthefirstrowofTable3.4canberearrangedas TCENS LOW SATU UNSA 1 4 0 11001 2 4 1010 1 1 2 1001 Assume that the rearranged data are saved in the text file ‘‘C: /p33RAT.DAT’’, whichcontainsthedatafromthe90ratsinfivecolumnsasaboveandthefivenumbersineachrowarespace-separated.Thisdatafileisreadyforalmostall270         of the statistical software packages for parametric survival analysis currently available,suchasSASandBMDP.Supposethatthetumor-freetimefollowstheWeibulldistributionandthefollowingWeibullregressionmodelisused: logT/p71/p58a/p15/p59a/p16SATU/p71/p59a/p17UNSA/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71(11.4.6 ) where /afii9830/p71has a double exponential distribution as defined in (11.3.2 )and (11.3.3 ).Notethatfrom (11.4.3 )and (11.4.2 ), logh(t,/afii9838/p71,/afii9828)/p58log/afii9838/p71/p59log(/afii9828t/p65/p92/p16) /p58/p57/afii9839/p71/afii9846/p59log(/afii9828t/p65/p92/p16) /p58/p57a/p15/p57a/p16SATU/p71/p57a/p17UNSA/p71/afii9846/p59log(/afii9828t/p65/p92/p16) (11.4.7) Denotethehazardfunctionofaratfedanunsaturated,saturated,andlow-fat diet ash/p83,h/p81, andh/p74, respectively. From (11.4.7 ), logh/p83/p58(/p57a/p15/p57a/p17)/ /afii9846/p59log(/afii9828t/p65/p92/p16), logh/p81/p58(/p57a/p15/p57a/p16)//afii9846/p59log(/afii9828t/p65/p92/p16), and log h/p74/p58/p57a/p15/ /afii9846/p59log(/afii9828t/p65/p92/p16). Thus, the logarithm of the hazard ratio of rats fed a low-fat dietandthosefedasaturatedfatdietislog (h/p74/h/p81)/p58a/p16//afii9846,andthesimilarratios of rats fed a low-fat diet and those an unsaturated fat diet, and of rats fed asaturated fat diet and those fed an unsaturated fat diet are, respectively,log(h/p74/h/p83)/p58a/p17//afii9846andlog (h/p81/h/p83)/p58(a/p17/p57a/p16)//afii9846. Theseratiosareconstantsand are independent of time. Therefore, to test the null hypothesis that the threediets have an equal effect on tumor-free time is equivalent to testing thefollowing three hypotheses: H/p15:h/p74/h/p81/p581o ra/p16/p580,H/p15:h/p74/h/p83/p581, ora/p17/p580, andH/p15:h/p81/h/p83/p581ora/p17/p58a/p16.ThestatisticdefinedinSection9.1.1canbeused to test the first two null hypotheses,and the statistic defined in (11.2.16 )can be used for the third one. Failure to reject a null hypothesis impliesthat thecorrespondinglog-hazard ratio is not statistically different from zero; that is,therearenostatisticallysignificantdifferencesbetweenthetwocorrespondingdiets. For example, failure to reject H/p15:a/p16/p580 means that there are no significantdifferencesbetweenthe hazardsfor ratsfeda low-fatdietandratsfedasaturatedfatdiet.Whenallthreehypotheses H/p15:a/p16/p580,H/p15:a/p17/p580,and H/p15:a/p17/p58a/p16are rejected, we conclude that the three diets have significantly different effects on tumor-free time. Furthermore, a positive (negative )es- timated implies that the hazard of a rat fed a low-fat diet is exp (a/p16//afii9846) times higher (lower )thanthat ofa rat feda saturatedfat diet. Similarly,a positive (negative )estimateda/p17and (a/p17/p57a/p16)imply, respectively, the hazard of a rat fed a low-fat diet is exp (a/p17//afii9846)times higher (lower )than that of a rat fed an unsaturated fat diet, and the hazard of a rat fed a saturated fat diet isexp[ (a/p17/p57a/p16)//afii9846] timeshigher (lower )thanthatofaratfedanunsaturatedfat diet.   271 To estimate the unknown coefficients, a/p16,a/p17,a/p15, and /afii9846, we construct the log-likelihood function by replacing /afii9839in(11.4.2 ),(11.4.4 ), and (11.4.5 )with (11.4.6 ).Next,placetheresulting f(t/p71,/afii9838/p71,/afii9828) andS(t/p71,/afii9838/p71,/afii9828) inthelog-likelihood function (11.2.10 ). The log-likelihood function for the observed 90 exact or right-censoredtumor-freetimes, t/p16,t/p17,...,t/p24/p15,inthethreedietgroupsis l(a/p15,a/p16,a/p17,/afii9828)/p58/p26log[f(t/p71,/afii9838/p71,/afii9828)]/p59/p26log[S(t/p71,/afii9838/p71,/afii9828)] /p58/p26[log/afii9828/p59(/afii9828/p571)logt/p71/p57/afii9828/afii9839/p71/p57t/p65/p71exp(/p57/afii9828/afii9839/p71)] /p59/p26[/p57t/p65/p71exp(/p57/afii9828/afii9839/p71)] /p58/p26/p43log/afii9828/p59(/afii9828/p571)logt/p71/p57/afii9828(a/p15/p59a/p16SATU/p9/p59a/p17UNSA/p9) /p57t/p65/p71exp[/p57/afii9828(a/p15/p59a/p16SATU/p9/p59a/p17UNSA/p9)]/p44 /p59/p26/p43/p57t/p65/p71exp[/p57/afii9828(a/p15/p59a/p16SATU/p9/p59a/p17UNSA/p9)]/p44 The first term in the log-likelihood function sums over the uncensored observations,andthesecondtermsumsovertheright-censoredobservations.TheMLE(a/p24/p16,a/p24/p17,a/p24/p15,/afii9846/p24)of(a/p16,a/p17,a/p15,/afii9846) where /afii9846/p581//afii9828isasolutionof (11.2.12 ) with the above log-likelihood function by applying the Newton —Raphson iterative procedure. The results from SAS are shown in Table 11.3, whereINTERCPT /p58a/p15and SCALE /p58/afii9846. The MLE /afii9846/p24/p580.43,a/p24/p16/p58/p570.394, a/p24/p17/p58/p570.739,anda/p24/p17/p57a/p24/p16/p58/p570.345.H/p15:a/p16/p580(orh/p74/h/p81/p581),H/p15:a/p17/p580(or h/p74/h/p83/p581),andH/p15:a/p17/p57a/p16/p580(orh/p81/h/p83/p581)arerejectedat significancelevel p/p580.0065,p/p580.0001, andp/p580.0038, respectively. The conclusion that the dataindicate significantdifferencesamong thethree diets is the same as thatobtained in Chapter 3 using the k-sample test. Furthermore, both a/p24/p16anda/p24/p17are negative and h/p19/p74/h/p19/p81/p58exp(a/p24/p16//afii9846/p24)/p58exp(/p570.916 )/p580.40,h/p19/p74/h/p19/p83/p58exp(a/p24/p17/ /afii9846/p24)/p58exp(/p571.719 )/p580.18, andh/p19/p81/h/p19/p83/p58exp((a/p24/p17/p57a/p24/p16)//afii9846/p24)/p58exp(/p570.802 )/p580.45. Thus,basedonthedataobserved,thehazardofratsfedalow-fatdietis40%and18%ofthehazardofratsasaturatedfatdietandanunsaturatedfatdiet,respectively, and the hazard of rats fed a saturated fat diet is 45%of that ofratsfedanunsaturatedfatdiet. Thesurvivorshipfunctionin (11.4.5 )canbeestimatedbyusing (11.4.2 )and theMLEofa/p15,a/p16,a/p17,and /afii9846: S/p19(t,/afii9838,/afii9828)/p58exp(/p57/afii9838/p19t /afii9828/p24) /p58exp/p7/p57exp/p3/p571 /afii9846/p24(a/p24/p15/p59a/p24/p16SATU /p59a/p24/p17UNSA )/p4t1//afii9846/p24/p8 /p58exp[/p57exp(/p5712.56 /p590.92/p59SATU /p591.72/p59UNSA )t/p17/p13/p18/p18] Based onS/p19(t,/afii9838/p71,/afii9828), we can estimate the probabilityof surviving a given time for rats fed with any of the diets. For example, for rats fed a low-fat diet,272         Table 11.3 Analysis Results for Rat Data in Table 3.4 Using a Weibull Regression Model Regression Standard Variable Coefficient Error X/p42pexp(a/p24/p71//afii9846/p24) INTERCPT(a/p24/p15)5.400 0.113 2297 .610 0.0001 TRTSA (a/p24/p16) /p570.394 0.145 7.407 0.0065 0.40 TRTUS (a/p24/p17) /p570.739 0.140 28.049 0.0001 0.18 SCALE (/afii9846/p24) 0.430 0.043 a/p24/p17/p57a/p24/p16/p570.345 0.119 8.355 0.0038 0.45 (SATU /p580andUNSA /p580),theprobabilityofbeingtumor-freefor200daysis S/p19/p42/p45/p53(200)/p58exp[/p57exp(/p5712.56 )(200)/p17/p13/p18/p18] /p58exp[/p570.00000353 (200)/p17/p13/p18/p18]/p580.132 and for rats fed an unsaturated fat diet, (SATU /p580 and UNSA /p581), the probabilityis0.011. FollowingistheSAScodeusedtoobtainTable11.3,basedontheWeibull regressionmodelin (11.4.6 ). dataw1; infile‘c: /p33rat.dat’missover; inputtcenslowsatuunsa; run;procliferegcovout; modelt*cens (0)/p58satuunsa/d /p58weibull; run; TherespectiveBMDPprocedure2Lcodebasedon (11.4.6 )is /input file /p58‘c:/p33rat.dat’. variables /p585. format /p58free. /print level /p58brief. /variable names /p58t,cens,low,satu,unsa. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58satu,unsa. accel /p58weibull. /end   273 11.5 LOGNORMAL REGRESSION MODEL Let/afii9830in(11.2.4 )be the standard normal random variable with the density functiong(/afii9830) andsurvivorshipfunction G(/afii9830), g(/afii9830)/p58exp(/p57/afii9830/p17/2) /p402/afii9843(11.5.1 ) G(/afii9830)/p581/p57/afii9818(/afii9830)/p581/p571 /p402/afii9843/p16/p67 /p92/p27e/p92/p86/p130/p30/p17dx (11.5.2 ) where /afii9818is the cumulative distribution function of the standard normal distribution. Then the model defined by (11.2.4 )for the survival time Tof individuali, logT/p71/p58a/p15/p59/p78/p26 /p73/p14/p16a/p73x/p73/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71 isthe lognormalregressionmodel. Thasthe lognormaldistributionwiththe densityfunction f(t,/afii9839/p71,/afii9846/p17)/p58exp[/p57(logt/p57/afii9839/p71)/p17/2/afii9846/p17] /p402/afii9843/afii9846t(11.5.3) andthesurvivorshipfunction S(t,/afii9839/p71,/afii9846/p17)/p581/p57/afii9818/p1logt/p57/afii9839/p71/afii9846 /p2(11.5.4 ) It can be shown that the hazard function h(t,/afii9846,a/p15,a/p16,...,a/p78)ofTwith covariatex/p16,x/p17,...,x/p78and unknown parameters and coefficients /afii9846,a/p15, a/p16,...,a/p78canbewrittenas logh(t,/afii9846,a/p15,a/p16,...,a/p78)/p58logh/p15[texp(/p57/afii9839)]/p57/afii9839 (11.5.5 ) whereh/p15(·)isthehazardfunctionofanindividualwithallcovariatesequalto zero. Equation (11.5.5 )indicates that h(t,/afii9846,a/p15,a/p16,...,a/p78)is a function of h/p15evaluatedattexp(/p57/afii9839), not independentof t. Thus, the lognormal regression modelisnotaproportionalhazardsmodel. Example 11.3 Considerthesurvivaltimedatafrom30patientswithAML inTable11.4.Twopossibleprognosticfactorsorcovariates,age,andcellular-274         Table 11.4 Survival Times and Data for Two Possible Prognostic Factors of 30 AML Patients SurvivalTime x/p16x/p17SurvivalTime x/p16x/p17 18 0 0 8 1 0 90 1 2 1 1 28/p59 00 2 6 /p59 10 31 0 1 10 1 1 39/p59 014 10 19/p59 013 10 45/p59 014 10 60 1 1 8 1 180 1 8 1 1 15 0 1 3 1 1 23 0 0 14 1 1 28/p59 003 10 70 1 1 3 1 1 12 1 0 13 1 1 91 0 3 5 /p59 10 itystatusareconsidered: x/p16/p58/p71 if patient is /p4650 years old 0 otherwise x/p17/p58/p71 if cellularity of marrowclot section is 100% 0 otherwise Letususethelognormalregressionmodel logT/p71/p58a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71/p59/afii9846/afii9830/p71(11.5.6 ) and /afii9839/p71/p58a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71(11.5.7) Theunknowncoefficientsandparameter a/p16,a/p17,a/p15,/afii9846needtobeestimated. Weconstructthelog-likelihoodfunctionbyreplacing /afii9839in(11.5.3 )and (11.5.4 ) with (11.5.7 ), then replacing f(t/p71,/afii9839,/afii9846/p17) andS(t/p71,/afii9839,/afii9846/p17) in the log-likelihood function (11.2.5 )with their expression (11.5.3 )and (11.5.4 ), respectively. The resultinglog-likelihoodfunctionfortheexactandright-censoredsurvivaltimes   275 Table 11.5 Asymptotic Likelihood Inference for Data on 30 AML Patients Using a Lognormal Regression Model Regression Standard Variable /p63 Coefficient Error X/p42p INTERCPT(a/p15) 3.3002 0.3750 77.4675 0.0001 x/p16(a/p16) /p571.0417 0.3605 8.3475 0.0039 x/p17(a/p17) /p570.2687 0.3568 0.5672 0.4514 SCALE (/afii9846) 0.9075 0.1409 /p63x/p16/p581 if patient /p4650 yearsold, and0 otherwise; x/p17/p581 if cellularityof marrowclot section is 100%,and0otherwise.observedfromthe30patientswithAMLis l(a/p15,a/p16,a/p17,/afii9846)/p58/p26/p3/p57(logt/p71/p57/afii9839/p71)/p17 2/afii9846/p17/p57log(/p402/afii9843/afii9846t/p71)/p4 /p59/p26/p7log/p31/p57/afii9818/p1logt/p71/p57/afii9839/p71/afii9846 /p2/p4/p8 /p58/p26/p7/p57[logt/p71/p57(a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71)]/p17 2/afii9846/p17/p57log(/afii9846t/p71/p402/afii9843)/p8 /p59/p26/p7log/p31/p57/afii9818(logt/p71/p57(a/p15/p59a/p16x/p16/p71/p59a/p17x/p17/p71) /afii9846 /p4/p8 The first term in the log-likelihood function sums over the uncensored observations,and the second sumsover the right-censoredobservations.TheMLE (a/p24/p16,a/p24/p17,a/p24/p15,/afii9846/p24)o f(a/p16,a/p17,a/p15,/afii9846) can be obtained by applying the Newton—Raphsoniterativeprocedure.The hypothesis-testingproceduresdis- cussed in Section 9.1.2 can be used to test whether the coefficients a/p16anda/p17areequaltozero.Table11.5showsthat a/p16issignificantly (p/p580.0039 )different fromzero,while a/p17isnot (p/p580.4514 ).Thesignsoftheregressioncoefficients indicatethatageover50yearshassignificantlynegativeeffectsonthesurvivaltime,whilea100%cellularityofmarrowclotsectionalsohasanegativeeffect;however,theeffectisnotofsignificantimportancetothesurvivaltime. LetTbethesurvivaltimeandCENSbeanindex (ordummy )variable with CENS /p580i fTis censored and 1 otherwise. Assume that the data are saved in a text file ‘‘C: /p33AML.DAT’’ with four numbers in each row, space- separated,whichcontainssuccessively T,CENS,x1,andx2. LetTbethesurvivaltimeandCENSbeanindex (ordummy )variablewith CENS /p580ifTiscensoredand1otherwise.Assumethatthedataaresavedin a text file ‘‘C: /p33AML.DAT’’ with four numbers in each row, space-separated, which contains successively T, CENS, x1, and x2. The following SAS code is usedtoobtaintheresultsinTable11.5.276         dataw1; infile‘c: /p33aml.dat’missover; inputtcensx1x2; run; proclifereg; model1:modelt*cens (0)/p58x1x2/d /p58lnormal; run; IfBMDPisused,thefollowing2Lcodeissuggested. /input file /p58‘c:/p33aml.dat’. variables /p584. format /p58free. /print level /p58brief. /variable names /p58t,cens,x1,x2. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58x1,x2. accel /p58lnormal. /end 11.6 EXTENDED GENERALIZED GAMMA REGRESSION MODEL Inthissectionweintroducea regressionmodelthat isbasedonanextended formofthegeneralizedgammadistributiondefinedinSection6.4.Assumethatthesurvivaltime Tofindividualiandcovariates x/p16,...,x/p78havetherelation- shipgivenin (11.4.1 ),where /afii9830hasthelog-gammadistributionwiththedensity functiong(/afii9830) andsurvivorshipfunction G(/afii9830): g(/afii9830)/p58/p34/afii9829/p34[exp (/afii9829/afii9830)//afii9829/p17]/p16/p30/p66/p130exp[/p57exp(/afii9829/afii9830)//afii9829/p17] /afii9772(1//afii9829/p17)(11.6.1 ) G(/afii9830)/p58/p7I/p3exp(/afii9829/afii9830) /afii9829/p17,1 /afii9829/p17/p4if/afii9829/p580 1/p57I/p3exp(/afii9829/afii9830) /afii9829/p17,1 /afii9829/p17/p4if/afii9829/p570/p57/p45/p58/afii9830/p58 /p59/p45(11.6.2 ) (11.6.3 ) This model is the extended generalized gamma regression model. It can be shown thatThas the extended generalized gamma distribution with the densityfunction f(t,/afii9825,/afii9838,/afii9828)/p58/p34/afii9825/p34/afii9828/p65/afii9838/p63/p65/p71t/p63/p65/p92/p16exp[/p57/afii9828(/afii9838/p71t)/p63] /afii9772(/afii9828)(11.6.4) andsurvivorshipfunction S(t,/afii9825,/afii9838,/afii9828)/p58/p7I(/afii9828(/afii9838/p71t)/p63,/afii9828) 1/p57I(/afii9828(/afii9838/p71t)/p63,/afii9828)if/afii9825/p580 if/afii9825/p570(11.6.5 ) (11.6.6 )     277 where /afii9838/p71/p58exp(/p57/afii9839/p71) /afii9825/p58/afii9829 /afii9846/afii9828/p581 /afii9829/p17(11.6.7 ) /afii9772(x) isthecompletegammafunctiondefinedin (6.2.9 ),I(a,x)istheincomplete gamma function defined in (6.4.4 ), and /afii9829is a shape parameter. We used the extended generalized gamma distribution in (11.6.4 )here because it is the distribution used in SAS. The derivation is left to the reader as an exercise (Exercise11.12 ). The estimation procedures for the parameters, regression coefficients, and the covariate adjusted survivorship function are similar to those discussed inSections11.3and11.4. Example 11.4 Consider the survival times (T)in days and a set of prognostic factors or covariates from 137 lung cancer patients, presented inAppendix I of Kalbfleisch and Prentice (1980 ). The covariates include the Karnofskymeasureofthe overallperformancestatus (KPS )ofthe patientat entry into the trial, time in months from diagnosis to entry into the trial(DIAGTIME ), age in years (AGE ), prior therapy (INDPRI, yes or no ), histological type of tumor, and type of therapy. There are four histologicaltypes of tumor: adeno, small, large, and squamous cell and two types oftherapies: standard and experimental. The values of KPS have the followingmeanings: 10—30 completely hospitalized, 40 —60 partial confinement, 70 —90 able to care for self. Assume that the survival time follows the extendedgeneralizedgammaregressionmodel,wewishtoidentifythe mostsignificantprognosticvariables. First we define several index (or dummy )variables for the categorical variablesandthecensoringstatus.LetCENS /p580whenthesurvivaltime Tis censoredand1otherwise;INDADE /p581,INDSMA /p581,andINDSQU /p581if the type of cancer cell is adeno, small, and squamous, respectively, and 0otherwise; INDTHE /p581 if the standard therapy is received and 0 otherwise; andINDPRI /p581ifthereisapriortherapyand0otherwise.Themodelis logT/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17AGE/p71/p59a/p18DIAGTIME/p71/p59a/p19INDPRI/p71 /p59a/p20INDTHE/p71/p59a/p21INDADE/p71/p59a/p22INDSMA/p71/p59a/p23INDSQU/p71/p59/afii9846/afii9830/p71 (11.6.8 ) wherethedensityfunctionof /afii9830/p71isdefinedin (11.6.1 ).Thus, /afii9839/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17AGE/p71/p59a/p18DIAGTIME/p71/p59a/p19INDPRI/p71/p59a/p20INDTHE/p71 /p59a/p21INDADE/p71/p59a/p22INDSMA/p71/p59a/p23INDSQU/p71(11.6.9 ) Toestimatea/p16,...,a/p23,/afii9829,a/p15,and /afii9846,weconstructthelog-likelihoodfunctionby replacing /afii9839in(11.6.7 )and (11.6.4 )—(11.6.6 )with (11.6.9 ),thenreplacing f(t/p71,b) andS(t/p71,b)inthelikelihoodfunction (11.2.10 )bythosein (11.6.4 )and (11.6.5 ) or(11.6.6 ). The MLE (a/p24/p16,...,a/p24/p23,/afii9829/p19,a/p24/p15,/afii9846/p24)of (a/p16,...,a/p23,/afii9829,a/p15,/afii9846) can be278         Table 11.6 Asymptotic Likelihood Inference on Lung Cancer Data Using a Generalized Gamma Regression Model Regression Standard Variable Coefficient Error X/p42p INTERCPT(a/p15)2 .176 0 .719 9 .143 0.003 INDADE(a/p21) /p570.759 0.286 7.034 0.008 INDSMA(a/p22) /p570.594 0.264 5.059 0.025 INDSQU(a/p23) 0.150 0.291 0.266 0.606 KPS(a/p16) 0.034 0.005 46.443 0.000 AGE(a/p17) 0.008 0.009 0.845 0.358 DIAGTIME(a/p18) 0.000 0.009 0.001 0.980 INDPRI(a/p19) /p570.089 0.216 0.171 0.679 INDTHE(a/p20) 0.168 0.185 0.823 0.364 SCALE (/afii9846) 1.000 0.071 SHAPE (/afii9829) 0.450 0.223 INTERCPT(a/p15) 2.748 0.396 48.247 0.000 INDADE(a/p21) /p570.766 0.280 7.492 0.006 INDSMA(a/p22) /p570.534 0.258 4.284 0.039 INDSQU(a/p23) 0.144 0.280 0.264 0.608 KPS(a/p16) 0.033 0.005 45.497 0.000 SCALE (/afii9846) 1.004 0.070 SHAPE (/afii9829) 0.473 0.206 obtained in a manner similar to that used in Examples 11.2 and 11.3. The hypothesis-testing procedure defined in Section 11.2 can be used to testwhetherthecoefficients a/p16,a/p17,...,a/p23areequaltozero.ThefirstpartofTable 11.6 shows the results from SAS (where INTERCPT /p58a/p15, SCALE /p58/afii9846, and SHAPE /p58/afii9829). Table 11.6 shows that a/p16,a/p21, anda/p22are significantly (p/p580.05)different fromzero,whereastheothercovariatesarenot (p/p570.05).Thatis,onlyKPS and the type of cancer cell have significant effects on the survival time. Inparticular, adeno cell carcinoma and small cell carcinoma have significantnegativeeffectsonsurvivaltime.PatientswhohavebetterKarnofskyperform-ance status have a longer survival time. If we wish to include only KPS andcelltypeinthemodel,thelowerpartofTable11.6givestheresults. Assume that the coded data are saved in ‘‘C: /p33LCANCER.DAT’’as a text file with 10 numbers in a row, space-separated, which contains data for T, CENS, KPS, AGE, DIAGTIME, INDPRI, INDTHE, INDADE, INDSMA,andINDSQU,inthatorder.TheSAScodeusedtoobtainTable11.6is dataw1; infile‘c: /p33lcancer.dat’missover; inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu; run;     279 proclifereg; Model1: modelt*cens (0)/p58kpsage diagtimeindpriindtheindadeindsmaindsqu / d/p58gamma; Model2:modelt*cens (0)/p58kpsindadeindsmaindsqu/d /p58gamma; run; 11.7 LOG-LOGISTIC REGRESSION MODEL Assumethattherelationshipbetweenthesurvivaltime T/p71forindividualiand aset of covariates, x/p16,...,x/p78canbe expressedbythe AFTmodel in (11.4.1 ), where /afii9830/p71hasalogisticdistributionwiththedensityfunction g(/afii9830)/p58exp(/afii9830) [1/p59exp(/afii9830)]/p17(11.7.1 ) andsurvivorshipfunction G(/afii9830)/p581 1/p59exp(/afii9830)(11.7.2 ) This model is the log-logistic regression model. Then Thas the log-logistic distribution defined in Section 6.5. The parameter /afii9825in the distribution is a functionofthecovariates: /afii9825/p71/p58exp/p1/p57/afii9839/p71/afii9846/p2/afii9828/p581 /afii9846(11.7.3) Substituting (11.7.3 )inthesurvivorshipfunctionin (6.5.2 ),weobtain logS(t,b) 1/p57S(t,b)/p58/p57log(/afii9825t/p65)/p58/afii9839 /afii9846/p57/afii9828logt(11.7.4) or logS(t,b) 1/p57S(t,b)/p58a/p15/afii9846/p591 /afii9846/p78/p26 /p73/p14/p16a/p73x/p73/p57/afii9828logt(11.7.5) whereb/p58(a/p15,a/p16,...,a/p78,/afii9846).SinceS(t/p71,b)istheprobabilityofsurvivinglonger thant,S(t/p71,b)/[1/p57S(t/p71,b)]istheoddsofsurvivinglongerthan t.LetOR/p71and OR/p72denote the odds of surviving longer than tfor individuals iandj, respectively.Thelogarithmoftheoddsratiois logOR/p71OR/p72/p581 /afii9846/p78/p26 /p73/p14/p16a/p73(x/p73/p71/p57x/p73/p72) (11 .7.6) Thisratio isindependentoftime.Therefore,thelog-logisticregressionmodel isaproportionaloddsmodel,notaproportionalhazardsmodel.280         Example 11.5 Wefitthelog-logisticregressionmodelabovetothedatain Example11.6.1usingonlyKPSandthethreecancercelltypeindexvariables.Thatis, logT/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71/p59/afii9846/afii9830/p71 (11.7.7 ) wherethedensityfunctionof /afii9830/p71isdefinedin (11.7.1 ).Thus, /afii9839/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71(11.7.8 ) Toestimate b/p58(a/p15,a/p16,...,a/p19,/afii9846)/p30,weconstructthelog-likelihoodfunctionby using the /afii9825and/afii9828in(11.7.3 )as parameters in the density and survivorship functions of the log-logistic distribution in Section 6.5. The resulting log-likelihoodfunctionforthe137observedexactorright-censoredsurvivaltimesis l(b)/p58/p26 /p7/p57/afii9839/p71/afii9846/p57log/afii9846/p591/p57/afii9846 /afii9846logt/p71/p572log /p31/p59exp/p1/p57/afii9839/p71/afii9846/p2t/p16/p30/p78/p71/p4/p8 /p59/p26/p7/p57log/p31/p59exp/p1/p57/afii9839/p71/afii9846/p2t/p16/p92/p78/p71/p4/p8 /p58/p26/p1/p571 /afii9846(a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71) /p57log/afii9846/p591/p57/afii9846 /afii9846logt/p71 /p572log /p71/p59exp/p3/p571 /afii9846(a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71 /p59a/p19INDSQU/p71)/p4t/p16/p30/p78/p71/p8/p2 /p57/p26/p1log/p71/p59exp/p3/p571 /afii9846(a/p15/p59a/p16KPS/p71/p59a/p17INDADE/p71/p59a/p18INDSMA/p71 /p59a/p19INDSQU/p71)/p4t/p16/p30/p78/p71/p8/p2 The first term in the log-likelihood function sums over the uncensored observations,and the second sumsover the right-censoredobservations.TheMLE(a/p24/p16,...,a/p24/p19,a/p24/p15,/afii9846/p24)o f(a/p16,...,a/p19,a/p15,/afii9846) aregiveninTable11.7,withtheir-   281 Table 11.7 Asymptotic Likelihood Inference on Lung Cancer Data Using a Log-Logistic Regression Model Regression Standard Variable Coefficient Error X/p42pexp(a/p71/a) INTERCPT(a/p15)2.451 0.344 50.911 0.000 — INDADE(a/p17) /p570.749 0.261 8.217 0.004 0.275 INDSMA(a/p18) /p570.661 0.240 7.565 0.006 0.321 INDSQU(a/p19)0.029 0.264 0.012 0.913 1.051 KPS(a/p16) 0.036 0.004 66.885 0.000 1.064 SCALE (/afii9846) 0.581 0.043 — — — standarderrors, likelihood ratio test statistics (X/p42), andpvalues. The results aresimilartothoseobtainedfromfittingthegeneralgammaregressionmodelinExample11.4. In addition, using (11.7.5 )and (11.7.6 ), we can obtain odds ratios for the covariates. For example, let the odds of surviving to time tfor four patients withthesameKPSbutdifferentcelltype (adeno,small,squamousandlarge ) bedenotedbyOR/p31/p34,OR/p49/p43,OR/p49/p47,andOR/p42/p31,respectively;thenthelog-odds ratiosoftheindividualswithadeno,small,andsquamouscelltypestotheonewithlargecelltypeare,respectively, logOR/p31/p34OR/p42/p31 /p58a/p17/afii9846logOR/p49/p43OR/p42/p31/p58a/p18/afii9846logOR/p49/p47OR/p42/p31/p58a/p19/afii9846 Replacinga/p17,a/p18,a/p19,and /afii9846withtheirestimates,wehave OR/p31/p34OR/p42/p31/p58exp/p1a/p24/p17/afii9846/p24/p2/p580.275 OR/p49/p43OR/p42/p31/p58exp/p1a/p24/p18/afii9846/p24/p2/p580.321 OR/p49/p47OR/p42/p31/p58exp/p1a/p24/p19/afii9846/p24/p2/p581.051 Theseresultsmeanthatinlungcancerpatients,personswithadenoandsmall cell type have odds of only about one-fourth and one-third, respectively, ofthose with large cell type. The odds of persons with large cell carcinoma arenotsignificantlydifferentfromthoseofpatientswithsquamouscellcarcinoma.Further,whenignoringcelltype,exp (a/p24/p16//afii9846/p24)representsanincrease (ordecrease ) in the odds for any 1-unit increase in the KPS measure. In this case,282         exp(a/p24/p16//afii9846/p24)/p581.064; thus for a 1-unit increase in the KPS measure, the odds increaseby6.4%.Theresultsare,ingeneral,consistentwiththoseobtainedinExample11.4. ThefollowingSAScodecanbeusedtoobtaintheresultsinTable11.7. dataw1; infile‘c: /p33lcancer.dat’missover; inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu; run;proclifereg; modelt*cens (0)/p58kpsindadeindsmaindsqu/d /p58llogistic; run; ThefollowingBMDP2Lcodeisalsoapplicable. /input file /p58‘c:/p33lcancer.dat’. variables /p5810. format /p58free. /print level /p58brief. /variable names /p58t,cens,kps,age,diagtime,indpri,indthe,indade,indsma,indsqu. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58kps,indade,indsma,indsqu. accel /p58llogistic. /end 11.8 OTHER PARAMETRIC REGRESSION MODELS Inthissectionwediscusstwomodelsinwhichthesurvivaltime Tisassumed tofollowtheexponentialdistributionwithdensityandsurvivorshipfunctionsasdefinedin (6.1.1 )and (6.1.3 ),respectively,andthemeansurvivaltime1/ /afii9838or hazardrate /afii9838hasthefollowinglinearrelationshipwiththecovariates: Model1:1 /afii9838/p71/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71 Model2: /afii9838/p71/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71 Model 1 is considered by Feigl and Zelen (1965 )and extended to include censoreddatabyZippinandArmitage (1966 ).Model2isusedbyByaretal. (1974 ).    283 Model 1 Supposethatnpatientsareenteredinastudy; rofthesedieand s/p58n/p57rare still alive at the end of the study. Let t/p16,...,t/p80be the exact survival times of therdeaths andt/p62/p16,...,t/p62/p81be thescensoring times. Furthermore, let x/p72/p71, i/p581,...,n,j/p581,...,p, be the observed value of the jth covariate of the ith patient.Themodelassumesthat themeansurvivaltime islinearlyrelatedtothecovariates: 1 /afii9838/p71/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71/p58/p78/p26 /p72/p14/p15a/p72x/p72/p71(11.8.1) wherex/p15/p71/p891.Theterma/p15representstheunderlyinghazardinthesensethat 1/a/p15isthehazardrate /afii9838/p71whencovariatesareignoredorall x/p72/p71’sarezero.Then thelikelihoodfunctionofthe nsurvivaltimesunderthemodel (11.8.1 )canbe writtenas L(a/p15,a/p16,...,a/p78)/p58/p80/p147 /p71/p14/p16/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2/p92/p16exp/p3/p57t/p71/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2/p92/p16/p4 /p59/p81/p147 /p73/p14/p16exp/p3/p57t/p62/p73/p1/p78/p26 /p72/p14/p15a/p72x/p72/p73/p2/p92/p16/p4(11.8.2) Thelog-likelihoodisthen l(a/p15,a/p16,...,a/p78)/p58/p57/p80/p26 /p71/p14/p16log/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2/p57/p80/p26 /p71/p14/p16t/p71/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2/p92/p16 /p57/p80/p26 /p73/p14/p16t/p62/p73/p1/p78/p26 /p72/p14/p15a/p72x/p72/p73/p2/p92/p16(11.8.3) The maximumlikelihood estimates of a/p72,j/p580, 1,...,p/p72, may be obtained bysolvingsimultaneouslythe p/p591equations: /p57/p80/p26 /p71/p14/p16/p1/p78/p26 /p72/p14/p16a/p72x/p72/p71/p2/p59/p80/p26 /p71/p14/p16t/p71/p1/p78/p26 /p72/p14/p16a/p72x/p72/p71/p2/p92/p17/p59/p81/p26 /p73/p14/p16t/p62/p73/p1/p78/p26 /p72/p14/p16a/p72x/p72/p73/p2/p92/p17/p580 /p57/p80/p26 /p71/p14/p16x/p72/p71/p1/p78/p26 /p72/p14/p16a/p72x/p72/p71/p2/p92/p16/p59/p80/p26 /p71/p14/p16t/p71x/p72/p71/p1/p78/p26 /p72/p14/p16a/p72x/p72/p71/p2/p92/p17 /p59/p81/p26 /p73/p14/p16t/p62/p73x/p72/p73/p1/p78/p26 /p72/p14/p16a/p72x/p72/p73/p2/p92/p17/p580j/p581,...,p (11.8.4) Again,thiscanbedonebyNewton —Raphsoniterativeproceduresdescribedin Section7.1.284         AfterobtainingtheMLE, a/p24/p72,j/p580,1,...,p,thelog-likelihoodfunctioncan beusedtotestthesignificanceofthecovariates.TheprocedureisexactlythesameasthoseusedinExample11.1. The survivorship function (for theith patient )adjusted for the covariates canbeobtainedfrom S/p19/p71(t)/p58exp(/p57/afii9838/p19/p71t) /p58exp/p3/p57t/p1/p78/p26 /p72/p14/p15a/p24/p72x/p72/p71/p2/p92/p16/p4(11.8.5 ) Model 2 Byaretal. (1974 )developedanotherexponentialmodelrelatingsurvivaltime toconcomitantinformationforprostatecancerpatientsinwhichtheindividualhazardislinearlyrelatedtothepossibleprognosticvariables: /afii9838/p71/p58a/p15/p59/p78/p26 /p72/p14/p16a/p72x/p72/p71/p58/p78/p26 /p72/p14/p15a/p72x/p72/p71(11.8.6) wherex/p15/p71/p581. Similar to the model of Feigl and Zelen (1965 ),a/p15is the underlyinghazardratewhen covariatesareignored,forceof mortalityortheintercept. Supposethatrofthenpatientsaredeadand s/p58n/p57rarestillaliveatthe endofthestudy;thenthelikelihoodfunctionis L(a/p15,a/p16,...,a/p78)/p58/p80/p147 /p71/p14/p16/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2exp/p3/p57/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2t/p71/p4 /p59/p81/p147 /p73/p14/p16exp/p3/p57/p1/p78/p26 /p72/p14/p15a/p72x/p72/p73/p2t/p62/p73/p4(11.8.7) Takingthelogarithmsof (11.8.7 ),weobtainthelog-likelihoodfunction l(a/p15,a/p16,...,a/p78)/p58/p80/p26 /p71/p14/p16/p3log/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2/p57/p1/p78/p26 /p72/p14/p15a/p72x/p72/p71/p2t/p71/p4/p57/p81/p26 /p73/p14/p16/p1/p78/p26 /p72/p14/p15a/p72x/p72/p73/p2t/p62/p73 (11.8.8) ToobtaintheMLEsofthe a/p72’s,weneedtosolvesimultaneouslythefollowing p/p591equations: /p80/p26 /p71/p14/p16/p1x/p72/p71/p26/p78/p72/p14/p15a/p24/p72x/p72/p71/p57x/p72/p71t/p71/p2/p57/p81/p26 /p73/p14/p16x/p72/p73t/p62/p73/p580j/p580,1,...,p(11.8.9 )    285 TheseequationscanbesolvedsimultaneouslybyusingtheNewton —Raphson iterativeprocedure. After the MLEs of a/p72,j/p580, 1,...,p, are obtained, the log-likelihood functioncanbeusedtotestthesignificanceofthecovariatesbyfollowingthesameprocedureasthatusedinExample11.1. Thesurvivorshipfunctionforthe ithindividualadjustedforthecovariates canbeestimatedfrom S/p19/p71(t)/p58exp(/p57/afii9838/p19/p71t) /p58exp/p1/p57t/p78/p26 /p72/p14/p15a/p24/p72x/p72/p71/p2(11.8.10) 11.9 MODEL SELECTION METHODS Toidentifyimportantriskfactors usinga parametricapproach,oneneedsto select a most appropriateparametric model and identify the most significantsubset of covariates. In this section we first discuss, for a given parametricmodel,howtochooseanoptimalsubsetofthecovariatesthathavestatisticallysignificant effects on the survival time. Second, we consider if the significantcovariates are known, how to determine which parametric model is mostappropriate.Third,wediscussamethodthatcanbeusedtocompareamongparametricmodelswithdifferentsubsetsofcovariates. 11.9.1 SelectionofMost SignificantCovariatesfor a Known ParametricModel For a known parametricmodel, the following methods can be used to select anoptimalsubsetofthecovariatesinthesensethatthesubsetselectedhasthemost statistically significant effects on the survival time among all subsets ofthe covariates. These methods include the forward, backward, stepwise, AIC,andBICselectionprocedurescommonlyusedinordinaryregressionanalyses.Wegiveonlyabriefoutlinehere.Interestedreadersare referredtobooksonordinaryregressionanalysis. Forward Selection Procedure Theforwardselectionprocedureis anadding processinwhichonecovariateisselectedandaddedtothemodelateverystep.First,wehavetoestimatethespecificparametersthatdefinetheparametricmodelandthecoefficientsoftheadjusting covariates, if any, that are forced into the model. For example, tohaveage-andgender-adjustedresults,ageandgendermustbeincludedinthemodel, whether or not they are significant. Then the adjusted chi-squarestatisticsforeachcovariatenotinthemodelarecomputedandthelargestofthesestatisticsisidentified.Ifthelargestchi-squarestatisticissignificantatthe286         /afii9825level specified (usually, /afii9825/p580.15)for entry, the corresponding covariate is addedtothemodel. Leta/p16bethevectoroftheparametersorcoefficientsofcovariatesalreadyin the model and l(·) be the log-likelihood function. The forward selection procedurewillselect x/p72,whichisnotyetinthemodel,toenterthemodelifthe difference between the log-likelihood values with x/p72and withoutx/p72is largest among all the x/p73’s that are not in the model. That is, the coefficient a/p72ofx/p72satisfies X/p42/p582[(l(a/p24/p72,a/p24/p16)/p57l(a/p24/p16/p72(0))] /p58max /p73/p432[l(a/p24/p73,a/p24/p16)/p57l(a/p24/p16/p73(0))],foranyx/p73thatisnotinthemodel /p44 (11.9.1 ) andX/p42/p57/afii9851/p17/p16/p11/p63wherea/p73isthecoefficientof x/p73notyetinthemodel,( a/p24/p73,a/p24/p16),is theMLEof (a/p73,a/p16),a/p24/p16/p73(0)istheMLEof a/p16givena/p73/p580,and /afii9851/p17/p16/p11/p63isthe /afii9825-level critical point of the chi-square distribution with 1 degree of freedom. In theforwardselectionprocedure,onceacovariateisenteredintothemodel,itwillnever be removed. The process is repeated until none of the remainingcovariatesmeetthelevel /afii9825specifiedforentryoruntilapredeterminednumber ofcovariateshavebeenentered. Backward Selection Procedure The backward selection procedure is an elimination process in which all thecovariatesareincludedinthemodelatthebeginningandareremovedonebyone according to a significance criterion. The specific parameters that definethe parametric model and the coefficients of all the covariates are estimatedfirst.ThentheWaldtestisusedtoexamineeachcovariate.Theleastsignificantcovariate that does not meet the specified level /afii9825(usually, /afii9825/p580.15)for stayinginthemodelisremoved.Thatis,covariate x/p72willberemovedfromthe modelif X/p53/p58a/p24/p17/p72v/p17/p72/p72 /p58min /p73/p1a/p24/p17/p73v/p17/p73/p73foranyx/p73thatisinthemodel /p2(11.9.2 ) andX/p53/p45/afii9851/p17/p16/p11/p63wherea/p72is the corresponding coefficient for x/p72andv/p17/p72/p72is the estimatedvarianceof a/p24/p72.Inthebackwardselectionprocedure,onceacovariate isremovedfromthemodel,itremainsexcluded.Theprocessisrepeateduntilallthecovariatesremainedinthemodelmeetthespecifiedsignificancelevel /afii9825 forstayingoruntilapredeterminednumberofcovariatesremaininthemodel.The advantages of the backward procedure have been discussed by Mantel(1970 ).   287 Stepwise Selection Procedure The stepwise selection procedure is a combinationof forward and backwardselectionprocedures. At first, it issimilar to the forwardselection procedure;however,covariatesalreadyinthemodeldonotnecessarilyremain.Covariatesalreadyinthemodelmayberemovedlateriftheyarenolongersignificant.Thestepwiseselectionprocessterminatesifnosignificantcovariatecanbeaddedtothe model or if the covariate just entered into the model is removed and nomorecovariatescanbeadded. Information Criterion (AIC and BIC) Procedures The Baysian information criterion (BIC)selection procedure discussed in Section 9.3 can be used to select the best parametric model with covariates.Thiscanbedoneeasilybyreplacingthelog-likelihoodfunction l(b/p19)in(9.3.1 ) withthelog-likelihoodfunctionwithsubsetsofcovariatesdefinedinprevioussections of this chapter. The subset of covariates that produces the largest r value in (9.3.1 )among all possible subsets is the choice. If the number of covariates is large, one may apply the forward, backward, and stepwiseselectionmethodfirstto reducethenumberofcandidatecovariates,thenusetheBICprocedure.TheAICcriterioncanbeappliedinasimilarmanner. 11.9.2 Selection of a Parametric Model with a Fixed Subset of Covariates Ifthemostsignificantsubsetofcovariatesisknown,selectionofanappropriate parametricmodelcanbecarriedoutbyusingaproceduresimilartothosebasedon the likelihood functions and discussed in Section 9.2. The procedures areexactlythesameexceptthatallthelikelihoodfunctionsarereplacedbythosewith covariates, for example, those given in (11.3.10 )and Examples 11.2 and 11.5.Withcomputersoftwarepackagesavailablecommercially,theprocedurecaneasilybeapplied.Thefollowingexampleillustratestheapplication. Example 11.6 Considerthe lungcancerpatientswhodid not receiveany prior therapy in Example 11.4. Assume that the three covariates KPS,INDADE,andINDSMAaremostsignificant.Forthesethreefixedcovariates,the log-likelihood values based on the exponential, Weibull, lognormal, log-logistic, and generalized gamma models are given in Table 11.8. From thistable,thelognormal,Weibullandexponentialmodels (relativetothegeneral- izedgammamodel ),withthethreecovariates,arerejectedat /afii9825/p580.0325,0.016 and 0.024, respectively. It appears that the exponentialmodel, relative to theWeibull, is not rejected (p/p580.194 ). However, since the exponential model belongs to the Weibull distribution family and the Weibull model has beenrejected,theexponentialmodelwiththethreecovariatesisnotappropriateforthe data, as noted earlier in Chapter 9. Thus, we conclude that none of thethree models (exponential, Weibull, and lognormal ), with covariates, provide anappropriatefittothedata.InExample11.7wewillseethatthelog-logisticmodelisthebestfitamongallthesemodels.288         Table 11.8 Goodness-of-Fit Tests Based on Asymptotic Likelihood Inference on Lung Cancer Data Distribution LL /p63LLR /p63p/p63BIC AIC Extended /p57132.793 — — /p57146.517 /p57144.793 generalizedgamma Log-logistic /p57131.230 — — /p57142.667 /p57141.230 Lognormal /p57135.022 4.459 /p640.035 /p57146.459 /p57145.022 Weibull /p57135.669 5.752 /p650.016 /p57147.106 /p57145.669 Exponential /p57136.512 7.438 /p660.024 /p57145.661 /p57144.512 Exponential /p57136.512 1.686 /p670.194 — — /p63LL,log-likelihood;LLR,log-likelihoodratiostatistic; p,probabilitythattherespectivechi-square randomvariable /p57LLR. /p64Lognormalrelativetoextendedgeneralizedgamma. /p65Weibullrelativetoextendedgeneralizedgamma. /p66Exponentialrelativetoextendedgeneralizedgamma. /p67ExponentialrelativetoWeibull. Using the data file ‘‘C: /p33LCANCER.DAT’’ described in Example 11.4, the followingSAScodecanbeusedtoobtainTable11.8. dataw1; infile‘c: /p33lcancer.dat’missover; inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu;ifindpri /p580; run; proclifereg; Model1:modelt*cens (0)/p58kpsindadeindsma/d /p58exponential; Model2:modelt*cens (0)/p58kpsindadeindsma/d /p58weibull; Model3:modelt*cens (0)/p58kpsindadeindsma/d /p58lnormal; Model4:modelt*cens (0)/p58kpsindadeindsma/d /p58gamma; Model5:modelt*cens (0)/p58kpsindadeindsma/d /p58llogistic; run; 11.9.3 Selection of a Parametric Model and an Optimal Subset of Covariates Simultaneously: AIC and BIC Procedures TheextendedAICandBICcriteria,whichincludecovariates,canbeapplied notonlytoselectthemostsignificantcovariatesforagivenparametricmodel,but also, simultaneously, to select the best parametric model. The proceduremaybetediousifthenumberofcovariatestobeconsideredislarge.However,in practice, the number of covariates worthy of consideration in a model isusuallyreducedafter univariateanalyses,as describedin Section11.1.There-fore,theAICorBICproceduremaynotbetoodifficulttoapply.Withtheaidof software packages, we can apply the forward, backward, and stepwise   289 selection methods in Section 11.9.1 first to fit different parametric regression models to the data and then include in the AIC or BIC procedure all or asubsetofthesignificantcovariatesidentifiedineachfit. Example 11.7 Consider the lung cancer data of Example 11.6. We apply themethodsofSection11.9.1toselectthebestsubsetofcovariatesseparatelyfor the exponential, Weibull, lognormal, log-logistic, and generalized gammamodels. The same three covariates (KPS, INDADE, and INDSMA )are selected (at the 0.05 level )as the most significantcovariatesin every of these parametricmodels.ThelastcolumnofTable11.8givesthe rvaluesoftheBIC for the different parametric models with the same three covariates. Based onthesevalues,thelog-logisticmodelwiththethreecovariatesshouldbeselectedasthefinalmodelforthedatasinceitsBICorAICvalueisthelargestamongallthemodels.However,itisnotknownifthelog-logisticmodelissignificantlybetterthantheothermodels. 11.9.4 Cox--Snell Residual Procedure with Covariates TheAFTmodelsinSections11.2to11.7assumethefollowinglinearrelation- shipbetweenlog Tandthepcovariates: logT/p71/p58a/p15/p59/p78/p26 /p73/p14/p16a/p73x/p73/p71/p59/afii9846/afii9830/p71/p58/afii9839/p71/p59/afii9846/afii9830/p71(11.9.3) where /afii9830/p71has survivalfunction G(/afii9830). Oncea specifiedparametricmodeland a subsetofcovariatesareselected,toassessthegoodnessoffitofthismodel,oneapproachistocomputetheregressionresiduals /afii9830/p24/p71/p58logt/p71/p57/afii9839/p24/p71/afii9846/p24 i/p581,2,...,n (11.9.4) where /afii9839/p24/p71/p58a/p24/p15/p59/p78/p26 /p73/p14/p16a/p24/p73x/p73/p71 anda/p24/p15,a/p24/p16,a/p24/p17,...,a/p24/p78and/afii9846/p24aretheMLEof a/p15,a/p16,a/p17,...,a/p78and/afii9846,respectively, andt/p71’s are observed survival times. An /afii9830/p24/p71is taken as censored if the corresponding t/p71is censored. If the model fitted is correct, the corresponding survivalfunction G(/afii9830) isthesurvivalfunctionofthefittedmodel.Forexample, ifTindeedfollowsthe log-logisticregressionmodel with a selectedsubset of covariates, the corresponding /afii9830/p24/p71’s should followthe log-logistic distribution. Moreover, if the fitted model is correct, the Cox —Snell residuals defined in (8.4.1 )are r/p71/p58/p57logG(/afii9830/p24/p71;d/p19)/p58/p57logG/p1logt/p71/p57/afii9839/p24/p71/afii9846/p24;d/p19/p2i/p581,2,...,n(11.9.5 )290         Figure 11.2 Cox—Snellresidualsplotfromthefittedexponentialmodelonlungcancer data. whered/p19istheMLEoftheparametersofthedistribution.Let S/p19(r) denotethe estimated survival function of r/p71’s. From Section 8.4, the graph of r/p71versus /p57logS/p19(r/p71),i/p581, 2,...,n, should be closed to a straight line with unit slope and zero intercept if the fitted model for the survival time Tis correct. This graphicalmethodcan be used to assessthe goodness of fit of the parametric regressionmodel. Example 11.8 Figures 11.2 to 11.6 showthe Cox —Snell residuals plots from fitting the exponential, Weibull, lognormal, log-logistic, and extendedgeneralized gamma models, respectively with the three covariates KPS, IN-DADE, and INDSMA, to the lung cancer data in Example 11.6. The fivegraphslooksimilar,andallareclosetoastraightlinewithunitslopeandzerointercept. No significant differences are observed in these graphs. The resultsobtained are similar to those from Examples 11.6 and 11.7. The differencesamongthe fivedistributionsaresmall withthelog-logisticdistributionbeingslightlybetterthantheothers. Using the same data file ‘‘C: /p33LCANCER.DAT’’ as in Example 11.6.1, the following SAS code can be used to obtain the Cox —Snell residuals based on the exponential, Weibull, lognormal, log-logistic, and generalized gammamodelwiththethreecovariates,KPS,INDADE,andINDSMA.   291 Figure 11.3 Cox—Snell residuals plot from the fitted Weibull model on lung cancer data. Figure 11.4 Cox—Snellresidualsplotfromthefittedlognormalmodelonlungcancer data.292         Figure 11.5 Cox—Snellresidualsplotfromthefittedlog-logisticmodelonlungcancer data. dataw1; infile‘c: /p33lcancer.dat’missover; inputtcenskpsagediagtimeindpriindtheindadeindsmaindsqu;ifindpri /p580; run; procliferegnoprint; a:modelt*cens (0)/p58kpsindadeindsma/d /p58exponential; outputout /p58wacdf /p58f; b:modelt*cens (0)/p58kpsindadeindsma/d /p58weibull; outputout /p58wbcdf /p58f; c:modelt*cens (0)/p58kpsindadeindsma/d /p58lnormal; outputout /p58wccdf /p58f; d:modelt*cens (0)/p58kpsindadeindsma/d /p58gamma; outputout /p58wdcdf /p58f; e:modelt*cens (0)/p58kpsindadeindsma/d /p58llogistic; outputout /p58wecdf /p58f; run; datawa; setwa;model /p58‘Exponential’; datawb;   293 Figure 11.6 Cox—Snell residuals plot from the fitted extended generalized gamma modelonlungcancerdata. setwb; model /p58‘Weibull’; datawc; setwc;model /p58‘LNnormal’; datawd; setwd;model /p58‘Gamma’; datawe; setwe;model /p58‘LLogistic’; dataw2; setwawbwcwdwe;rcs/p58/p57log(1/p57f); run;procsort; bymodel; run; proclifetestnotableouts /p58wsnoprint; timercs*cens (0); bymodel;294         run; dataws; setws; mls/p58/p57log(survival ); run;title‘Cox-SnellResiduals (rcs)and-log (estimatedsurvivalfunctionofrcs )(mls)’; procprintdata /p58ws; varmodelrcsmls; run; Bibliographical Remarks Anexcellentexpositorypaperonstatisticalmethodsfortheidentificationand use of prognostic factors has been written by Armitage and Gehan (1974 ). Manystudiesofprognosticfactorshavebeenpublished,includingSirottetal.(1993 ), Brancato et al. (1997 ), Linka et al. (1998 ), and Lassarre (2001 ). The accelerated failure time (AFT )model was introduced by Cox (1972 ). The detailedstatisticalinferenceoftheAFTmodelsandthetheoreticalaspectsofmodel-selectingmethodsareincludedintheworkscitedinthebibliographicalremarks at the end of Chapter 9 and in the papers and books cited in thischapter. EXERCISES 11.1Consider the data given in Exercise Table 3.1. In addition to the five skintests,ageandgendermayalsohaveprognosticvalue.Examinetherelationshipbetweensurvivalandeachofthesevenpossibleprognosticvariables as in Table 3.12. For each variable, group the patientsaccording to different cutoff points. Estimate and drawthe survivalfunctionforeachsubgroupbytheproduct-limitmethodandthenusethemethodsdiscussedinChapter5 tocomparesurvivaldistributionsofthesubgroups.PrepareatablesimilartoTable3.12.Interpretyourresults. Is there a subgroup of any variable that shows significantlylongersurvivaltimes? (Fortheskintestresults,usethelargerdiameter ofthetwo. ) 11.2Consider the seven variables in Exercise 11.1. Use the Weibull re- gressionmodeltoidentifythemostsignificantvariables.CompareyourresultswiththatobtainedinExercise11.1. 11.3ConsiderthedatagiveninExerciseTable3.3.Examinetherelationship between remission duration and survival time and each of the ninepossibleprognosticvariables:age,gender,familyhistoryofmelanoma,andthesixskintests.Groupthepatientsaccordingtodifferentcutoff 295 points. Estimate and drawremission and survival curves for each subgroup. Compare the remission and survival distributions of sub-groups using the methods discussed in Chapter 5. Prepare tablessimilartoTable3.8. 11.4Perform the following analyses: (1)Use the exponential, Weibull, lognormal,generalizedgamma,orlog-logisticregressionmodelssepa-ratelytoidentifythesignificantvariablesinExercise11.3fortheirrelative importancetoremissiondurationandsurvivaltime. (2)Selectamodel amongthesefinalmodelsusingtheBICorAICmethod. (3)Calculate separatelytherespectivelikelihoodfortheexponential,Weibull,lognor-mal,orgeneralizedgammaregressionmodelwiththefixedvariablesinthemodelselectedinstep (2),thenusethemethodinSection11.9.2tochoose amodelandseewhetherthismodelisthemodelselectedinstep (2). 11.5Performthesameanalysesas inExercise11.4for survivaltimein the 157diabeticpatientsgiveninExerciseTable3.4. 11.6Using the notations in Example 11.2, show that if we use the follow- ing model to replace the model defined in (11.4.6 ), logT/p71/p58a/p15/p59a/p16LOW/p9/p59a/p17UNSA/p9/p59/afii9846/afii9830/p9/p58/afii9839/p9/p59/afii9846/afii9830/p9, the hypothesis H/p15:h/p81/p58h/p83is equivalenttoH/p15:a/p17/p580. 11.7Following Examples 11.2 and 11.3, obtain the log-likelihood function basedon (11.6.8 )forthe137observedexactandright-censoredsurvival timesfromthelungcancerpatients. 11.8Using the same notation as in Example 11.5, showthat if w e use the model log T/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDLAR/p71/p59a/p18INDSMA/p71/p59a/p19INDSQU/p71/p59/afii9846/afii9830/p71,whereINDLAR /p581ifthetypeofcancerislarge, and0otherwise,toreplacethemodeldefinedin (11.7.8 ),thehypotheses H/p15:a/p18/p580 andH/p15:a/p19/p580 are equivalent to H/p15:OR/p49/p43/p58OR/p31/p34and H/p15:OR/p49/p47/p58OR/p31/p34, respectively. Moreover, if we use the model logT/p71/p58a/p15/p59a/p16KPS/p71/p59a/p17INDLAR/p71/p59a/p18INDADE/p71/p59a/p19INDSQU/p71/p59/afii9846/afii9830/p71toreplacethemodeldefinedin (11.7.8 ),thehypothesis H/p15:a/p19/p580 isequivalentto H/p15:OR/p49/p47/p58OR/p49/p43. 11.9Let/afii9830beasurvivaltimewiththedensityfunction g(/afii9830)/p58exp[/afii9830/p57exp(/afii9830)]. Showthat the survival time Tdefined by log T/p58/afii9839/p59/afii9846/afii9830has the Weibulldistributionwith /afii9838/p58exp(/p57/afii9839//afii9846)and/afii9828/p581//afii9846byapplyingthe densitytransformationrulein (11.2.11 ). 11.10Let/afii9830beasurvivaltimewiththestandardnormaldistribution N(0,1). Showthat the survival time Tdefined by log T/p58/afii9839/p59/afii9846/afii9830has the lognormaldistributionbyapplyingthe densitytransformationrule in(11.2.11 ).296         11.11Letubeasurvivaltimewiththedensityfunction f(u), f(u)/p58exp[u//afii9829/p17/p57exp(u)] /afii9772(1//afii9829/p17) where /afii9772(·)isthegammafunctiondefinedin (6.2.8 ). (a)Showthat the survival time /afii9830defined by /afii9830/p58/afii9839//afii9829/p59(log/afii9829/p17)//afii9829has thefollowingdensityfunction, g(/afii9830)/p58/p34/afii9829/p34[exp (/afii9829/afii9830)//afii9829/p17]/p16/p30/p66/p130exp[/p57exp(/afii9829/afii9830)//afii9829/p17] /afii9772(1//afii9829/p17)/p57/p45/p58/afii9830/p58 /p59/p45 andsurvivalfunction, G(/afii9830)/p58/p7I/p1exp(/afii9829/afii9830) /afii9829/p17,1 /afii9829/p17/p2if/afii9829/p580 1/p57I/p1exp(/afii9829/afii9830) /afii9829/p17,1 /afii9829/p17/p2if/afii9829/p570 whereI(·,·)istheincompletegammafunctiondefinedasin (6.4.4 ). (b)Showthat the survival time Tdefined by log T/p58/afii9839/p59/afii9829/afii9830has the extendedgammadensityfunctiondefinedin (11.6.4 ). 11.12If/afii9830hasalogisticdistributionwiththedensityfunction g(/afii9830)/p58exp(/afii9830) [1/p59exp(/afii9830)]/p17 showthatthesurvivaltime TdefinedbylogT/p58/afii9839/p59/afii9846/afii9830hasthelog-logistic distribution with /afii9825/p58exp(/p57/afii9839//afii9846)and/afii9828/p581//afii9846by applying the density transformationrulein (11.2.11 ). 297 CHAPTER 12 Identification of Prognostic Factors Related to Survival Time:Cox Proportional Hazards Model In Chapter 11 we discussed parametric survival methods for model fitting andforidentifyingsignificant prognosticfactors. Thesemethodsare powerfulif theunderlying survival distribution is known. The estimation and hypothesistesting of parameters in the models can be conducted by applying standardasymptotic likelihood techniques. However, in practice, the exact form of theunderlying survival distribution is usually unknown and we may not be ableto find an appropriate model. Therefore, the use of parametric methods inidentifying significant prognostic factors is somewhat limited. In this chapterwediscussamostcommonlyused model,theCox (1972 )proportionalhazards model, and its related statistical inference. This model does not requireknowledge of the underlying distribution. The hazard function in this modelcan take on any form, including that of a stepfunction, but the hazardfunctions of different individuals are assumed to be proportional and indepen-dentof time.Theusual likelihoodfunctionis replacedby thepartial likelihoodfunction.Theimportantfactisthatthestatisticalinferencebasedonthepartiallikelihood function is similar to that based on the likelihood function. 12.1 PARTIALLIKELIHOODFUNCTIONFORSURVIVALTIMES The Cox proportional hazards model possesses the property that different individuals have hazard functions that are proportional, i.e., [ h(t/p34x/p16)/h(t/p34x/p17)], the ratio of the hazardfunctions for two individuals with prognostic factors orcovariatesx/p16/p58(x/p16/p16,x/p17/p16,...,x/p78/p16)/p30, andx/p17/p58(x/p16/p17,x/p17/p17,...,x/p78/p17)/p30is a constant (does not vary with time t). This means that the ratio of the risk of dying of two individuals is the same no matter how long they survive. In Sections 11.3 298 and 11.4, we showed that the exponential and Weibull regression models possess this property. This property implies that the hazard function given aset of covariates x/p58(x/p16,x/p17,...,x/p78)/p30can be written as a function of an underlying hazard function and a function, say g(x/p16,...,x/p78), of only the covariates, that is, h(t/p34x/p16,...,x/p78)/p58h/p15(t)g(x/p16,...,x/p78)o rh(t/p34x)/p58h/p15(t)g(x)(12.1.1 ) The underlying hazard function, h/p15(t), represents how the risk changes with time,andg(x)representsthe effect of covariates. h/p15(t) can be interpreted as the hazard function when all covariates are ignored or when g(x)/p581, and is also called thebaseline hazard function . The hazard ratio of two individuals with different covariates x/p16andx/p17is h(t/p34x/p16) h(t/p34x/p17)/p58h/p15(t)g(x/p16) h/p15(t)g(x/p17)/p58g(x/p16) g(x/p17)(12.1.2) which is a constant, independent of time. The Cox (1972 )proportional hazard model assumes that g(x)in(12.1.1 )is an exponential function of the covariates, that is, g(x)/p58exp/p1/p78/p26 /p72/p14/p16b/p72x/p72/p2/p58exp(b/p30x) and the hazard function is h(t/p34x)/p58h/p15(t) exp /p1/p78/p26 /p72/p14/p16b/p72x/p72/p2/p58h/p15(t) exp (b/p30x)( 12.1.3 ) whereb/p58(b/p16,...,b/p78)denotes the coefficients of covariates. These coefficients can be estimated from the data observed and indicate the magnitude of theeffects of their corresponding covariates. For example, if there is only onecovariate treatment, let x/p16/p580 if a person receives placebo and x/p16/p581i fa personreceivestheexperimentaldrug.Thehazardratioofthepatientreceivingthe experimental drug and the one receiving placebo based on (12.1.2 )and (12.1.3 )is h(t/p34x/p16/p581) h(t/p34x/p16/p580) /p58exp(b/p16) Thus, the two treatments are equally effective if b/p16/p580 and the experimental drugintroduceslower (higher )riskforsurvivalthanplaceboif b/p16/p580(b/p16/p570). It can be shown that (12.1.3 )is equivalent to S(t/p34x)/p58[S/p15(t)]exp(/afii9814/p78/p72/p14/p16b/p72x/p72)/p58[S/p15(t)]exp(b/p30x)(12.1.4 )      299 Thus the covariates can be incorporated into the survivorship function. The use of (12.1.3 )can be exemplified as follows. 1.Two-sample problems. Suppose that p/p581; that is, there is only one covariate,x/p16, which is an indicator variable: x/p16/p71/p58/p70 if theith individual is from group0 1 if theith individual is from group1 Then according to (12.1.3 ), the hazard functions of groups 0 and 1 are, respectively, h/p15(t) andh/p16(t)/p58h/p15(t) exp(b/p16). Thehazardfunctionofgroup 1 is equal to the hazard function of group0 multip lied by a constantexp(b/p16), or the two hazard functions are proportional. In terms of the survivorshipfunction, S(t)/p58[S/p15(t)]/p65 wheretheconstant c/p58exp(b/p16)(Nadas,1970 ).Thetwo-sampletestdevelop- ed from (12.1.3 )is the Cox—Manteltest discussedin Chapter 5. It is now apparent that the test is based on the assumption of a proportionalhazard between the two groups. 2.Two-sampleproblemswithcovariates. The covariates in (12.1.3 )can either be indicator variables such as x/p16in the two-groupp roblem above or prognosticfactors. Havingone or more covariates representing prognos-tic factors in (12.1.1 )enables us to examine the relation between two groups, adjusting for the presence of prognostic factors. 3.Regression problems. Dividing both sides of (12.1.3 )byh/p15(t) and taking its logarithm, we obtain logh/p71(t) h/p15(t) /p58b/p16x/p16/p71/p59b/p17x/p17/p71/p59/p37/p59b/p78x/p78/p71/p58/p78/p26 /p72/p14/p16b/p72x/p72/p71/p58b/p30x/p71(12.1.5 ) wherethex’sarecovariatesforthe ithindividual.Theleftsideof (12.1.5 )is a function of hazard ratio (or relative risk )and the right side is a linear function of the covariates and their respective coefficients. As mentioned earlier, h/p15(t) is the hazard function when all covariates are ignored.If thecovariates are standardizedaboutthe meanand the modelusedis logh/p71(t) h/p15(t)/p58b/p16(x/p16/p71/p57x/p21/p16)/p59b/p17(x/p17/p71/p57x/p21/p17)/p59/p37/p59b/p78(x/p78/p71/p57x/p21/p78)/p58b/p30(x/p71/p57x/p21) (12.1.6 )300         wherex/p21/p30/p58(x/p21/p16,x/p21/p17,...,x/p21/p78)andx/p21/p72is the average of the jth covariate for all patients, the left side of (12.1.6 )is the logarithm of the ratio of risk of failure for a patient with a given set of values x/p30/p71/p58(x/p16/p71,x/p17/p71,...,x/p78/p71)to that for an average patient who has an average value for every covariate. In this chapter we focus on the use of (12.1.5 ), and the main interest here is to identify important prognostic factors. In other words, we wish to identifyfrom thepcovariates a subset of variables that affect the hazard more significantly, and consequently, the length of survival of the patient. We are concerned with the regression coefficients. If b/p71is zero, the corresponding covariateis notrelatedto survival.If b/p71is notzero,it representsthe magnitude of the effect of x/p71on hazard when the other covariates are considered simultaneously. To estimate the coefficients, b/p16,...,b/p78, Cox (1972 )proposes a partial likelihoodfunctionbasedon aconditionalprobabilityoffailure,assumingthattherearenotiedvaluesinthesurvivaltimes.However,inpractice,tiedsurvivaltimes are commonly observed and Cox’s partial likelihood function wasmodified to handle ties (Kalbfleisch and Prentice, 1980; Breslow, 1974; Efron, 1977 ). In the following we describe the estimation procedure without and with ties. 12.1.1 EstimationProcedureswithoutTiedSurvivalTimes Suppose that kof the survival times from nindividuals are uncensored and distinct, and n/p57kare right-censored. Let t/p7/p16/p8/p58t/p7/p17/p8/p58/p37/p58t/p7/p73/p8be the ordered kdistinct failure times with corresponding covariates x/p7/p16/p8,x/p7/p17/p8,...,x/p7/p73/p8. Let R(t/p7/p71/p8) be the risk set at time t/p7/p71/p8.R(t/p7/p71/p8)consists of all persons whose survival times are at least t/p7/p71/p8. For the particular failure at time t/p7/p71/p8, conditionallyon the risksetR(t/p7/p71/p8),theprobabilitythatthefailureisontheindividualasobservedis exp/p37 /p26/p78/p72/p14/p16b/p72x/p72/p7/p71/p8/p38 /p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74) /p1/p58exp(b/p30x/p7/p71/p8) /p26l/p43R(t/p7/p71/p8)exp(b/p30x/p74)/p2 Each failure contributes a factor and hence the partial likelihood function is L(b)/p58/p73/p147 /p71/p14/p16exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p7/p71/p8/p38 /p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74) /p1/p58/p73/p147 /p71/p14/p16exp(b/p30x/p7/p71/p8) /p26l/p43R(t/p7/p71/p8)exp(b/p30x/p74)/p2(12.1.7 ) and the log-partial likelihood is l(b)/p58logL(b)/p58/p73/p26 /p71/p14/p16/p78/p26 /p72/p14/p16b/p72x/p72/p71/p57/p73/p26 /p71/p14/p16log/p3/p26 l/p43R(t/p7/p71/p8)exp/p1/p78/p26 /p72/p14/p16b/p72x/p72/p74/p2/p4 /p58/p73/p26 /p71/p14/p16/p7b/p30x/p7/p71/p8/p57log/p3/p26 l/p43R(t/p7/p71/p8)exp(b/p30x/p74)/p4/p8(12.1.8 )      301 The maximum partial likelihood estimator (MPLE )b/p19ofbcan be obtained by the steps shown in (7.1.2 )—(7.1.4 ). That is,b/p19/p16,...,b/p19/p78are obtained by solving the following simultaneous equations: /p42(l(b)) /p42b/p580 or /p42l(b) /p42b/p83/p58/p73/p26 /p71/p14/p16[x/p83/p7/p71/p8/p57A/p83/p71(b)]/p580u/p581, 2,...,p(12.1.9) where A/p83/p71(b)/p58/p26l/p43R(t/p7/p71/p8)x/p83/p74exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74) /p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74/p38/p58/p26l/p43R(t/p7/p71/p8)x/p83/p74exp(b/p30x/p74) /p26l/p43R(t/p7/p71/p8)exp(b/p30x/p74)(12.1.10 ) by applying the Newton —Raphson iterated procedure. The second partial derivatives of l(b)with respective to b/p83andb/p84,u,v/p581, 2,...,p, in the Newton—Raphson iterative procedure are I/p83/p84(b)/p58/p42/p17l(b) /p42b/p83/p42b/p84/p58/p57/p73/p26 /p71/p14/p16C/p7/p83/p84/p71/p8(b/p16,...,b/p78)/p58/p57/p73/p26 /p71/p14/p16C/p7/p83/p84/p71/p8(b)u,v/p581, 2,...,p (12.1.11) where C/p7/p83/p84/p71/p8(b)/p58/p26l/p43R(t/p7/p71/p8)x/p83/p74x/p84/p74exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74) /p26l/p43R(t/p7/p71/p8)exp/p37/p26/p78/p72/p14/p16b/p72x/p72/p74)/p57A/p83/p71(b)A/p84/p71(b) (12.1.12) The covariance matrix of the MPLE b/p19, defined similarly as V/p19(b) defined in (7.1.5 ),i s V/p19(b/p19)/p58Cov/p19(b/p19)/p58/p3/p57/p42/p17l(b/p19) /p42b/p42b/p30/p4/p92/p16(12.1.13 ) where /p57/p42/p17l(b/p19)//p42b/p42b/p30is called the observedinformationmatrix with /p57I/p83/p84(b/p19)as its(u,v)element and where I/p83/p84(b)is defined in (12.1.11 ). Let the (i,j) element ofV/p19(b/p19)in(12.1.13 )bev/p71/p72; then the 100 (1/p57/afii9825)% confidence interval for b/p71is, according to (7.1.6 ), /p37b/p19/p71/p57Z/p63/p30/p17/p40v/p71/p71,b/p19/p71/p59Z/p63/p30/p17/p40v/p71/p71/p38 (12.1.14) 12.1.2 EstimationProcedurewithTiedSurvivalTimes Suppose that among the nobserved survival times there are kdistinct uncensored times t/p7/p16/p8/p58t/p7/p17/p8/p58/p37/p58t/p7/p73/p8. Letm/p7/p71/p8denote the number of people302         who fail att/p7/p71/p8or the multiplicity of t/p7/p71/p8;m/p7/p71/p8/p571 if there are more than one observation with value t/p7/p71/p8;m/p7/p71/p8/p581 if there is only one observation with value t/p7/p71/p8. LetR(t/p7/p71/p8)denote the set of people at risk at time t/p7/p71/p8[i.e.,R(t/p7/p71/p8)consists of those whose survival times are at least t/p7/p71/p8] andr/p71be the number of such persons.Forexample,in thefollowingset ofsurvivaltimes fromeightsubjects,/p4315, 16 /p59, 20, 20, 20, 21, 24, 24 /p44,n/p588,k/p584,t/p7/p16/p8/p5815,t/p7/p17/p8/p5820,t/p7/p18/p8/p5821, t/p7/p19/p8/p5824,m/p7/p16/p8/p581,m/p7/p17/p8/p583,m/p7/p18/p8/p581, andm/p7/p19/p8/p582. ThenR(t/p7/p16/p8) includes all eight subjects. R(t/p7/p17/p8)/p58/p43those subjects with survival times 20, 21, and 24 /p44, R(t/p7/p18/p8)/p58/p43those subjects with survival times 21 and 24 /p44, andR(t/p7/p19/p8)/p58/p43those subjects with survival time 24 /p44; thus,r/p16/p588,r/p17/p586,r/p18/p583, andr/p19/p582. To discuss the methods for ties, we introduce a few additional notations. From everyR(t/p7/p71/p8), we can randomly select m/p7/p71/p8subjects. Donate each of these m/p7/p71/p8selections byu/p7/p72/p8. There are/p80/p71C/p75 /p7/p71/p8/p58r/p71!/[m/p7/p71/p8!(r/p71/p57m/p7/p71/p8)!] possibleu/p7/p72/p8’s. Let U/p71denote the set that contains all the u/p7/p72/p8’s. For example, from R(t/p7/p17/p8), we can randomlyselectany m/p7/p17/p8/p583 out of the 6 (r/p17/p586) subjects. There are a total of /p21C/p18/p5820 such selections (or subsets ), and one ofu/p7/p72/p8’s is, for example, /p43three subjects with survival times 20, 20, and 24 /p44.U/p17/p58/p43u/p7/p16/p8,u/p7/p17/p8,...,u/p7/p17/p15/p8/p44contains all 20 subsets. Now let us focus on the tied observations. Letx/p73/p58(x/p16/p73,x/p17/p73,...,x/p78/p73)/p30denote the covariates of the kth individual, z u(j)/p58/p26k/p43u/p7/p72/p8x/p73/p58(z1u(j),z2u(j),...,zpu(j))/p30,wherezlu(j)isthesumof the lth covariate of them/p7/p71/p8persons who are in u/p7/p72/p8. Letu*/p7/p71/p8denotes the set of m/p7/p71/p8people who failed at time t/p7/p71/p8, andzu*(i)/p58/p26k/p43u*/p7/p71/p8x/p73/p58(z*1u*(i),z*2u*(i),...,z*pu*(i))/p30, wherez*lu*(i)be the sum of the lth covariate of the m/p7/p71/p8persons who are in u*/p7/p71/p8(failed at time t/p7/p71/p8). For example, for the set of survival times above, z*1u*(2)equals the sum of the first covariate values of three persons who failed at time 20. With thesenotations we are ready to introduce the following method for ties. Continuous Time Scale In the case of a continuous time scale, for the m/p7/p71/p8persons failing at t/p7/p71/p8,i ti s reasonable to say that the survival times of the m/p7/p71/p8people are not identical sincethe ties are most likelyto be the resultsof imprecisemeasurements.If theprecisemeasurementscouldbemade,these m/p7/p71/p8survivaltimescouldbeordered and we could use the likelihood function in (12.1.7 ). In the absence of knowledge of the true order (the real case ), we have to consider all possible orders of these observed m/p7/p71/p8tied survival times. For each t/p7/p71/p8, the observed m/p7/p71/p8tied survival time can be ordered in m/p7/p71/p8!(m/p7/p71/p8factorial )different possible ways. For each of these possible orders we will have a product as in (12.1.7 )for the corresponding m/p7/p71/p8survival times. Therefore, when the survival time is meas- ured at a continuous time scale, construction and computation of the exactpartial likelihood function is a very tedious task if m/p7/p71/p8is larger. Readers interested in the details of the exact partial likelihood function are referred toKalbfleischandPrentice (1980 )andDelongetal. (1994 ).Theformulaprovided by Delong et al. makes computation of the partial likelihood function for tiedcontinuous survival times more feasible. We will not discuss the exact partiallikelihood function further due to its complexity. Among the statistical sof-      303 twarepackages,SASincludesaprocedurebasedontheexactpartiallikelihood function. Use of this procedure is illustrated in Example 12.3. To approximate the exact partial likelihood function, the following two likelihood functions can be used when each m/p7/p71/p8is small compared to r/p71. Breslow (1974 )provided the following approximation: L/p32(b)/p58/p73/p147 /p71/p14/p16exp(z/p30u*(i)b) [/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b)]m/p7/p71/p8(12.1.15 ) An alternative approximation was provided by Efron (1977 ): L/p35(b)/p58/p73/p147 /p71/p14/p16exp(z/p30u*(i)b) /p147/p75/p7/p71/p8/p72/p14/p16[/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b)/p57[(j/p571)/m/p7/p71/p8]/p26l/p43u*/p7/p71/p8exp(x/p30/p74b)](12.1.16 ) Discrete Time Scale If survival times are observed at discrete times, the tied observations are trueties: that is, these events really happen at the same time. Cox (1972 )proposed the following logistic model: h/p71(t)dt 1/p57h/p71(t)dt/p58h/p15(t)dt 1/p57h/p15(t)dtexp/p1/p78/p26 /p72/p14/p16b/p72x/p72/p71/p2/p58h/p15(t)dt 1/p57h/p15(t)dtexp(b/p30x/p71) This model reduces to (12.1.3 )in the continuous time scale. Using the model and replacing the ith term in (12.1.7 )with the following term with tied observations at t/p7/p71/p8: exp(z/p30u*(i)b) /p26u/p7/p72/p8/p43U/p71exp(z/p30u(j)b) the partial likelihoodfunction with tied observations at a discrete time scale is L/p66(d)/p58/p73/p147 /p71/p14/p16exp(z/p30u*(i)b) /p26u/p7/p72/p8/p43U/p71exp(z/p30/p83/p7/p72/p8b)(12.1.17 ) Theith term in this expression represents the conditional probability of observing the m/p7/p71/p8failures given that there are m/p7/p71/p8failures at time t/p7/p71/p8and the risk setR(t/p7/p71/p8)att/p7/p71/p8. The number of terms in the denominator of the ith term is/p80/p71C/p75/p7/p71/p8/p58r/p71!/[m/p7/p71/p8!(r/p71/p57m/p71)!], as noted earlier, and will be very large if the m/p7/p71/p8is large. Fortunately, a recursive algorithm proposed by Gail et al. (1981 ) makes the calculation manageable. Equation (12.1.17 )can also be considered as an approximation of the partial likelihood function for continuous survivaltimes with ties by assuming that the ties are true as if they were observed at adiscrete time scale.304         Asshowninmanypapersinliterature,inmostpracticalsituations,thethree partial likelihood functions above are reasonably good approximations of theexact partial likelihood function for continuous survival time with ties. Whenthere are no ties on the event times (i.e.,m/p7/p71/p8/p891), (12.1.15)—(12.1.17 )reduce to (12.1.7 ). The maximum partial likelihood estimate of bin(12.1.15 )—(12.1.17 ) can be estimated using procedures similar to those in (12.1.8 )—(12.1.14 ). Once the coefficients are estimated, relative risk (or relative hazard )in (12.1.2 )or(12.1.5 )can be obtained.For example, if x/p16represents hypertension and is defined as x/p16/p58 /p71 if patient is hypertensive 0 otherwise thehazardrateforhypertensivepatientsis exp (b/p19/p16)timesthat fornormotensive patients. That is, the risk associated with hypertension is exp (b/p19/p16)adjusting for the other covariates in the model. A 100 (1/p57/afii9825)% confidence interval for the relative risk can be obtained by using the confidence interval for b/p16. Let (b/p16/p42,b/p16/p51)be the 100 (1/p57/afii9825)% confidence interval for b/p16; a 100 (1/p57/afii9825)% confidence interval for the relative risk is (exp(b/p16/p42),exp(b/p16/p51))according to (7.1.8 ). This application of the proportional hazards model has been used extensively, particularly by epidemiologists. The following three examples illustrate the use of Cox’s regression model. Example 12.1 Consider the survival data from 30 patients with AML in Table 11.4. Recall that the two possible prognostic factors are x/p16/p58/p71 if patient is /p4650 years 0 otherwise x/p17/p58/p71 if cellularity of marrow clot section is 100% 0 otherwise We fit the Cox proportional hazard model to the data. The results are presented in Table 12.1 In this case, Breslow’s approximation in (12.1.15 )is usedtohandleties.Thepositivesignsoftheregressioncoefficientsindicatethatthe older patients (/p4650 years )and patients with 100% cellularity of the marrow clot section have a higher risk of dying. Furthermore, age is signifi-cantly related to survival after adjustment for cellularity. The results areconsistent with those from fitting the lognormal regression model in Example11.3. The coefficients of the binary covariates can be interpreted in terms ofrelative risk. The estimated risk of dying for patients at least 50 years of age is2.75 times higher than that for patients younger than 50. Patients with 100%cellularity have a 42%higher risk of dying than patients with less than 100%cellularity.      305 Table12.1 ResultsofaProportionalHazardsRegressionAnalysisofDatain Table11.4 Regression Standard Covariate Coefficient Error pValue exp (coefficient ) x/p16(age) 1.01 0.46 0013 2.75 x/p17(cellularity ) 0.35 0.44 0.212 1.42 The95%confidenceintervalsfor b/p16(age)andb/p17(cellularity )are1.01 /p601.96 (0.46)or(0.11, 1.91 )and 0.35 /p601.96 (0.44)or(/p570.51, 1.21 ), respectively. Consequently, the 95% confidence intervals for the relative risks are(e/p15/p13/p16/p16,e/p16/p13/p24/p16)or(1.12, 6.75 )and (e/p92/p15/p13/p20/p16,e/p16/p13/p17/p16)or(0.60, 3.35 ), respectively. The small number of patients (30)may have contributed to the large standard errors ofb/p19/p16andb/p19/p17and consequently, the wide confidence intervals. The lower bound of the confidence interval for age is only slightly above 1. This suggeststhat the importance of age should be interpreted carefully. In general, if thenumber of subjects is small and the standard errors of the estimates are large,the estimates may be unreliable. When the two covariates are considered simultaneously, the risk for a patientwith x/p16/p581 andx/p17/p581 relative to patientswith x/p16/p580 andx/p17/p580 can be estimated. The relative risk is estimated as exp (1.01/p590.35)/p583.90 for a patient who is over 50 years of age and whose cellularity is 100%, comparedto patients who are younger than 50 and whose cellularity is less than 100%. Using the same data set ‘‘C: /p33AML.DAT’’ defined in Example 11.3, the following SAS code can be used to obtain the results in Table 12.1. data w1; infile ‘c: /p33aml.dat’ missover; input t cens x1 x2; run;proc phreg; model t*cens (0)/p58x1 x2 / rl; run; If BMDP 2L is used, the following code is applicable. /input file /p58‘c:/p33aml.dat’ . variables /p584. format /p58free. /print cova./variable names /p58t, cens, x1, x2. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58x1, x2.306         If SPSS is used, the following code suffices. data list file /p58‘c:/p33aml.dat’ free / t cens x1 x2. coxreg t with x1 x2 /status /p58cens event (1) /print /p58all. Example 12.2 In a study (Buzdar et al., 1978 )to evaluate a combination of 5-flourouracil, adramycin, cyclophosphamide, and BCG (FAC-BCG )as adjuvant treatment in stage II and III breast cancer patients with positiveaxillary nodes, 131 patients receiving FAC-BCG after surgery and radiationtherapy were compared with 151 patients receiving surgery and radiationtherapy only (control group ). Cox’s regression model was used to identify prognostic factors and to evaluate the comparability of the two treatment groups. The model was fittedto the data from 151 patients to determine the variables related to length ofremission. The possible prognostic variables considered were age (years ), menopausal status (1, premenopausal; 0, other ), size of primary tumor (2,/p583 cm; 4, 3—5 cm; 7, /p575c m ), state of disease (2, stage II; 3, stage III ), location of surgery (1, M. D. Anderson Hospital; 0, other ), number of nodes involved (2, /p584; 7, 4—10; 12, /p5710), and race (1, Caucasian; 2, other ). The covariates were selected by the forward selection method outlined in Section 11.9. Threevariables—number of nodes involved, state of disease, and menopausalstatus—were selected for use in the model,all related significantly (p/p580.1)to disease-free time. The regression equation including these variables only is logh/p71(t) h/p15(t) /p580.111(number of nodes /p576.16)/p590.8122 (stage /p572.39) /p590.872 (menopausal /p570.26) Table 12.2 gives the details of the fit. Relative risk was taken as h/p71(t)/h/p15(t), the ratio of the risk of death per unit of time for a patient with a given set ofprognostic variables to the risk for a patient whose prognostic variables wereaverage in value. The relative risk for each variable was calculated byconsidering favorable or unfavorable values of that variable, assuming thatother variables were at their average value. Note that the risk of relapse perunit time for a patient with 12 positive nodes is 3.04 (ratio or risk )times that fora patientwith only twopositive nodes.The riskof relapseper unittime fora stage III patient was 2.25 times that of a stage II patient. The Cox’s regression model was also fitted to the combined groupof FAC-BCG and control patients, including type of treatment (0, control; 1, FAC-BCG ),menopausalstatus,sizeofprimarytumor,andnumberofinvolved nodes as potential prognostic variables. The regression equation with three      307 Table12.2 PatientCharacteristicsRelatedtoDisease-FreeTimeinCox’sRegression ModelFittoControlPatients Maximum Relative Risk /p63 Prognostic Regression Significance Log Ratio of Variable Coefficient Level ( p) Likelihood Favorable Unfavorable Risks Number of nodes 0.1110 /p580.01 /p57257.407 0.63 1.91 3.04 Stage 0.8122 0.016 /p57254.533 0.73 1.64 2.25 Menopausal status 0.8720 /p580.1 /p57250.576 0.80 1.91 2.39 Source:Buzdar et al. (1978 ). Reprinted by permission of the editor. /p63Favorable variables: number of nodes /p582, stage II, postmenopausal. Unfavorable variables: number of nodes /p5812, stage III, premenopausal. Table12.3 PatientCharacteristicsRelatedtoSurvival,TreatmentIncluded Maximum Relative Risks /p63 Prognostic Regression Significance Log Ratio of Variable Coefficient Level ( p) Likelihood Favorable Unfavorable Risks Treatment /p571.8792 /p580.01 /p57201.200 0.37 2.42 6.55 Menopausal 0.9644 0.01 /p57197.719 0.73 1.91 2.62 status Size of 0.1611 0.05 /p57195.865 0.72 1.61 2.24 primary tumor Source:Buzdar et al. (1978 ). Reprinted by permission of the editor. /p63Favorable variables: treatment—FAC-BCG, postmenopausal, size of primary tumor 2cm. Unfavorable variables: no adjuvant treatment, premenopausal, size of primary tumor 7cm.significant (p/p580.05)variables obtained was as follows: logh/p71(t) h/p15(t)/p58/p571.8792 (treatment /p570.47)/p590.9644 (menopausal status /p570.33) /p590.1611 (size of primary tumor /p574.04) Table12.3givesthedetailsofthefit.Themostimportantvariableinpredicting survival time was the type of treatment (FAC-BCG favorable ); other signifi- cantly important variables were menopausalstatus and size of primary tumor.Therisk ofdeathperunitof timefora patientreceivingno adjuvanttreatment(control group )was 6.55 times that for a patient receiving the treatment, showing that FAC-BCG can prolong life considerably.308         Example 12.3 Suppose that demographic, personal, clinical, and labora- tory data are collected from an interview and physical examination of 200participants in a study of cardiovascular disease (CVD ). These participants, aged 50—79 years and free of CVD at the time of the baseline examination, are then followed for 10 years. During the follow-upp eriod, 96 of the 200participants develop or die of CVD. We use this set of simulated data toillustrate further the use of the proportional hazards model in identifyingimportant risk factors. Table 12.4 gives a subset of the simulated data of 68 participants. The event time Tof interest is CVD-free time, which is defined as the time in years from baseline examination to the first time that a participant wasdiagnosed as having CVD or confirmed as a CVD death. CVD includescoronary heart disease (CHD )and stroke. The covariates of interest are age (AGE ), gender (SEX /p581 if male and /p580 if female ); smoking status (SMOKE /p581 if current smoker, and 0 otherwise ); body mass index (BMI /p58weight in kilograms divided by height in meter squared ); systolic blood pressure (SBP ); logarithm of ratio of urinary albumin and creatinine (LACR ); logarithm of triglycerides (LTG ); hypertension status (HTN /p581i f SBP/p46140 mmHg or DBP /p4690 mmHg or under treatments of hypertension, and /p580 otherwise ); and diabetes status (DM /p581 if fasting glucose /p46126 mg/dL or under the treatments of diabetes, and /p580 otherwise ). For the CVD outcome of interest, we let DG denote the type of CVD. DG /p580 if the participant is free of CVD at the end of the study or confirmed as a non-CVDdeath (thus the CVD-freetime is censored ),/p581 if the participanthad a stroke, /p582 if the participant had a CHD, and /p583 if the participant had other CVDs. ItisofinteresttocomparetheriskofCVDamongthethreeagegroups:50 —59, 60—69, and70—79. We create twodummy variables:AGEA /p581 if aged 50—69, /p580 otherwise; and AGEB /p581 if aged 60—69, and /p580 otherwise. Thus for a 70 to79-year-old,AGEA /p580andAGEB /p580.Wealsocreateavariabletodenote the censoring status: CENS /p580 if t is censored, and /p581 if uncensored. Toillustratethedifferentmethodstohandleties,wefittheCoxproportional hazards model with the following six covariates: AGEA, AGEB, SEX,SMOKE, BMI, and LACR. The approximated partial likelihood functiondefined in (12.1.15 )—(12.1.17 )as well as the exact partial likelihood function (Delong et al., 1994 )are applied. As noted in Sections 11.3 and 11.4, the exponential and Weibull regression models are also proportional hazardmodels. Therefore, for comparisons we also fit an exponential and a Weibullregression model with the same covariates to the data. The estimated re-gression coefficients obtained from the proportional hazards model withapproximated discrete, Breslow, Efron, and exact partial likelihood functionsas well as those from the exponential and Weibull regression models are givenin Table 12.5. All of the estimates based on the Cox model and an approxi-mated partial likelihood function are very closed to those based on the exactpartial likelihood. Those based on Efron’s approximation are almost identicalto those (different only at the fourth decimal place )based on the exact partial      309 Table12.4 ASubsetoftheSimulatedDataforaCardiovascularDiseaseStudyinExample12.3 /p63 ID T CENS DG AGEA AGEB SEX SMOKE BMI SBP LACR LTG AGE HTN DM 1 7.4 0 0 0 0 0 0 31.78 141 4.23 3.94 77.8 1 0 2 7.9 0 0 0 0 0 0 25.02 124 4.31 4.66 76.9 0 1 3 6.4 0 0 0 0 0 1 26.05 111 4.38 4.27 76.3 0 04 7.1 0 0 0 0 0 1 26.92 140 1.11 4.51 72.2 1 0 5 6.0 0 0 0 0 0 1 34.30 146 1.19 4.82 76.0 1 0 6 6.5 0 0 0 0 0 1 31.76 142 1.20 4.88 74.5 1 07 8.3 0 0 0 0 1 0 25.01 154 3.53 4.10 70.7 1 1 8 7.9 0 0 0 0 1 0 28.21 136 3.73 4.12 75.2 1 0 9 7.6 0 0 0 1 0 0 28.13 127 2.92 4.24 64.9 0 0 10 8.4 0 0 0 1 0 0 25.68 118 2.47 4.41 60.2 0 0 11 7.4 0 0 0 1 0 0 34.34 118 2.37 4.46 64.4 0 1 12 7.7 0 0 0 1 0 0 28.92 127 3.58 4.55 68.8 1 113 6.9 0 0 0 1 0 1 24.68 100 2.11 4.33 64.4 0 0 14 7.2 0 0 0 1 0 1 21.93 121 3.39 4.64 60.8 0 1 15 6.3 0 0 0 1 0 1 29.47 98 1.96 4.69 64.4 0 0 16 7.4 0 0 0 1 0 1 28.65 150 2.59 4.95 61.6 1 0 17 4.5 0 0 0 1 1 0 32.28 128 2.99 4.73 65.3 0 0 18 7.0 0 0 0 1 1 0 29.21 117 2.17 4.91 65.7 0 119 2.8 0 0 0 1 1 0 28.82 136 4.04 4.92 65.4 0 1 20 7.2 0 0 0 1 1 0 30.58 121 2.84 4.94 64.5 0 1 21 7.4 0 0 1 0 0 0 27.83 95 1.85 4.44 52.0 0 022 5.2 0 0 1 0 0 0 26.61 128 2.87 4.51 50.7 0 0 23 7.7 0 0 1 0 0 0 30.32 96 2.41 4.60 52.5 0 1 24 7.8 0 0 1 0 0 0 30.41 130 1.45 4.73 55.9 0 025 7.6 0 0 1 0 0 1 29.98 140 1.88 4.51 53.4 1 0 26 7.9 0 0 1 0 0 1 26.00 118 2.34 4.53 51.0 0 0 27 7.3 0 0 1 0 0 1 29.05 110 1.44 4.67 50.6 0 0 310 288.2 0 0 1 0 0 1 27.21 131 2.50 4.68 57.7 0 0 29 3.8 0 0 1 0 1 0 36.97 141 4.60 4.25 58.7 1 0 30 6.9 0 0 1 0 1 0 29.44 115 2.89 4.26 53.6 0 131 6.1 0 0 1 0 1 0 33.85 154 3.48 4.48 51.2 1 0 32 7.2 0 0 1 0 1 0 32.13 122 2.92 4.48 55.2 0 0 33 8.4 0 0 1 0 1 1 27.52 135 2.39 4.42 53.7 0 034 5.0 0 0 1 0 1 1 30.64 114 1.39 4.45 54.9 1 0 35 6.5 0 0 1 0 1 1 29.94 120 2.96 4.49 50.7 0 0 36 6.4 0 0 1 0 1 1 29.89 115 1.68 4.52 51.3 0 037 2.6 1 1 0 0 0 0 30.88 189 5.38 4.72 73.9 1 1 38 2.7 1 1 0 0 0 1 25.05 200 3.37 4.86 77.2 1 1 39 2.7 1 1 0 0 1 0 26.80 130 2.31 5.10 73.5 0 040 3.3 1 1 0 0 1 1 21.67 111 3.53 4.18 71.1 0 0 41 2.9 1 1 0 1 0 0 36.83 114 2.64 4.52 68.2 0 0 42 0.2 1 1 0 1 0 1 21.49 125 4.61 4.69 67.3 0 043 2.1 1 1 0 1 1 0 31.05 131 1.38 4.48 69.1 0 0 44 6.8 1 1 0 1 1 1 26.78 134 4.36 4.90 61.0 1 0 45 5.7 1 1 1 0 0 0 35.78 132 9.93 5.11 52.5 0 146 1.1 1 1 1 0 0 1 28.44 134 3.54 4.32 55.7 0 0 47 6.6 1 1 1 0 1 0 24.38 124 4.16 4.00 51.8 0 1 48 1.3 1 1 1 0 1 1 34.13 126 5.87 3.95 53.1 0 149 4.6 1 2 0 0 0 0 43.23 128 5.08 5.25 72.2 0 1 50 6.3 1 2 0 0 0 1 38.67 126 5.16 4.50 76.8 1 1 51 2.0 1 2 0 0 1 0 34.49 130 2.69 3.95 76.7 1 152 4.2 1 2 0 0 1 1 20.78 127 4.40 4.54 73.1 0 0 53 3.6 1 2 0 1 0 0 28.40 118 5.43 4.66 69.3 1 1 54 3.2 1 2 0 1 0 1 28.73 154 1.94 5.24 68.9 1 155 4.5 1 2 0 1 1 0 44.25 97 2.01 4.40 68.6 0 1 56 4.5 1 2 0 1 1 1 32.46 141 0.74 4.39 63.5 1 0 57 6.1 1 2 1 0 0 0 39.72 118 2.39 3.93 52.6 0 1 (Continuedoverleaf ) 311 Table12.4 Continued ID T CENS DG AGEA AGEB SEX SMOKE BMI SBP LACR LTG AGE HTN DM 58 3.0 1 2 1 0 0 1 27.90 117 7.45 5.61 56.0 0 1 59 2.1 1 2 1 0 1 0 27.77 119 7.03 4.71 54.3 0 1 60 1.3 1 2 1 0 1 1 31.03 151 3.94 4.43 59.2 1 161 4.9 1 3 0 0 0 0 25.22 129 6.69 3.90 75.4 1 0 62 2.5 1 3 0 0 0 1 45.29 130 2.46 4.40 75.7 0 1 63 3.8 1 3 0 0 1 0 25.03 188 6.25 5.63 71.7 1 164 5.0 1 3 0 1 1 0 46.76 96 3.93 4.12 65.6 1 0 65 1.5 1 3 0 1 1 1 28.53 126 3.09 4.65 68.6 0 1 66 4.1 1 3 1 0 0 0 23.63 144 8.24 4.82 59.4 1 167 0.5 1 3 1 0 1 0 31.39 134 6.96 4.11 54.2 1 0 68 2.7 1 3 1 0 1 1 30.29 115 4.70 4.98 59.1 1 1 /p63ID, participant id number; T, CVD event time (CVD-free time );C E N S /p580 if censored, and /p581 if uncensored; DG /p580i fn o n - C V Da t the end of the study or non-CVD death, /p581i fs t r o k e , /p582 if coronary heart disease (CHD ),a n d /p583 if the other CVDs; AGEA /p581i f aged 50—59 and /p580o t h e r w i s e ;A G E B /p581i fa g e d6 0—69 and /p580o t h e r w i s e ;S E X /p581i fm a l ea n d /p580 if female; SMOKE /p581i fc u r r e n t smoker and 0 otherwise; BMI, body mass index; SBP, systolic blood pressure; LACR, logarithm of the ratio of urinary albumin and creatinine; LTG, logarithm of triglycerides; HTN /p581i fS B P /p46140mmHg or DBP (diastolic blood pressure )/p4690mmHg and /p580 otherwise; DM /p581 if fasting glucose /p46126mg/dL and /p580o t h e r w i s e . 312 Table12.5 ResultsfromFittingaCoxProportionalHazardsModelBasedonDifferent MethodsforTiesontheCVDData Regression Coefficient Variable Breslow Discrete Efron Exact Exponential Weibull AGEA /p571.3478 /p571.3662 /p571.3558 /p571.3560 /p571.2550 /p571.0436 AGEB /p570.7709 /p570.7828 /p570.7753 /p570.7755 /p570.7107 /p570.5966 SEX 0.7134 0.7233 0.7187 0.7189 0.6862 0.5659 SMOKE 0.3762 0.3810 0.3776 0.3776 0.3440 0.2855BMI 0.0253 0.0256 0.0255 0.0255 0.0233 0.0194 LACR 0.1735 0.1759 0.1739 0.1740 0.1658 0.1357 likelihood function. The estimated regression coefficients based on the two parametric models, particularly the exponential regression model, are alsoclose to those based on the Cox hazards model. From the signs of thecoefficients,we see that men, current smokers,and persons with high BMIandalbumin—creatinine ratios have a higher hazard (risk)of CVD and shorter CVD-free time. The coefficients of the two age variables are both negative,indicating that persons in the younger age groups have a lower hazard (risk) of CVD. Supposethat‘‘C: /p33EX12d2d1.DAT’’containseightsuccessivecolumns,forT, CENS,AGEA,AGEB,SEX,SMOKE,BMI,andLACR,andthatthenumbersin each row are space-separated.The following code for the SAS PHREG andLIFEREG procedures can be used to obtain the results in Table 12.5. data w1; infile ‘c: /p33ex12d2d1.dat’ missover; input t cens agea ageb sex smoke bmi lacr; run;proc phreg; model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58breslow; run;proc phreg; model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58discrete; run;proc phreg; model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58efron; run;proc phreg; model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58exact; run;proc lifereg; Model a: model t*cens (0)/p58agea ageb sex smoke bmi lacr / d /p58exponential; Model b: model t*cens (0)/p58agea ageb sex smoke bmi lacr / d /p58weibull; run;      313 12.2 IDENTIFICATIONOFSIGNIFICANTCOVARIATES As noted earlier, one principal interest is to identify significant prognostic factors or covariates. This involves hypothesis testing and covariate selectionprocedures, similar to those discussed in Chapter 11 for parametric methods.The differences are that the Cox proportional hazard model has a partiallikelihood function in which the only parameters are the coefficientsassociated with the covariates. However, statistical inference based on the partial likelihood function has asymptotic properties similar to those basedon the usual likelihood. Therefore, the estimation procedure (discussed in Section 12.1 )is similar to those in Section 7.1, and the hypothesis-testing procedures are similar to those in Sections 9.1 and 11.2. For example, theWald statistic in (9.1.4 )can be used to test if any one of the covariates has no effect on the hazard, that is, to test H/p15:b/p71/p580. By replacing the log-likelihood function with the log partial likelihood function, the log-likelihood ratiostatistic, the Wald statistic, and the score statistic in (9.1.10 ),(9.1.11 ), and (9.1.12 )can be used to test the null hypothesis that all the coefficients are equal to zero, that is, to test H/p15:b/p16/p580,b/p17/p580,...,b/p78/p580 orH/p15:b/p580in(9.1.9 ). Similarly the forward, backward, and stepwise selection procedures discussed in Section 11.9.1 are applicable to the Cox proportionalhazard model. The following example, using the SAS PHREG procedure, illustrates these procedures. Example 12.4 We use the entire CVD data set in Example 12.3 to demonstrate how to identify the most important risk factors among all thecovariates. Suppose that the effects of age, gender, and current smoking statusonCVDriskareoffundamentalinterestandwewishtoincludethesevariablesin the model. In epidemiology this is often referred to as adjusting for thesevariables. Thus, AGEA, AGEB, SEX, and SMOKE are forced into the modeland we are to select the most important variables from the remainingcovariates (BMI,SBP,LACR,LTG,HTN,andDM ),adjustingforage,gender, and current smoking status. The SAS procedure PHREG is used with Breslow’s approximation for ties (default procedure )and three variable selection methods (forward, backward, and stepwise ). Two covariates, BMI and LACR, are selected at the 0.05 significance level by all three selection methods. The final model, in the formof(12.1.5 ),includingonlythefourcovariatesthatwepurposefullyincludedand the two most significant ones identified by the selection method, is314         Table12.6 AsymptoticPartialLikelihoodInferenceontheCVDDatafromtheFinal CoxProportionalHazardsModel /p63 95%Confidence Interval Regression Standard Wald Relative Variable Coefficient Error Statistic pHazards Lower Upper FinalModelfortheCohortCVDData AGEA /p571.3558 0.2712 24.9910 0.0001 0.258 0.151 0.439 AGEB /p570.7753 0.2618 8.7709 0.0031 0.461 0.276 0.769 SEX 0.7187 0.2193 10.7457 0.0010 2.052 1.335 3.153 SMOKE 0.3776 0.2208 2.9235 0.0873 1.459 0.946 2.249 BMI 0.0255 0.0124 4.2113 0.0402 1.026 1.001 1.051LACR 0.1739 0.0446 15.2112 0.0001 1.190 1.090 1.299 b/p16/p57b/p17/p570.580 4 .9443 0.0262 0.560 b/p18/p59b/p191.096 11 .5409 0.0007 2.993 b/p18/p57b/p190.341 1 .3001 0.2542 1.407 HypothesisTestingResults (H/p15:allb/p71/p580) Log-partial-likelihood ratio statistic 42.1130 0.0001 Score statistic 43.1750 0.0001Wald statistic 41.3830 0.0001 /p63The covariates, except AGEA, AGEB, SEX, and SMOKE, in the final model are selected among BMI, SBP, LACR, LTG, HTN, and DM.logh(t/p71) h/p15(t/p71)/p58b/p16AGEA/p71/p59b/p17AGEB/p71/p59b/p18SEX/p71/p59b/p19SMOKE/p71 /p59b/p20BMI/p71/p59b/p21LACR/p71 /p58/p571.3558AGEA/p71/p5707753AGEB/p71/p590.7187SEX/p71 /p590.3776SMOKE/p71/p590.0255BMI/p71/p590.1739LACR/p71(12.2.1 ) The regression coefficients, their standard errors, the Wald test statistics, p values, and relative hazards (relative risks as they are termed by many epidemiologists )are given in Table 12.6. The estimated regression coefficients b/p19/p71,i/p581, 2,...,6, are solutions of (12.1.9 )using the Newton —Raphson iterated procedure (Section 7.1 ). The estimated variances of b/p19/p71,i/p581, 2,...,6, are the respective diagonal elements of the estimated covariance matrix defined in(12.1.13 ).The square rootsof these estimatedvariancesare thestandarderrors in the table. The Wald statistics are for testing the null hypothesis that thecovariate is not related to the risk of CVD or H/p15:b/p71/p580,i/p581,...,6, respect- ively. For example, the Wald statistic equals 10.7457 for gender with a pvalue    315 of 0.0010 and b/p580.7187. It indicates that after adjusting for all the variables in the model (12.2.1 ), gender is a significant predictor for the development of CVD,withmenhavingahigherriskthanwomen.Therelativehazard (orrisk ) isexp(b/p19/p16),andforthecovariategender,it isexp (0.7187 )/p582.052,whichimplies that men aged 50 —79 years have about twice the risk of developingCVD in 10 years. The 95%confidence interval for the relative risk is (1.335, 3.153 ), which is calculated according to (7.1.8 ). For a continuous variable, exp( b/p19/p71)represents the increase in risk corresponding to a 1-unit increase in the variable. For example,for BMI,exp (0.0255 )/p581.026; that is, for every unitincreasein BMI, the risk for CVD increases 2.6%. To compare hazards among different age groups, between genders, or between smokers and nonsmokers, let h/p31/p37/p35/p31(t),h/p31/p37/p35/p32(t),h/p31/p37/p35/p33(t),h/p43/p31/p42(t), h/p36/p35/p43(t),h/p49/p43(t), andh/p44/p49/p43(t) denote hazard functions for participants that are 50—59, 60—69, 70—79 years old, male, female, current smoker, and not current smoker, respectively. The log hazard ratio of a person in the 50 to 59-yearage group to a person in the 70 to 79-year group assuming the two people areof same gender and the same current smoking status, BMI and LACR, islog[h/p31/p37/p35/p31(t)/h/p31/p37/p35/p33(t)]/p58b/p16; similarly, log[ h/p31/p37/p35/p32(t)/h/p31/p37/p35/p33(t)]/p58b/p17and log[h/p31/p37/p35/p31(t)/h/p31/p37/p35/p32(t)]/p58b/p16/p57b/p17. Assuming that the two people are in the same age groupand have the same BMI and LACR, the log hazard ratio ofmale to females is logh/p43/p31/p42(t) h/p36/p35/p43(t) /p58b/p18 Similarly,assuming thatthe two people arein the same agegroup,of the same gender, and have the same BMI and LACR, the hazard ratio of a smoker to anonsmoker is logh/p49/p43(t) h/p44/p49/p43(t)/p58b/p19 Thus, testing whether risk of CVD are the same among different age groups is equivalent to testing H/p15:b/p16/p580,H/p15:b/p17/p580, andH/p15:b/p16/p57b/p17/p580. Similarly, to test if the risk of CVD is the same between males and females or betweensmokersandnonsmokersisequivalenttotastingthenullhypothesis H/p15:b/p18/p580 orH/p15:b/p19/p580, respectively. To consider more than one covariate, we also can formulate the null hypothesis by using (12.2.1 ). For example, if we wish to compare male nonsmokers to female smokers, from (12.2.1 ), logh/p43/p31/p42/p92/p44/p49/p43h/p36/p35/p43/p92/p49/p43 /p58b/p18/p57b/p19316         assuming that they are in the same age groupand have the same BMI and LACR. Thus to test if these two groups of people have the same risk of CVD,we test the null hypothesis H/p15:b/p18/p57b/p19/p580. Similarly, to compare male smokerstofemalenonsmokers,wecantestthenullhypothesis H/p15:b/p18/p59b/p19/p580. Thesenullhypothesesareintheformoflinearcombinationsofthecoefficients.Using the notations in Section 11.2, the hypotheses H/p15:b/p16/p57b/p17/p580 and H/p15:b/p18/p59b/p19/p580 are the hypotheses in (11.2.13 )withc/p580,L/p58(1/p5710000 ), andL/p58(001100 ), respectively. The Wald statistics in Table 12.6 are calculated according to (11.2.14 ). By assuming that the patients have the same BMI and LACR, we can construct hypotheses to compare subgroups definedby age groups, gender, and current smoking status. The last part of Table 12.6 shows the results of testing the null hypothesis that none of these covariates have any effect on the development of CVD. Thelog partial likelihood ratio, Wald, and score statistics, X/p42,X/p53, andX/p49are calculated according to (9.1.10 ),(9.1.11 ), and (9.1.12 ), respectively. Table 12.6 indicates that the hypotheses, H/p15:b/p16/p580,H/p15:b/p17/p580,H/p15:b/p16/p57b/p17/p580, H/p15:b/p18/p580,H/p15:b/p20/p580,H/p15:b/p21/p580,andH/p15:b/p18/p59b/p19/p580 are rejected ata signifi- cance level of p/p580.05. However, the hypotheses H/p15:b/p19/p580 and H/p15:b/p18/p57b/p19/p580 are not rejected at a 0.05 level. The null hypothesis H/p15:allb/p71/p580,i/p581,...,6, is rejected with p/p580.0001 by using any of these tests. Assuming that the other covariates are the same, based on the relative hazards shown in the table, we conclude that (1)participants aged 50 —59 and 60—69 have, respectively, about 25% and 50% lower CVD risk than those aged 70 —79(H/p15:b/p16/p580 andH/p15:b/p17/p580 are rejected );(2)participants aged50—59have50%lowerCVDriskthanthoseaged60 —69(H/p15:b/p16/p57b/p17/p580 is rejected );(3)men’s CVD risk is twice as high as that of women (H/p15:b/p18/p580 is rejected );(4)BMI and LACR have a significant effect on CVD risk (H/p15:b/p20/p580 andH/p15:b/p21/p580 are rejected )and the risk increases about 3% and 19%,respectively,forevery1-unitincreaseinBMIandLACR,respectively; (5) male smokers have a CVD risk three times higher than that of femalenonsmokers (H/p15:b/p18/p59b/p19/p580 is rejected );(6)male nonsmokers have CVD risk similartothatoffemalesmokers (H/p15:b/p18/p57b/p19/p580isnotrejected );(7)consider- ing current smoking status alone, smokers had similar CVD risk as non-smokers (H/p15:b/p19/p580 is not rejected ). This example is solely for the purpose of illustrating the use of the proportional hazards model and the interpretationof its results. Other hypotheses of interest can be constructed in a similarmanner. The construction of null hypotheses for comparisons among sub-groups defined by AGEGROUP*SEX*SMOKE are left to the reader asexercises. Suppose that ‘‘C: /p33EX12d4d1.DAT’’ is a text data file that contains 12 successive columns for T, CENS, AGEA, AGEB, SEX, SMOKE, BMI,LACR,SBP,LTG,HTN,andDM.ThefollowingSAScodeisusedtoobtainedthe results in Table 12.6.    317 data w1; infile ‘c: /p33ex12d4d1.dat’ missover; input t cens agea ageb sex smoke bmi lacr sbp ltg htn dm; run; proc phreg data /p58w1; model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm / include /p584 selection /p58f; run; proc phreg data /p58w1; model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm / include /p584 selection /p58b; run;proc phreg data /p58w1 outest /p58wcov covout; model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm / include /p584 selection /p58s; run; proc phreg data /p58w1; model t*cens (0)/p58agea ageb sex smoke bmi lacr sbpltg htn dm / include /p584 selection /p58score best /p583; run;data wcov; set wcov; if-type-/p58‘cov’; keepagea ageb sex smoke bmi lacr sbpltg htn dm; run;title ‘The estimated covariance of the estimated coefficients’;proc print data /p58wcov; run; The following SPSS code can be used to select an optimal subset of covariates among all covariates by the forward and backward selectionmethods defined in Section 11.9.1 and to obtain the estimated coefficients andthe other results in Table 12.6. data list file /p58‘c:/p33ex12d4d1.dat’ free / t cens agea ageb sex smoke bmi lacr sbpltg htn dm. coxreg t with agea ageb sex smoke bmi lacr sbpltg htn dm /status /p58cens event (1) /method /p58fstepbmi lacr sbpltg htn dm /criteria pin (0.05)pout (0.05) /print /p58all. coxreg t with agea ageb sex smoke bmi lacr sbpltg htn dm /status /p58cens event (1) /method /p58bstepbmi lacr sbpltg htn dm /criteria pin (0.05)pout (0.05) /print /p58all.318         If BMDP 2L is used, the following code is applicable when selecting an optimal subset of covariates among all covariates by the stepwise selectionmethod defined in Section 11.9.1 and to obtain the results in Table 12.6. /input file /p58‘c:/p33ex12d4d1.dat’ . variables /p5812. format /p58free. /print cova. /variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, htn, dm. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, htn, dm. Step /p58phh. Example 12.5 If we do not force age, gender, and current smoking status on the model and are not interested in the three age groups, we can fit theproportional hazard model with age as a continuous variable and the othercovariates: SEX, SMOKE, BMI, SBP, LACR, LTG, HTN, and DM. UsingBreslow’s method for ties, the stepwise selection method, and the SAS pro-cedure PHREG, the final model with significant (p/p580.05)covariates is logh(t) h/p15(t) /p580.697AGE /p590.7528SEX /p590.1111LACR /p590.3987LTG (12.2.2 ) The details are given in Table 12.7; all four covariates in the model have positive coefficients, indicating that the risk of developing CVD increases withage, gender, albumin/creatinine ratio, and triglyceride values. The relativehazards represent the increase in risk of CVD per unit increase in thecovariates. For example, for every 1-unit increase in log (albumin/creatinine ), the risk of developing CVD increases 12%after adjusting for age, gender, andlog triglyceride. Men have more than twice the risk of CVD as women. Theglobal null hypothesis that all four coefficients equal zero ( H/p15:allb/p71/p580)is rejected by all three tests, as given in the lower part of Table 12.7. 12.3 ESTIMATIONOFTHESURVIVORSHIPFUNCTIONWITH COVARIATES Whenparametricregressionmodels (Chapter11 )areused,wecanestimatethe survivorshipfunctionsimplybyreplacingtheparametersandcoefficientsinthe survival function with their estimates. This is not the case when the Cox       319 Table12.7 AsymptoticPartialLikelihoodInferenceontheCVDDatafromtheFinal CoxProportionalHazardsModelSelectedbytheStepwiseModelSelectionMethod /p63 95%Confidence Interval for Relative Hazards Regression Standard Chi-Square Relative Variable Coefficient Error Statistic pHazards Lower Upper AGE 0.0697 0.0136 26.1393 0.0001 1.07 1.04 1.10 SEX 0.7528 0.2192 11.7893 0.0006 2.12 1.38 3.26LACR 0.1111 0.0459 5.8602 0.0155 1.12 1.02 1.22 LTG 0.3987 0.1976 4.0722 0.0436 1.49 1.01 2.20 H/p15:All coefficients equal zero Log-partial-likelihood ratio statistic 44.002 0.0001 Score statistic 44.278 0.0001 Wald statistic 42.527 0.0001 /p63The covariates in the final model are selected among AGE, SEX, SMOKE, BMI, LACR, LTG, HTN, and DM using the stepwise selection method. proportional hazards model is used since we do not know the exact form of the baseline hazard function or the survival function. In this section weintroduce briefly two estimators of the survival function, one proposed byBreslow (1974 )and the other by Kalbfleisch and Prentice (1980 ). These estimates are available in commercialsoftware packages. Readers interested indetails are referred to the corresponding publications. As indicated earlier, under the Cox model, the survivorshipfunction with covariatesx/p72’s is S(t,x)/p58[S/p15(t)] exp(/afii9814/p78/p72/p14/p16b/p72x/p72)(12.3.1) Once the regression coefficients, the b/p72’s, are estimated, we need only estimate the underlying survivorshipfunction, S/p15(t). From the estimated survivorship function,wecaneasilyestimatetheprobabilityofsurvivinglongerthanagiventime for a patient with a given set of covariates x/p16,...,x/p78. Byassumingthatthebaselinehazardfunctionisconstantbetweeneachpair of successive observed failure times, Breslow has proposed the followingestimator of the baseline cumulative hazard function: H/p19/p15(t)/p58/p26 t/p7/p71/p8/p45tm/p7/p71/p8/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)(12.3.2 )320         Following (2.15), the baseline survival function can be estimated as S/p19/p15(t)/p58exp[ /p57H/p19/p15(t)]/p58/p147 t/p7/p71/p8/p45t/p7exp/p3m/p7/p71/p8/p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)/p4/p8(12.3.3 ) and the survivorshipfunction for a p erson with a set of covariates x/p58(x/p16,...,x/p78)is S/p19(t,x)/p58[S/p19/p15(t)]exp(/afii9814/p78/p72/p14/p16b/p19/p72x/p72)/p58[S/p19/p15(t)]exp(b/p19/p30x)(12.3.4 ) Under mild assumptions, S/p19(t,x)has an asymptotic normal distribution with meanS(t,x). SinceS(t,x)/p58exp[ /p57H(t,x)], the variance estimator Var /p19(S/p19(t,x)) ofS/p19(t,x)is Var/p19(S/p19(t,x))/p60[S/p19(t,x)]/p17Var/p19(H/p19(t,x)) We will not give H/p19(t,x)here because of its complexity. The asymptotic confidence bands for the survivorshipfunction is /p37S/p19(t,x)/p57Z/p63/p30/p17/p40Var/p19(S/p19(t,x)),S/p19(t,x)/p59Z/p63/p30/p17/p40Var/p19(S/p19(t,x))/p38 (12.3.5 ) whereZ/p63/p30/p17is the upper 100 (1/p57/afii9825/2)percentile point of the standard normal distribution. An alternative estimator has been suggested by Kalbfleisch and Prentice in whichthebaselinesurvivorshipfunction S/p15(t) isestimatedtobeastepfunction and S/p19/p15(t)/p58/p71/p92/p16/p147 /p72/p14/p15/afii9825/p24/p72t/p7/p71/p92/p16/p8/p58t/p45t/p7/p71/p8,i/p581,...,k/p591 (12.3.6 ) where /afii9825/p24/p15/p891and /afii9825/p24/p16,/afii9825/p24/p17,...,/afii9825/p24/p73arethesolutionofthefollowing ksimultaneous equations: /p26 j/p43u*/p7/p71/p8exp(x/p30/p72b/p19) 1/p57/afii9825/p24/p71exp(x/p30/p72b/p19)/p58/p26 l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)i/p581,...,k(12.3.7) When there are no ties, /afii9825/p24/p71/p58/p31/p57exp(x/p30/p7/p71/p8b/p19) /p26l/p43R(t/p7/p71/p8)exp(x/p30/p74b/p19)/p4exp(/p57x/p30/p7/p71/p8b/p19) i/p581,...,k(12.3.8) and S/p19/p15(t)/p58/p71/p92/p16/p147 /p72/p14/p15/p31/p57exp(x/p30/p7/p72/p8b/p19) /p26l/p43R(t/p7/p72/p8)exp(x/p30/p74b/p19)/p4exp(/p57x/p30/p7/p72/p8b/p19) t/p7/p71/p92/p16/p8/p45t/p58t/p7/p71/p8i/p581,...,k/p591       321 Thus, S/p19(t,x)/p58[S/p19/p15(t)]exp(b/p19/p30x)(12.3.9 ) Undermildassumptions,theKalbfleischandPrenticeestimatorin (12.3.9 )also follows an asymptotic normal distribution with mean S(t,x)and a variance thatcanbeestimated.Thusconfidencebandsforthesurvivorshipfunctioncanalso be constructed. Using (12.3.4 )withS/p15(t)i n(12.3.3 )or(12.3.6 ),the survivorshipfunctioncan be estimated with any given values of x/p16,...,x/p78. If the observed average of every covariate, x/p21/p16,...,x/p21/p78is used, the estimated survivorshipfunction can be interpreted as the survivorship function of an ‘‘average’’ person. Both the Breslow and Kalbfleisch —Prentice estimators are available in the SAS procedure PHREG. The Breslow estimator is also available in BMDP(program 2L )and SPSS (program COXREG ). The following example illus- trates the procedures. Example 12.6 Again, we use the CVD data in the Example 12.3, the data set‘‘C: /p33EX12d2d1.DAT’’,andtheSASprocedurePHREG.Weusetheaverage of each of the covariates in (12.2.1 ), and therefore the estimated survivorship function is for an average person. The Kalbfleisch —Prentice and Breslow estimates of the survival function, defined in (12.3.9 )and (12.3.4 )(Efron adjustment for ties is used ), and the lower and upper 95% confidence bands, calculated based on (12.3.5 ), are shown in Figures 12.1 and 12.2. These estimatedsurvivalfunctions,using all the covariatesin the model with averagevalues, are often referred to as the global covariate —adjusted survivorship functions. The two figures are almost identical, which indicates that the twomethods produce very similar results for this set of data. From Figure 12.1 itappears that the global covariates —adjusted survivorshipfunction decreases somewhatmore rapidly after 3.5 years. This means that the process to developCVD accelerates after 3.5 years. Using the data set ‘‘C: /p33EX12d2d1.DAT’’ defined in Example 12.3, the SAS code used for this example is the following. data w1; infile ‘c: /p33ex12d2d1.dat’ missover; input t cens agea ageb sex smoke bmi lacr; run; proc phreg data /p58w1 noprint; model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58efron; baseline out /p58base1 survival /p58survival l /p58lowb u /p58uppb / method /p58pl; run;title ’K-P estimate of the survival function and its lower and upper bands’; proc print data /p58base1; var t survival lowb uppb; run;322         Figure 12.1 Kalbfleisch—Prentice estimate of survivorshipfunction and its 95% confidence bands at the averages of the covariates from the fitted Cox proportional hazards model on the CVD data. proc phreg data /p58w1 noprint; model t*cens (0)/p58agea ageb sex smoke bmi lacr / ties /p58efron; baseline out /p58base1 survival /p58survival l /p58lowb u /p58uppb / method /p58ch; run; title ’Breslow estimate of the survival function and its lower and upper bands’;proc print data /p58base1; var t survival lowb uppb; run; The following SPSS code can be used to obtain the Breslow estimate of the survival function and its standard error at each uncensored observation. Theconfidence bands can then be calculated according to (12.3.5 ). data list file /p58‘c:/p33ex12d2d1.dat’ free / t cens agea ageb sex smoke bmi lacr. coxreg t with agea ageb sex smoke bmi lacr /status /p58cens event (1) /print /p58all.       323 Figure 12.2 Breslow estimate of the survivorshipfunction and its 95% confidence bands at the averages of the covariates from the fitted Cox proportionalhazards model on the CVD data. The corresponding BMDP 2L code is /input file /p58‘c:/p33ex12d2d1.dat’ . variables /p588. format /p58free. /print cova. Survival. /variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58agea, ageb, sex, smoke, bmi, lacr. In addition to the global covariates —adjusted survivorshipfunction defined asS/p19(t,x/p21),wherex/p21/p58(x/p21/p16,x/p21/p17,...,x/p21/p78),thesurvivorshipfunctioncanbeestimated with any specific values of one or more of the covariates and interactions. Wecan also estimate the probability of surviving longer than a given time forindividuals with a given set of values for covariates. The following is anexample.324         Figure12.3 Breslow estimate of survivorshipfunctions at the averages of BMI and LACRfromSEX*SMOKERsubgroupsin aged 70 —79 participantsfrom thefitted Cox proportional hazards model on the CVD data. Example 12.7 ForthesamemodelasinExample12.6,we canestimatethe covariate-specific survivorship function for female nonsmokers, femalesmokers,malesmokers,andmalenonsmokers.Let ususethe70 —79agegroup and assume that BMI and LACR are at the average of the respectiveSEX—SMOKE subgroup. Thus, the specific covariate vector (AGEA, AGEB, SEX, SMOKE, BMI, LACR )for female nonsmokers is (0, 0, 0, 0, 30.69, 4.62 ), where 30.69 and 4.62 are the average values of BMI and LACR for femalenonsmokers. Similarly, the specific covariate vectors for female smokers, malenonsmokers, and male smokers are, respectively, (0, 0, 0, 1, 31.19, 2.67 ),(0, 0, 1, 0, 28.19, 3.43 ), and (0, 0, 1, 1, 25.76, 3.47 ). The estimated survival curves are shown in Figure 12.3. Similarly, Figures 12.4 and 12.5 give the estimatedsurvival curves of the four groups in persons aged 60 —69 years and 50 —59 years, respectively. The groups show that in all these age groups, females havea lower risk of developing CVD (longer CVD-free time )than males. Female nonsmokershave a slightly lower risk than female smokers and the differencesincrease as age decreases. However, among males, the differences in the risk ofCVD between smokers and nonsmokers are almost negligible in the youngestgroupandmuchlargerinthetwooldergroups.Malesmokershavethehighestrisk of developing CVD (shortest CVD-free time )among the four groups.       325 Figure12.4 Breslow estimate of survivorshipfunctions at the averages of BMI and LACRfromSEX*SMOKERsubgroupsin aged 60 —69 participantsfrom thefitted Cox proportional hazards model on the CVD data. 12.4 ADEQUACYASSESSMENTOFTHEPROPORTIONAL HAZARDSMODEL Thevalidityofstatisticalinferencesthatleadstotheidentificationofimportant risk or prognostic factors depends largely on the adequacy of the modelselected. The proportional hazards model is used widely in medical andepidemiologicalstudies. The adequacyof this model, includingthe assumptionof proportional hazards and the goodness of fit, needs to be assessed. In thissection we introduce several methods for this purpose. A major reason forselectingthese methodsto present here is the availabilityof computer softwarethat can perform the calculations. 12.4.1 CheckingtheProportionalHazardsAssumption The proportional hazards models defined in (12.1.1 )and (12.1.3 )assume that the hazard ratio of two people is independent of time. This requires thatcovariates not be time-dependent. If any of the covariates varies with time, theproportional hazards assumption is violated. This fact can be used to test theassumption by including a time —covariate interaction term in the model and326         Figure12.5 Breslow estimate of survivorshipfunctions at the averages of BMI and LACRfromSEX*SMOKERsubgroupsin aged 50 —59 participantsfrom thefitted Cox proportional hazards model on the CVD data. testing if the coefficient for interaction is significantly different from zero. For example, we can add an interaction term x/p71torx/p71logtin the model, that is, logh(t) h/p15(t)/p58b/p16x/p16/p59/p37/p59b/p71x/p71/p59b/p71/p16x/p71t/p59b/p71/p62/p16x/p71/p62/p16/p59/p37/p59b/p78x/p78 or logh(t) h/p15(t)/p58b/p16x/p16/p59/p37/p59b/p71x/p71/p59b/p71/p16x/p71logt/p59b/p71/p62/p16x/p71/p62/p16/p59/p37/p59b/p78x/p78 With the added interaction term, the partial likelihood function becomes more complicated. Fortunately, computer software is available to carry out thecalculations. Testing procedures similar to those discussed earlier (e.g., the Waldtest ), can be used to testthe null hypothesis H/p15:b/p71/p16/p580. IfH/p15is rejected, we conclude that Cox’s proportional hazard model is not appropriate for thedata. The interaction term with log tcan be included in the model for each of the covariates separately. If none of the corresponding pnull hypotheses H/p15:b/p71/p16/p580isrejected,wemayconcludethattheproportionalhazardsassump- tion is appropriate.       327 Table12.8 AsymptoticPartialLikelihoodInferenceontheCVDDatafromtheCox ProportionalHazardsModelwithTime-DependentCovariate 95%Confidence Interval for Relative Hazards Regressor Regressor Standard Wald Relative Variable Coefficient Error Statistic pHazards Lower Upper (a) AGE 0.068 0.014 25.249 0.0001 1.07 1.04 1.1 SEX 0.759 0.218 12.056 0.0005 2.14 1.39 3.28LACR 0.111 0.046 5.781 0.0162 1.12 1.02 1.22 LTG 0.915 0.435 4.420 0.0355 2.50 1.06 5.86 LTG* /p570.390 0.298 1.710 0.1910 0.68 0.38 1.22 log(t/p591) (b) AGE 0.071 0.014 26.635 0.0001 1.07 1.05 1.1 SEX 0.741 0.220 11.327 0.0008 2.10 1.36 3.23LACR /p570.087 0.120 0.519 0.4714 0.92 0.72 1.16 LTG 0.395 0.199 3.917 0.0478 1.48 1 2.19 LACR* 0.143 0.079 3.269 0.0706 1.15 0.99 1.35 log(t/p591) (c) AGE 0.038 0.033 1.330 0.2488 1.04 0.97 1.11 SEX 0.764 0.220 12.020 0.0005 2.15 1.39 3.31LACR 0.111 0.046 5.888 0.0152 1.12 1.02 1.22 LTG 0.417 0.197 4.469 0.0345 1.52 1.03 2.24 AGE*log(t /p591) 0.023 0.023 1.046 0.3064 1.02 0.98 1.07Example 12.8 Considerthefittedproportionalhazardsmodelin (12.2.2 )for the CVD data. To check the proportional hazards assumption, we add a termLTG/p59log(t/p591) to the model. We use t/p591 instead of tto avoid negative values. Table 12.8 (a)gives the results. The pvalue for the interaction term is 0.1910. Similarly, the results in Table 12.9 (b)and (c)suggest that LACR/p59log(t/p591) andAGE /p59log(t/p591) arenotsignificanteither.Sincegender is time-independent, we may conclude that the data satisfy the proportionalhazards assumption since every covariate in the model is time-independent. Anothermethodtochecktheproportionalhazardsassumptionis tostratify the data based on some values of a covariate, fit a stratified Cox proportionalhazards model (this is discussed in Chapter 13 ), and then construct the survivorship function separately for the each stratum and plot log(/p57log(S/p19/p72(t;x/p21/p72)))j/p581, 2,...,m328         Figure 12.6 Log[ /p57log(S(t))] plots for the age-stratified Cox proportional hazards model on the CVD data.againsttimet,wheremisthenumberofstratadefinedbythecovariate, x/p21/p72isthe vectoroftheaveragevaluesoftheothercovariatesforthe jthstratum,and S/p19/p72(t;x/p21/p72) istheestimatedsurvivorshipfunctionofthe jthstratumevaluatedat tandx/p21/p72.If thehazardsareproportional,the mcurvesshouldbeparallel.Nonparallelcurves indicatedeparturefromtheproportionalhazardsassumption.Thisisbecauseifhazard functions from any two people are proportional, it can be shown from(12.1.1 )that,forany j/p34kand1 /p45j,k/p45m,thereexistsaconstant d/p72/p73suchthat S/p19/p72(t;x/p21/p72)/p58(S/p19/p73(t;x/p21/p73))/p66 /p72/p73 (12.4.1 ) Taking the logarithm twice, we have log[/p57log(S/p19/p72(t;x/p21/p72))]/p58logd/p72/p73/p59log[/p57log(S/p19/p73(t;x/p21/p73))] (12.4.2 ) Thusthecurvesoflog[ /p57log(S/p19/p72(t;x/p21/p72))]andlog[ /p57log(S/p19/p73(t;x/p21/p73))]versustshould be parallel. Example 12.9 Consider again the fitted model in (12.2.2 ); using the stratified analysis (more details are given in Chapter 13 ), we plot log[/p57logS/p19/p72(t;x/p21/p72)] againsttfor two age strata (50—64 and 65—79 years )and two gender strata separately, where x/p21/p72denotes the average values of the other covariatesfor the jth stratum. These graphs are givenin Figures 12.6 and 12.7,       329 Figure 12.7 Log[ /p57log(S(t))] plots for gender-stratified Cox proportional hazards model on the CVD data. respectively.ThetwocurvesinFigure12.6areroughlyparallel.Thetwocurves in Figure 12.7 are also parallel over time. The results suggest that theproportional hazards assumption holds. InChapter11wediscussedseveralparametricmodels.Amongthesemodels, the exponential and the Weibull are proportional hazards models, but theothers are not. Thus, if one of the other models providesa good fit to data, wewould know that the data do not meet the proportional hazards assumption.This procedure can also be served as an alternative for checking the propor-tional hazards assumption. 12.4.2 AssessingGoodnessofFitbyResiduals There are several other graphicalmethods availablefor assessing the goodness of fit of a proportional hazards model. These graphical methods are based onresiduals and are often used as diagnostic tools. In multiple regressionmethods, residuals are referred to as the difference between the observed andthepredictedvalues (basedontheregressionmodel )ofthedependentvariable. However,whencensoredobservationsarepresentandonlyapartiallikelihoodfunction is used in the proportional hazards model, the usual concept ofresiduals is not applicable. In the following we introduce three different types330         ofresiduals:theextendedCox —Snell,deviance,andSchoenfeldresiduals.These canbe plottedversusthesurvivaltime ora covariate.Thepatternofthegraphprovides some information about the appropriateness of the proportionalhazards model. It also provides information about outliers and other patterns.Similarto other graphical methods, interpretation of the residual plots may besubjective. The Cox—Snell method discussed in Section 8.4 can easily be extended to the proportional hazards model. The extended Cox —Snell residual, R/p71, for the ith individual with observed survival time tand covariates at values x/p71is defined asR/p71/p58/p57logS/p19(t/p71;x/p71), which is the estimated accumulated hazard based on the proportional hazards model. If the t/p71observed is censored, the corresponding R/p71is also censored. If the proportional hazards model is appropriate, the plot of R/p71and its Kaplan —Meier estimate of survival function (S/p19(R)) would appear as a 45° straight line. The Cox —Snell residual method is useful in assessing the goodness of fit of a parametric model (Section 11.9.4 ). However, it is not so desirable for a proportional hazards model where apartial likelihood function is used and the survivorship function is estimatedby nonparametric methods. The deviance residuals (Therneau et al., 1990 )are defined as R/p34/p71/p58sign(R/p43/p71)/p402[/p57R/p43/p71/p57/afii9829/p71log(/afii9829/p71/p57R/p43/p71)] i/p581, 2,...,n(12.4.3) where sign (·)is the sign function, which takes value 1 if its argument is positive, 0 if zero, and /p571 if negative, R/p43/p71is the martingale residual (Fleming and Harrington, 1991 )for theith individual, R/p43/p71/p58/afii9829/p71/p57R/p71i/p581,...,n and/afii9829/p71/p581 if the observed survival time t/p71is uncensored and 0 otherwise. The martingale residuals have a skewed distribution with mean zero (AndersonandGill,1982 ).Thedevianceresidualsalsohaveameanofzerobut are symmetrically distributed about zero when the fitted model is adequate.Devianceresidualsare positiveforpersonswho survivefora shorter timethanexpectedandnegative forthose whosurvive longer.Thedevianceresidualsareoften used in assessing the goodness of fit of a proportional hazards model. Another residual method was proposed by Schoenfeld (1982 )and modified by Grambsch and Therneau (1994 ). The original Schoenfeld residuals are definedforeachpersonandeachcovariateandarebasedonthefirstderivativeof the log-likelihood function in (12.1.9 ). A Schoenfeld residual for the jth covariate of the ith person with the observed survival time t/p71is R/p72/p71/p58/afii9829/p71 /p3x/p72/p71/p57/p26l/p43R(t/p7/p71/p8)x/p72/p74exp(b/p19/p30x/p74) /p26l/p43R(t/p7/p71/p8)exp(b/p19/p30x/p74)/p4j/p581, 2,...,p;i/p581, 2,...,n (12.4.4)       331 whereb/p19is the maximum partial likelihood estimator of b. The Schoenfeld residuals are defined only at uncensored survival times; for censored observa-tions they are set as missing. Since b/p19is the solution of (12.1.9 ), the sum of the Schoenfeld residuals for a covariate is zero. Thus asymptotically, the Schoen-feld residuals have a mean of zero. It can also be shown that these residualsare not correlated with one another. Grambsch and Therneau (1994 )suggested that the Schoenfeld residuals be weighted by the inverse of the estimated covariance matrix of R/p71/p58 (R/p16/p71,...,R/p78/p71)/p30denoted byV/p19(R/p71), that is, R*/p71/p58[V/p19(R/p71)]/p92/p16R/p71(12.4.5 ) The weighted Schoenfeld residuals have better diagnostic power and are used moreoftenthantheunweightedresidualsinassessingtheproportionalhazardsassumption. To simplify the computations, Grambsch and Therneau (1994 ) suggested an approximation of [ V/p19(R/p71)]/p92/p16in(12.4.5 ): [V/p19(R/p71)]/p92/p16/p60rV/p19(b/p19) whereris thenumberofeventsorthenumberofobserveduncensoredsurvival times andV/p19(b/p19)is the estimated covariance matrix of b/p19in(12.1.13 ). With this approximation, the weighted Schoenfeld residuals in (12.4.5 )can be approxi- mated by R*/p71/p58rV/p19(b/p19)R/p71(12.4.6 ) The graphs of deviance and Schoenfeld residuals against survival time or a covariatecanbeusedtochecktheadequacyoftheproportionalhazardsmodel.The presence of certain patterns in these graphs may indicate departures fromthe proportional hazards assumption, while extreme departures from the maincluster indicate possible outliers or potential stability problems of the model. Example 12.10 Consider the proportional hazards model (12.2.2 )for the CVD data. Using the estimated survivorshipfunction with covariates, weobtain the extended Cox —Snell residual R/p71values and plot the Kaplan —Meier estimateof the survivorshipfunctionof the R/p71’s. Figure12.8 givesthe extended Cox—Snellresidualplot.Theconfigurationisveryclosetoa45°line,indicating that the proportional hazards model (12.2.2 )provides a reasonable fit to the data. Figure 12.9 plots the deviance residuals against t. Roughly speaking, the residualsaredistributedsymmetricallyaroundzerobetween /p573and3 withno peculiar patterns. Larger positive (negative )residuals are associated with smaller (larger )tvalues. The deviance residuals suggest that the proportional hazards model provides a reasonable fit to the data.332         Figure12.8 Cox—Snell residuals plot from the fitted Cox proportional hazards model on the CVD data. The weighted Schoenfeld residuals versus AGE, LACR, and LTG are given in Figures 12.10 to 12.12. In all these graphs, the residuals are distributedsymmetrically around zero except that in Figure 12.12, there are two outliersin the upperright corner. These extremelylarge residualsare from people with exceptionallyhigh values of triglyceride. A large number of the residuals equalzero or are very close to zero, particularly those for AGE and LACR,suggestingthat the model is accuratein predicting the risk of developing CVDfor these people. We also fit several parametric models to the data. Table 12.9 gives the goodnessoffitassessmentsforfiveparametricmodels.Thelikelihoodratiotestresults suggest that the Weilbull regression model provides an adequate fit(p/p580.2534 ). The Weilbull fit also gives the largest BIC and AIC values, suggesting that the Weilbull fit is best among these five models. As mentionedearlier, the Weilbull model is a proportional hazards model. Thus, theparametric model fitting provides additional evidence that the proportionalhazards model is adequate. Usingthedataset‘‘C: /p33EX12d4d1.DAT’’inExample12.4,thefollowingSAS code is used to obtain the Cox —Snell, deviance, and weighted Schoenfeld residuals for AGE, LTG, and LACR in Example 12.10.       333 Figure12.9 DevianceresidualsfromthefittedCoxproportionalhazardsmodelonthe CVD data. data w1; infile ‘c: /p33ex12d4d1.dat’ missover; input t cens agea ageb sex smoke bmi lacr sbp ltg age htn dm; run;proc phreg data /p58w1 noprint; model t*cens (0)/p58age sex lacr ltg / ties /p58efron; output out /p58out1 logsurv /p58ls resdev /p58rdev wtressch /p58rage r2 rlacr rltg; run; data out1; set out1;rcs/p58-ls; run;proc lifetest data /p58out1 notable outs /p58ws noprint; time rcs*cens (0); run;data ws; set ws;mls/p58-log(survival ); run; title ‘Cox-Snell Residuals (rcs)and -log (estimated survival function of rcs )(mls)’; proc print data /p58ws; var rcs mls;334         Figure12.10 Weighted Schoenfeld residuals from the fitted Cox proportional hazards model on the CVD data. Figure12.11 Weighted Schoenfeld residuals from the fitted Cox proportional hazards model on the CVD data.       335 Figure12.12 Weighted Schoenfeld residuals from the fitted Cox proportional hazards model on the CVD data. run; title ‘Deviance residuals (rdev)and weighted Schoenfeld residuals for AGE, LACR and LTG’;proc print data /p58out1; var t age lacr ltg rage rlacr rltg rdev; run; The following SPSS code can be used to obtain Cox —Snell and Schoenfeld residuals for AGE, and LACR and LTG. data list file /p58‘c:/p33ex12d4d1.dat’ free / t cens agea ageb sex smoke bmi lacr sbpltg age htn dm coxreg t with age sex lacr ltg /status /p58cens event (1) /print /p58all /save /p58hazard resid presid. BibliographicalRemarks An excellent expository paper on statistical methods for the identification and336         Table12.9 Goodness-of-FitTestsBasedonAsymptoticLikelihoodInferenceinFitting theCVDData /p63 Model LL LLR p BIC AIC Generalized gamma /p57198.842 — — /p57217.113 /p57212.842 Log-logistic /p57203.322 — — /p57218.983 /p57215.322 Lognormal /p57206.017 14.3505 0.0002 /p57221.678 /p57218.017 Weibull /p57199.494 1.3046 0.2534 /p57215.155 /p57211.494 Exponential /p57203.061 8.4385 /p640.0147 /p57216.112 /p57213.061 Exponential /p57203.061 7.1339 /p650.0076 /p57216.112 /p57213.061 /p63LL,loglikelihood;LLR,log-likelihoodratio statistic; p,probabilitythatthe respectivechi-square random variable /p57LLR. /p64Compared to the generalized gamma fit. /p65Compared to the Weibull fit. use of prognostic factors is that of Armitage and Gehan (1974 ). Many studies of prognostic factors have been published. A few recent ones are cited here:Well et al. (1998 ), Shipley et al. (1999 ), Marrison and Siu (2000 ), Seaman and Bird (2001 ), Bolard et al. (2001 ), Vasan et al. (2001 ), Young et al. (2001 ), Meisinger et al. (2002 ), Feskanich et al. (2002 ), Williams et al. (2002 ), and Bliwise et al. (2002 ). Cox’s regression model has stimulated the interest of many statisticians. A large number of papers on this model and related areas have been publishedsince 1972. In addition to the articles cited earlier, the following are a fewexamples: Sasieni (1996 ), Alioum and Commenges (1996 ), Farrington (2000 ), Vaida and Xu (2000 ), and Zhang and Klein (2001 ). Survival data analysis methodsarecloselyrelatedtocountingprocesses,particularlytheproportionalhazards model and residual analysis. The counting process approach requiresa strong background in probability theory and stochastic processes and isbeyond the scope of this book. Interested readers are referred to Fleming andHarrington (1991 )and Andersen et al. (1993 ). EXERCISES 12.1 (a) Consider the data in Exercise Table 3.1. In addition to the five skin tests, age and gender may also have prognostic values. Examine therelationship between survival and each of seven possible prognosticvariables, as in Table 3.8. For each variable, groupthe p atientsaccording to different cutoff points. Estimate and draw the survivalfunctionforeachsubgroupusingtheproduct-limitmethod andthenuse the methods discussed in Chapter 5 to compare the survival 337 distribution of the subgroups. Prepare a table similar to Table 3.8. Interpretyourresults.Isthereasubgroupofanyvariablethatshowssignificantly longer survival times? (For the skin test results, use the larger diameter of the two. ) (b)Considerthe seven variables in part (a). Use Cox’smodel to identify the most significant variables. Compare your results with thoseobtained in part (a). 12.2 (a) Consider the data givenin Exercise Table3.3. Examine the relation- shipbetween remission duration and survival time for each of thenine possible prognostic variables: age, gender, family history ofmelanoma, and the six skin tests. Groupthe p atients according todifferent cutoff points. Estimate and draw remission and survivalcurves for each subgroup. Compare remission and survival distribu-tionofsubgroupsusingthemethodsdiscussedin Chapter5.Preparetables similar to Table 3.8. (b)Use Cox’s regression model to identify the significant variables in part (a)for their relative importance to remission duration and survivaltime.Checktheappropriatenessoftheproportionalhazardsmodel using the significant variables identified and the stratifiedanalysis and weighted Shoenfeld residuals. Interpret the results. 12.3Use the proportional hazards model to identify the most important factors related to survival time in the 157 diabetic patients in ExerciseTable3.4.Checktheappropriatenessof themodelusing allthe methodsdiscussed in Section 12.4 and interpret the results. 12.4 (a) Construct a table similar to Table 3.8 using the data given in Table 3.6. (b)Use the proportional hazards model to identify the most important factors related to survival time. (c)Is the proportional hazards model appropriate for this data set? 12.5Using the data given in Table 12.4, perform similar analyses as in Examples 12.3 to 12.10 and discuss the results obtained.338         CHAPTER 13 Identification of Prognostic Factors Related to Survival Time:Nonproportional Hazards Models In Chapter 12 we discussed the proportional hazards model for the identifica-tion of important prognostic factors, in which the covariates are assumed tobe independent of time. We also assume that there is only one cause of failure;that is, the event or failure is allowed to occur only once for each person, andthere is no correlation among failure times of different persons. However, inpractice,thecovariatesmay beobservedmore thanonceduring thestudy,andtheir values change with time, failure may be due to more than one event orcause, the same event or failure may recur during a follow-upstudy, and theevent or failure time observed may be from related persons in a family or fromthe same person at different times. In this chapter we discuss several modelsfor these situations. The first two models are extensions of the proportionalhazards model to handle time-dependent covariates and to perform stratifiedanalysis. Other models introduced in this chapter are for multiple causes offailure, recurrent events, and related observations. 13.1 MODELSWITHTIME-DEPENDENTCOVARIATES In the Cox proportional hazards model, the ratio of hazard functions for any two persons is assumed to be independent of time t, or the covariates are not time-dependent. However, it is common in practice that a study include bothtime-dependent and time-independent covariates. For example, in a longitudi-nal study of heart disease, certain demographic variables, such as gender andrace, do not change with time and are usually collected only once at thebaseline examination. Other variables, such as lipids, may vary with time andare often collected in subsequent examinations.The partial likelihoodfunctionallowingtime-dependentcovariateshasthesameformasthatin (12.1.7 )except 339 that the covariates are now a function of time. That is, the partial likelihood function with time-dependent covariates is L(b)/p58/p73/p147 /p71/p14/p16exp[/p26/p78/p72/p14/p16b/p72x/p72/p7/p71/p8(t/p7/p71/p8)] /p26l/p43R(t/p7/p71/p8)exp[/p26/p78/p72/p14/p16b/p72x/p72/p74(t/p7/p71/p8)]/p3/p58/p73/p147 /p71/p14/p16exp[b/p30x/p7/p71/p8(t/p7/p71/p8)] /p26l/p43R(t/p7/p71/p8)exp[b/p30x/p74(t/p7/p71/p8)]/p4(13.1.1 ) wherekisthenumberofdistinctfailuretimes, R(t/p7/p71/p8)istherisksetthatcontains all persons at risk at time t/p7/p71/p8,x/p74(t/p7/p71/p8)/p58(x/p16/p74(t/p7/p71/p8),x/p17/p74(t/p7/p71/p8),...,x/p78/p74(t/p7/p71/p8))/p30denotes thecovariatesobservedfromperson lattheordereduncensoredeventtime t/p7/p71/p8, andb/p30/p58(b/p16,b/p17,...,b/p78)/p30denotes the unknown coefficients. For covariates that are not time varying, their values are constant over time. For example, let x/p16/p73denotegenderofperson k,thenx/p16/p73(t)/p58x/p16/p73(0)/p58x/p16/p73forallt.Thus,inpractice, we usually have a mixture of non-time-dependent and time-dependent covari-atesinthelikelihoodfunction.Theestimationprocedureforthecoefficients, b/p72, is similar to that discussed in Chapter 12. We can also apply the modelselection methods mentioned in Chapter 11 to select the optimal subset ofcovariates as the most important prognostic or risk factors. There are two kinds of time-dependent covariates: (1)covariates that are observed repeatedly at different follow-up time points prior to the occurrenceof the event or the end of a study or the censored time; and (2)covariates that change with time according to a known mathematical function and covariatesthat have different values due to therapy, age, or the changes in medicalconditions. The following example illustrates how the Cox proportional hazards model is extended to fit observed survival or event time data with the first kind oftime-dependentcovariates, that is, covariates observed several times before theevent. Example 13.1 A study was conducted to examine whether biomarker profiles could be used for risk assessment and bladder cancer detection in acohort of workers occupationally exposed to benzidine and at risk of bladdercancer (Hemstreet et al., 2001 ). These workers were free of bladder cancer at the time of initial (or baseline )examinationand were reexaminedat least once based on their risk assessments in a seven-year period. The event timeconsidered in this study is the cancer-free time from baseline examination tolast follow-up. To simplify the analysis, we consider only four covariates: age,level of benzidine exposure, and two biomarkers, M1 and M2. The level ofbenzidine exposure (LEX )is scored based on the worker’s job position in the factory and is considered fixed (time independent ). In addition, age (AGEB ) and the two biomarkers M1B (/p580 is negative, /p581 if positive )and M2B (/p580 ifnegative, /p581ifpositive )weremeasuredatbaselineexamination (theyarenot changed with time ). At subsequent examinations, age (AGET )and the two biomarkers,M1TandM2T,weremeasuredagain (theyarechangedwithtime ) with the status of bladder cancer and the cancer-free time from baselineexamination to subsequent examination (TR). We selected a subset of 61340         persons from this study for this example. The data reproduced in Table 13.1 are solely for the purpose of illustrating the proportional hazards model withtime-dependent covariates. Thus, the results should not be interpreted as thetrue findings of this large study. Table 13.1 gives the baseline and follow-up data from the 61 participants selected from the study. We use ID numbers to distinguish the data observedfrom different participants. For example, the person in the table with ID /p584 had LEX /p5836, diagnosed as M1 positive and M2 negative (M1B /p581 and M2B /p580), and was 47.82 years old (AGEB /p5847.82 )at the baseline examin- ation (time 0 ). He was diagnosed with negative M1 and M2 (M1T /p580 and M2T /p580)and without cancer at 42.94 months (TR/p5842.94 )from the baseline examination and at 51.39 years of age (AGET /p5851.39 )(thus 42.94 was considereda censored event time, CS /p580). His third examinationwasconduc- ted at 67.06 months (TR/p5867.06 )and he was still cancer free with both M1 and M2 negative (M1T /p580 and M2T /p580)at 53.40 years old (AGET /p5853.40 ) (thus 67.06 was considered a censored event time, CS /p580). In other words, for this person, AGET /p5851.39, M1T /p580, and M2T /p580 during the time interval (0,42.94]andAGET /p5853.40,M1T /p580,andM2T /p580duringthetimeinterval (42.94, 67.06]. The event time TR was censored at the end of the first time interval (TR/p5842.94 months, CS /p580)and also at the end of the second time interval (TR/p5867.06 months, CS /p580). The left endpoint of a time interval is denoted as TL in the table. Thus, in this example, covariates LEX, AGEB,M1B, and M2Bare fixedfor all time intervals, but AGET,M1T, and M2T aretime-dependent covariates, which may change from one interval to another. To facilitate better understanding of (13.1.1. ), we use only the data from the first six people to illustrate how to construct the likelihood function (13.1.1. ). If we have only the data fromthe first six people,thereare two ( k/p582) distinct uncensored cancer-free times, t/p7/p16/p8/p5814.65(observed from the person with ID/p582)andt/p7/p17/p8/p5824.61(observed from the persons with ID /p581).A tt/p7/p16/p8, all six people are at risk and R(t/p7/p16/p8) contains all six. At time t/p7/p17/p8, only four people (ID/p581,3,4,and6 )areatriskand R(t/p7/p17/p8) containsthesefour.Thepersonwith ID/p585 is censored at 14.78 months, prior to t/p7/p17/p8. Table 13.2 gives those in the risk sets fort/p7/p16/p8andt/p7/p17/p8with values of the seven covariates. Let x/p74(t/p7/p71/p8)/p58(LEX/p74, AGEB/p74, M1B/p74, M2B/p74, AGET/p74(t/p7/p71/p8), M1T/p74(t/p7/p71/p8), M2T/p74(t/p7/p71/p8))/p30 denote the covariates from person l(ID/p58l)evaluated at the ordered uncen- soredeventtime t/p7/p71/p8,i/p581,2;l/p581,2,...,6, andb/p30/p58(b/p16,b/p17,...,b/p22)/p30denotethe unknown coefficients. Then the first term in (13.1.1 )fort/p7/p16/p8is exp[b/p30x/p17(t/p7/p16/p8)] /p26/p21/p74/p14/p16exp[b/p30x/p74(t/p7/p16/p8)]   -  341 Table13.1 Cancer-FreeTimesforWorkersExposedtoSomeChemicalElements /p63 ID LEX AGEB M1B M2B AGET M1T M2T CS TR TL 1 180 58.64 0 0 60.70 1 0 1 24.61 0.00 2 69 40.99 0 0 42.21 1 0 1 14.65 0.00 3 36 57.14 0 0 60.72 0 0 0 42.97 0.00 4 36 47.82 1 0 51.39 0 0 0 42.94 0.004 36 47.82 1 0 53.40; 0 0 0 67.06 42.94 5 36 34.85 1 0 36.08 0 0 0 14.78 0.00 6 15 64.24 0 0 67.66 0 1 0 41.03 0.007 15 60.72 0 0 64.14 0 0 0 41.00 0.00 8 15 58.97 0 0 61.54 0 0 0 30.82 0.00 8 15 58.97 0 0 62.01 1 0 0 36.40 30.828 15 58.97 0 0 62.41 1 0 0 41.26 36.40 8 15 58.97 0 0 63.00 0 0 0 48.33 41.26 8 15 58.97 0 0 63.54 0 0 0 54.83 48.338 15 58.97 0 0 64.03 1 0 0 60.71 54.83 8 15 58.97 0 0 64.49 0 0 0 66.17 60.71 8 15 58.97 0 0 65.06 1 0 0 73.07 66.179 15 49.95 0 0 49.95 0 0 0 41.00 0.00 10 15 69.19 0 0 72.61 0 0 0 41.03 0.00 11 15 48.98 0 0 52.41 0 0 0 41.20 0.0012 15 65.52 0 0 68.95 0 0 0 41.17 0.00 13 15 47.86 0 0 47.86 0 0 0 41.43 0.00 14 15 47.82 0; 0 51.28 0 1 0 41.43 0.0014 15 47.82 0 0 52.41 0 0 0 54.97 41.43 15 15 43.49 1 0 46.53 0 0 0 36.50 0.00 15 15 43.49 1 0 46.94 0 0 0 41.43 36.5015 15 43.49 1 0 47.53 0 0 0 48.46 41.43 15 15 43.49 1 0 48.56 0 0 0 60.85 48.46 16 15 41.28 0 0 44.74 0 1 0 41.56 0.0016 15 41.28 0 0 45.86 0 0 0 54.93 41.56 17 15 49.09 0 0 52.54 0 0 0 41.43 0.00 18 15 46.03 0 0 49.45 0 0 0 41.03 0.0019 15 64.41 0 0 67.85 0 0 0 41.23 0.00 20 164 52.52 0 0 53.54 1 1 1 12.32 0.00 21 15 61.51 0 0 64.94 0 0 0 41.10 0.0022 144 64.59 0 0 68.01 1 0 0 41.13 0.00 22 144 64.59 0 0 68.60 1 1 1 48.16 41.13 23 192 62.26 0 1 64.88 1 0 0 31.47 0.0023 192 62.26 0 1 65.27 0 0 1 36.17 31.47 24 54 57.56 0 0 57.95 1 0 1 4.67 0.00 25 264 60.03 0 0 73.03 1 0 0 36.07 0.0025 264 60.03 0 0 73.71 1 0 0 44.19 36.07 25 264 60.03 0 0 64.17 1 0 0 49.68 44.19 25 264 60.03 0 0 65.15 1 0 1 61.44 49.6826 40 44.30 0 0 45.48 1 0 0 14.13 0.00 26 40 44.30 0 0 46.49 1 0 1 26.25 14.13 27 265 52.84 0 1 53.98 0 0 0 13.73 0.0027 265 52.84 0 1 55.43 0 0 0 31.18 13.73 27 265 52.84 0 1 55.83 0 0 0 35.91 31.18 27 265 52.84 0 1 56.42 0 0 1 42.97 35.91342         Table13.1Continued ID LEX AGEB M1B M2B AGET M1T M2T CS TR TL 28 132 68.19 0 1 69.31 0 0 1 13.50 0.00 29 24 62.22 1 1 64.39 0 0 0 26.02 0.00 29 24 62.22 1 1 64.85 0 0 0 31.54 26.02 29 24 62.22 1 1 65.22 0 0 0 36.01 31.5429 24 62.22 1 1 65.82 0 1 0 43.27 36.01 29 24 62.22 0 1 66.89 1 0 1 56.02 43.27 30 132 68.27 0 0 70.12 0 0 1 22.14 0.0031 178 64.07 0 0 64.07 1 0 1 21.95 0.00 32 50 65.88 0 0 65.88 0 0 0 25.43 0.00 33 50 70.82 0 1 74.40 0 0 0 42.97 0.0034 50 60.53 0 1 63.54 0 0 0 36.14 0.00 34 50 60.53 0 1 64.67 0 0 0 49.68 36.14 34 50 60.53 0 1 66.18 0 0 0 67.88 49.6835 50 62.99 0 0 66.00 0 0 0 36.11 0.00 36 50 63.01 1 1 65.15 0 0 0 25.76 0.00 36 50 63.01 1 1 66.01 0 0 0 36.04 25.7636 50 63.01 1 1 66.60 0 0 0 43.07 36.04 36 50 63.01 1 1 67.68 1 0 0 56.05 43.07 37 50 63.86 0 0 66.89 0 0 0 36.40 0.0038 50 61.15 0 0 62.33 0 0 0 14.16 0.00 38 50 61.15 0 0 63.32 0 0 0 26.02 14.16 38 50 61.15 0 0 63.78 0 0 0 31.57 26.0238 50 61.15 0 0 64.75 0 0 0 43.20 31.57 38 50 61.15 0 0 65.30 1 0 0 49.87 43.20 39 50 61.02 0 0 64.03 1 0 0 36.14 0.0040 50 61.08 0 0 61.08 0 0 0 36.17 0.00 41 50 49.50 0 1 52.51 0 0 1 36.14 0.00 42 50 49.81 0 0 52.81 0 0 0 35.94 0.0043 50 49.09 0 0 52.10 0 0 0 36.17 0.00 44 50 47.07 0 0 50.08 0 0 0 36.14 0.00 45 50 63.69 0 1 64.84 0 0 0 13.90 0.0045 50 63.69 0 1 66.30 0 0 0 31.41 13.90 45 50 63.69 0 1 66.69 0 0 0 36.01 31.41 45 50 63.69 0 1 67.28 0 0 0 43.10 36.0146 50 55.77 0 0 58.77 0 0 0 36.01 0.00 47 50 60.84 0 1 61.99 1 0 0 13.83 0.00 47 50 60.84 0 1 64.98 0 0 0 49.71 13.8348 50 50.09 1 1 51.24 1 0 0 13.90 0.00 48 50 50.09 1 1 52.70 0 0 0 31.41 13.90 48 50 50.09 1 1 53.09 0 0 0 36.01 31.4148 50 50.09 1 1 54.23 0 0 0 49.77 36.01 48 50 50.09 1 1 54.76 0 0 0 56.05 49.77 48 50 50.09 1 1 55.75 0 0 0 67.98 56.0549 50 62.41 1 0 63.53 0 0 0 13.50 0.00 49 50 62.41 1 0 65.38 0 0 0 35.61 13.50 50 50 73.88 0 1 78,03 0 0 0 49.81 0.0051 50 44.68 0 0 47.68 1 0 0 35.98 0.00 51 50 44.68 0 0 49.36 0 0 0 56.12 35.98 52 50 62.67 0 0 65.66 0 0 0 35.91 0.00 (Continued overleaf )   -  343 Table13.1 Continued ID LEX AGEB M1B M2B AGET M1T M2T CS TR TL 53 275 74.28 0 1 75.34 1 1 0 12.75 0.00 53 275 74.28 0 1 77.28 0 0 0 36.04 12.75 53 275 74.28 0 1 77.70 1 0 0 41.07 36.04 53 275 74.28 0 1 78.26 1 0 1 47.80 41.0754 57 39.52 0 0 43.50 0 0 0 47.80 0.00 5 57 76.22 1 0 79.23 1 0 0 36.07 0.00 5 57 76.22 1 0 79.64 0 0 0 41.10 36.075 57 76.22 1 0 80.21 0 0 0 47.84 41.10 5 57 76.22 1 0 81.24 0 0 0 60.29 47.84 56 57 62.41 0 0 65.83 0 0 0 41.10 0.0057 57 67.64 0 0 71.06 0 0 0 41.10 0.00 58 57 80.61 0 0 84.03 0 1 0 41.10 0.00 58 57 80.61 0 0 85.14 0 0 0 54.37 41.1059 57 67.78 1 0 67.68 1 0 0 72.12 0.00 60 0 47.35 0 1 47.35 0 1 0 13.83 0.00 60 0 47.35 0 1 49.84 0 0 0 43.70 13.8361 0 40.98 1 0 42.13 0 0 0 13.83 0.00 61 0 40.98 1 0 43.59 0 0 0 31.38 13.83 61 0 40.98 1 0 44.62 0 0 0 43.70 31.3861 0 40.98 1 0 46.55 0 0 0 66.92 43.70 /p63ID,participantIDnumber;LEX,levelofexposure;AGEB,ageatthebaselineexamimation;M1B and M2B, index functions of measure 1 and 2 at the baseline; M1B /p581 if measure 1 is positive and 0 if not; M2B /p581 if measure 2 is positive and 0 if not; AGET, age at the end of each time interval; M1T and M2T, index functions of measure 1 and 2 at the end of each time interval;M1T /p581ifmeasure1ispositiveand0ifnot;M2T /p581ifmeasure2ispositiveand0ifnot;CS /p580 if censored and 1 if not; TR, cancer-free time in months (or the right endpoint of time interval ); TL, left endpoint of time interval. where x/p17(t/p7/p16/p8)/p58(69, 40.99, 0, 0, 42.21, 1, 0) /p30 is the column vector of covariates from person 2, whose cancer-free time is t/p7/p16/p8/p5814.65.Thex/p74(t/p7/p16/p8)’sinthedenominatorarethecovariatevectorsobserved for the six people in the risk set R(t/p7/p16/p8) and are listed in Table 13.2. For example,x/p18(t/p7/p16/p8)/p58(36, 57.14, 0, 0, 60.72, 0, 0).The second and also the last term in (13.1.1 )fort/p7/p17/p8/p5824.61 is exp[b/p30x/p16(t/p7/p17/p8)] exp[b/p30x/p16(t/p7/p17/p8)]/p59exp[b/p30x/p18(t/p7/p17/p8)]/p59exp[b/p30x/p19(t/p7/p17/p8)]/p59exp[b/p30x/p21(t/p7/p17/p8)] wherex/p16(t/p7/p17/p8)/p58(180, 58.64, 0, 0, 60.70, 1, 0) and x/p74(t/p7/p17/p8)’s in the denominator are the observed covariate vectors from the four persons (ID/p581, 3, 4, and 6 )344         Table 13.2 Construction of the Partial Likelihood for the Cancer-Free Times from the First Six People with Time-Dependent Covariates Ordered t/p7/p16/p8/p5814.65 t/p7/p17/p8/p5824.61 Event time (observed from individual with ID /p582)( observed from individual with ID /p581) ID LEX AGEB M1B M2B AGET M1T M2T LEX AGEB M1B M2B AGET M1T M2T 1 180 58.64 0 0 60.70 1 0 180 58.64 0 0 60.70 1 0 2 69 40.99 0 0 42.21 1 0 3 36 57.14 0 0 60.72 0 0 36 57.14 0 0 60.72 0 04 36 47.82 1 0 51.39 0 0 36 47.82 1 0 51.39 0 05 36 34.85 1 0 36.08 0 06 15 64.24 0 0 67.66 0 1 15 64.24 0 0 67.66 0 1 345 Table13.3 AsymptoticPartialLikelihoodInferenceonCancer-FreeTime Datafrom FittedModelwithTime-DependentCovariates 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper LEX 0.007 0.003 5.593 0.018 1.01 1.00 1.01 M1T 1.361 0.645 4.449 0.035 3.90 1.10 13.81 inR(t/p7/p17/p8)(see Table 13.2 for details ). The partial likelihood function for this reduced data set is the product of these two terms. The partial likelihood function for the entire data set in Table 13.1 can be constructed in a similar way and estimates of the coefficients can be obtainedusing the Newton —Raphson method. The data format style in Table 13.1 is referred to as a counting process data format. The results of fitting this modelwith time-dependent covariates and a stepwise selection method are given inTable 13.3. The coefficients indicate that high levels of exposure and positiveM1 at follow-upexamination are p ositively related to the risk of a shortcancer-free time. Assuming that other measures are the same, a person with apositiveM1atfollow-upexaminationwillhave3.9timeshigherrisktodevelopbladder cancer than will someone with a negative M1. For every 1-unitincrease in LEX, the risk will increase by 1%. Suppose that the text data file ‘‘C: /p33EX1311.DAT’’ contains the data in Table 13.1 and the successive 11 columns give ID, LEX, AGEB, M1B, M2B,AGET, M1T, M2T, CS, TR, and TL. The following SAS code can be used toobtain the results in Table 13.3. data w1; infile ‘c: /p33ex13d1d1.dat’ missover; input id lex ageb m1b m2b aget m1t m2t cs tr tl; run; title ‘‘Selected Cox proportional hazards model with time dependent covariates’’;proc phreg data /p58w1; model (tl,tr)*cs(0)/p58lex ageb m1b m2b aget m1t m2t / rl ties /p58efron selection /p58s; where tl /p58tr; run; For the second type of time-dependent covariate (i.e., the covariate known to change with time according to a mathematical function ), we simply use the known mathematical function to replace the covariate. Following is a hypo-thetical example to illustrate the use of SAS, SPSS, and BMDP.346         Example 13.2 Suppose that we wish to fit the proportional hazards model to a set of survival data that has been saved in a text file ‘‘C: /p33EX1312.DAT’’. This set of data consists of survival time t, an indicator variable CENS (/p581 for an uncensored observation and 0 for a censored observation )and three covariates, X1, X2, and X3. Furthermore, assume that X3 is known tochange with time according to the function X3 /p59log(t/p591). In this case, the following SAS, SPSS, and BMDP code can be used to incorporate thistime-dependent covariate with known mathematical relationship with time into the model. data w1; infile ‘c: /p33ex1312.dat’ missover; input t cens x1 x2 x3; run;proc phreg data /p58w1; model t*cens (0)/p58x1 x2 z/ rl ties /p58efron; z/p58x3*log (t/p591) run; If the SPSS COXREG procedure is used, the code is data list file /p58‘c:/p33ex1312.dat’ free / t cens x1 x2 x3. time program Compute z /p58x3*log (t/p591). coxreg t with x1 x2 z /status /p58cens event (1) /print /p58all /save /p58hazard resid presid. For the BMDP 2L procedure, the code is /input file /p58‘c:/p33ex1312.dat’ . variables /p586. format /p58free. /print cova./variable names /p58t,cens, x1, x2, x3. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58x1, x2, z. add/p58z. /function z /p58x3*ln (time/p591).   -  347 13.2 STRATIFIEDPROPORTIONALHAZARDSMODELS The proportional hazards model in (12.1.3 )assumes that the ratio of the hazard functions of any two people with prognostic variables x/p16andx/p17is a constant, independent of time. This assumption may not always be met inpractical situations. To accommodate the nonproportional cases, Cox’s modelcanbegeneralizedusingtheconceptofstratification (KalbfleischandPrentice, 1980 ).Thedatacanbestratifiedbyacovariate:forexample,age.Ifweconsider two strata, say age /p4650 and /p5850 years, the model in (12.1.3 )becomes two models: h/p71(t/p34x)/p58h/p15/p71(t) exp /p1/p26b/p72x/p72/p2/p58h/p15/p71(t) exp (b/p30x/p71)( 13.2.1 ) wherei/p581, 2 for the two age strata. Notice that the underlying hazard functionh/p15/p71(t) is assumed to be different for the two strata; however, the regression coefficients are the same for all strata. That is, we assume that thehazards for patients may be proportional within each stratum but not amongdifferent strata (or levels ). The partial (marginal )likelihood function for all observations from the mstrata is defined as L(b)/p58/p75/p147 /p72/p14/p16L/p72(b)( 13.2.2 ) whereL/p72(b)is the partial (marginal )likelihood function for the jth stratum. The regression coefficients bcan be estimated by the Newton —Raphson method. For stratified models, the baseline survivorshipfunction for eachstratum is estimated separately based on the estimated regression coefficientsb/p19andthedatain thatstratumalonebyusingthemethodsdiscussedin Section 12.3. Example 13.3 Consider the data givenin Example12.1.1. Suppose that we are not sure if the risk of dying for patients at least 50 years of age isproportionaltothatforpatientslessthan50yearsanddecidetodoastratifiedanalysis. Two regression equations are therefore assumed: h/p16(t/p34x/p17)/p58h/p15/p16(t) exp (b/p17x/p17) h/p17(t/p34x/p17)/p58h/p15/p17(t) exp(b/p17x/p17) whereh/p16(t/p34x/p17), the hazard function for patients under 50 years of age, and h/p17(t/p34x/p17), the hazard function for patients at least 50 years, are functions of cellularity,and h/p15/p16(t) andh/p15/p17(t) aretheunderlyinghazardfunctionsforthetwo groups.Theresultsofthestratifiedanalysis, b/p19/p17/p580.22,SE (b/p19/p17)/p580.44,p/p580.31, andexp (b/p19/p17)/p581.24,areclosetothoseobtainedearlierintheunstratifiedmodel.348         Table13.4 AsymptoticPartialLikelihoodInferenceonCVD-freeTimeDatafrom FittedModels 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper UnstratifiedModelforAllCVDs AGE 0.070 0.014 26.139 0.0001 1.07 1.04 1.10 SEX 0.753 0.219 11.789 0.0006 2.12 1.38 3.26 LACR 0.111 0.046 5.860 0.0155 1.12 1.02 1.22 LTG 0.399 0.198 4.072 0.0436 1.49 1.01 2.20 Gender-StratifiedModelforAllCVDs AGE 0.063 0.013 23.430 0.0001 1.07 1.04 1.09 LACR 0.149 0.043 11.828 0.0006 1.16 1.07 1.26 Gender-SpecificProportionalHazardsModels Female SBP 0.022 0.006 13.962 0.0002 1.02 1.01 1.03 DM 0.986 0.373 6.973 0.0083 2.68 1.29 5.57 Male AGE 0.069 0.018 14.453 0.0001 1.07 1.03 1.11 LACR 0.125 0.058 4.555 0.0328 1.13 1.01 1.27 However, this may not always be the case. Because the model is stratified by age groupand no sp ecific relationshipis assumed between the hazard ratio ofpatients at least 50 years old and those under 50, tests of significance of theregression coefficients for the other variables are adjusted for age. Example 13.4 In Example 12.5 we used the stepwise selection method to fit the proportional hazards model to the CVD data in Example 12.3. Wereanalyze the data using a gender-stratified proportional hazards model.Results from the unstratified model and the stratified model (with a stepwise selection procedure )are given in Table 13.4. The unstratified model identifies AGE, SEX, LACR (logarithm of the ratio of urinary albumin and creatinine ), and LTG (logarithm of triglycerides )as significant covariates for the time to CVD. The gender-stratified model withthe stepwise selection procedure identifies AGE and LACR as the mostsignificant covariates. The coefficients are close to those obtained in theunstratified model. The log[-log (S(t))] at the averages of covariates AGE and LACR for the two strata are plotted in Figure 12.7. The two curves look    349 paralleltoeachother.Itsuggeststhatstratificationforthissetofdatadoesnot provide more information for the study. Moreover, the sex-specific propor-tional hazards model (at the bottom of Table 13.4 )show that systolic blood pressure (SBP )and diabetes are significant covariates related to the risk of CVDinwomenandAGEandLACRin men.Thus,thegender-specificmodelsprovide more information and suggest that there are differences in CVD riskfactors among men and women. The method of stratification is useful in cases when the observations from different strata are considered independent, conditional on the stratifiedvariable, or one is not interested in the effect of the stratified variable itself onthe outcome but in the interactions of the stratified variable with the othercovariatesin the model and does not knowthe exact forms of the interactions.It is clear that modeling observations from different strata separately canprovide more information than either stratification or unstratification if thesample size in each stratum is large enough. Using the data file ‘‘C: /p33EX12d4d1.DAT’’ defined in Example 12.4, the following SAS code can be used to obtain the results in Table 13.4 andFigure 12.7. data w1; infile ‘c: /p33ex12d4d1.dat’ missover; input t cens agea ageb sex smoke bmi lacr sbp ltg age htn dm; run;title ‘‘Unstratified model’’;proc phreg data /p58w1; model t*cens (0)/p58age sex bmi lacr sbpltg smoke htn dm / selection /p58b ties/p58efron; run;title ‘‘gender stratified model’’;proc phreg data /p58w1; model t*cens (0)/p58age bmi lacr sbpltg smoke htn dm / selection /p58b ties/p58efron; strata sex; run;proc phreg data /p58w1 noprint; model t*cens (0)/p58age lacr / ties /p58efron; strata sex;baseline out /p58bas1 loglogs /p58lmls; run; title ‘‘Log-logS from fitting a gender stratified model’’;proc print data /p58bas1; var sex age lacr t lmls; run;proc sort data /p58w1; by sex; run;title ‘‘gender-specific models’’;350         proc phreg data /p58w1; model t*cens (0)/p58age bmi lacr sbpltg smoke htn dm / selection /p58b ties/p58efron; by sex; run; The followingSPSS code can also be used. In this case, the data for women and men are assumed to be in the files ‘‘C: /p33EX12d4d1a.DAT’’ and ‘‘C:/p33EX12d4d1b.DAT’’ separately. data list file /p58‘‘c:/p33ex12d4d1.dat’’ free / t cens agea ageb sex smoke bmi lacr sbpltg age htn dm. coxreg t with age sex bmi lacr sbpltg smoke htn dm /status /p58cens event (1) /method /p58bstepage sex bmi lacr sbpltg smoke htn dm /criteria pin (0.05)pout (0.05) /print /p58all coxreg t with age bmi lacr sbpltg smoke htn dm /status /p58cens event (1) /strata /p58sex /method /p58bstepage bmi lacr sbpltg smoke htn dm /criteria pin (0.05)pout (0.05) /print /p58all coxreg t with age lacr /status /p58cens event (1) /strata /p58sex /print /p58all /save /p58lml. data list file /p58‘‘c:/p33ex12d4d1a.dat’’ free / t cens agea ageb sex smoke bmi lacr sbpltg age htn dm. coxreg t with age bmi lacr sbpltg smoke htn dm /status /p58cens event (1) /method /p58bstepage bmi lacr sbpltg smoke htn dm /criteria pin (0.05)pout (0.05) /print /p58all data list file /p58‘‘c:/p33ex12d4d1b.dat’’ free / t cens agea ageb sex smoke bmi lacr sbpltg age htn dm. coxreg t with age bmi lacr sbpltg smoke htn dm /status /p58cens event (1) /method /p58bstepage bmi lacr sbpltg smoke htn dm /criteria pin (0.05)pout (0.05) /print /p58all For the BMDP 2L procedure, the following code can be used. /input file /p58‘c:/p33ex12d4d1.dat’ . variables /p5813. format /p58free.    351 /print cova. Survival. /variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, age, htn, dm. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58age, smoke, bmi, lacr, sbp, ltg, htn, dm. strata /p58sex. step/p58phh. /input file /p58‘c:/p33ex12d4d1a.dat’ . variables /p5813. format /p58free. /print cova. Survival. /variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, age, htn, dm. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58age, smoke, bmi, lacr, sbp, ltg, htn, dm. step/p58phh. /input file /p58‘c:/p33ex12d4d1b.dat’ . variables /p5813. format /p58free. /print cova. Survival. /variable names /p58t,cens, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, age, htn, dm. /form time /p58t. status /p58cens. response /p581. /regress covariates /p58age, smoke, bmi, lacr, sbp, ltg, htn, dm. step/p58phh. 13.3 COMPETINGRISKSMODEL All the methods for prognostic factor analysis discussed so far deal with a single type of failure time for each study subject. This may be a perfectlyacceptable way to proceed in many cases. However, in some situations, failureon an person may be due to several distinct causes. It may be desirable todistinguish different kinds of events that may lead to failure and treat themdifferently in the analysis. For example, to evaluate the efficacy of hearttransplants, one would certainly want to treat deaths due to heart failuredifferently from deaths due to other causes, such as accident and cancer. In a352         mortality study, it may be more interesting to study separately deaths due to heartdisease,diabetes,cancer,andothersthantocombineallthecauses.Thesedifferent causes of failure are considered as competing events, which introducecompeting risks. Thus, problems arising in the analysis of data with multiplecauses are commonly referred to as competingrisk problems. We will see later that competing risk analysis, in general, requires no inference methods otherthan those introduced in Chapters 11 and 12. We focus on using theproportional hazards model to identify significant prognostic or risk factors when competing risks are present. Readers interested in additional details arereferred to Kalbfleisch and Prentice (1980 ). LetTbe the survival time, xthe covariate vector, and Jthe type or cause of failure. We define a type- or cause-specific hazard function h/p72(t;x) h/p72(t;x)/p58lim /p9/p82/p29/p15P(t/p45T/p58t/p59/afii9773t,J/p58j/p34T/p46t,x) /afii9773t ,j/p581,...,m (13.3.1 ) In words,h/p72(t;x)is the instantaneous failure rate of cause jat timetgivenx and in the presence of other (m/p571)causes of failure. The only difference between (13.3.1 )and the hazard function defined in Chapter 2 is the appear- ance ofJ/p58j. Equation (13.3.1 )is a type- or cause-specific hazard function, which is very much the same as the ordinary hazard function except that theevent is of a specific type. The overall hazard of failure is the sum of all thetype-specific hazards, that is, h(t;x)/p58/p26 /p72h/p72(t;x)( 13.3.2 ) provided that the failure types are mutually excluded. Based on (2.15), we can define the function . S/p72(t;x)/p58exp /p3/p57/p16/p82 /p15h/p72(u;x)du/p4,j/p581,...,m(13.3.3) However, these functions cannot, in general, be interpreted as survivorship functions when m/p571. Let t/p72/p16/p58t/p72/p17/p58/p37/p58t/p72/p73/p72denote the failure times for failures of type j,j/p581,...,m. Assuming proportional hazards, the hazard function in (13.3.1 )can be written as h/p72(t;x)/p58h/p15/p72(t) exp(b/p30/p72x),j/p581,...,m (13.3.4 ) which can be generalized for time-dependent covariates by replacing xwith x(t), that is, h/p72(t;x)/p58h/p15/p72(t) exp[b/p30/p72x(t)],j/p581,...,m (13.3.5 )   353 The partial likelihood function for the model in (13.3.5 )is L/p58/p75/p147 /p72/p14/p16/p73/p72/p147 /p71/p14/p16exp[b/p30/p72x/p72/p71(t/p72/p71)] /p26l/p43R(t/p72/p71)exp[b/p30/p72x/p74(t/p72/p71)](13.3.6 ) whereR(t/p72/p71)is the risk set at t/p72/p71. The estimation of the coefficients and identification of significant covariates can be carried out exactly the same way asdescribedinChapters11and12bytreatingfailuretimesoftypesotherthanjas censored observations. This is perhaps the most important concept in competing risks analysis. It is because the basic assumption for a competingrisksmodelisthattheoccurrenceofonetypeofeventremovesthepersonfromrisk of all other types of events and the person will no longer contributeto thesuccessiveriskset. Furthermore,thereis nothingtopreventone fromchoosingdifferent types of models for different h/p72(t;x)’s. For example, in a mortality study we might choose a proportional hazards model for heart disease and aparametric model for diabetes. The coefficient vector b/p72in(13.3.6 )indicates the effects of the covariates for eventtypej. Ifany covariatesarenotrelatedtoaparticulartypeorcause, they may be set to 0. If b/p72are the same for all j, the model in (13.3.5 )reduces to the proportional hazards model in Chapter 12. The following example illustratesthe proportional hazards model with competing risks. Example 13.5 Let us again use the CVD data in Example 12.3. The event typesarenon-CVD (DG/p580),stroke (DG/p581),CHD (DG/p582),andtheother CVDs (DG/p583). If one is interestedin all CVD nomatter whether it isstroke, CHD,ortheotherCVDs,thecompetingrisksmodelreducestoageneralCVDevent model, the times (T)to CVD for DG /p581, 2, 3 are uncensored event times, and the other times are censored (DG/p580). An indicator variable, CS, can be used to indicate the censoring status; that is, CS /p581i fD G /p581, 2, 3, and CS /p580 otherwise. The result from fitting the proportional hazards model with the backward selection method is given in section (a)of Table 13.5. If one considers strokes only, the indicator variable CS has to be defined differently; that is, CS /p581i fD G /p581 and CS /p580i fD G /p580, 2 and 3. This means that in addition to non-CVD, the event time of CHD and the otherCVDsaretreatedas censoredobservations.Notethatwewillremoveapersonfrom the risk set after his or her first CVD event time in constructing thelikelihoodfunctionforthemodifieddataeveniftheeventwasnotafatalevent.For stroke, age is the only significant variable [section (b)in Table 13.5]. We call this model a marginal model for strokes. Similarly, if only CHD, otherCVDs,or either stroke or CHD are of interest,the respective modificationwillbe CS /p581i fD G /p582 and CS /p580 otherwise (CHD only );C S/p581i fD G /p583 and CS /p580 otherwise (other CVD only );o rC S /p581i fD G /p581, 2 and CS /p580 otherwise (either stroke or CHD ). The results of these three fits with the backwardselectionmethodareshowninTable13.5 (c)—(e).Theresultssuggest354         Table13.5 AsymptoticPartialLikelihoodInferenceonCVDEventTimeDatafrom theFittedCompetingRisksModels 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper (a) Model for All CVDs AGE 0.070 0.014 26.139 0.0001 1.07 1.04 1.10 SEX 0.753 0.219 11.789 0.0006 2.12 1.38 3.26 LACR 0.111 0.046 5.860 0.0155 1.12 1.02 1.22 LTG 0.399 0.198 4.072 0.0436 1.49 1.01 2.20 (b) Marginal Model for Strokes AGE 0.072 0.021 12.092 0.0005 1.08 1.03 1.12 (c) Marginal Model for CHDs AGE 0.069 0.020 11.622 0.0007 1.07 1.03 1.12 SEX 0.970 0.329 8.716 0.0032 2.64 1.39 5.02 BMI 0.040 0.017 5.162 0.0231 1.04 1.01 1.08 LTG 1.106 0.266 17.234 0.0001 3.02 1.79 5.09 (d) Margial Model for Other CVDs AGE 0.087 0.033 6.874 0.0087 1.09 1.02 1.17 SEX 1.100 0.555 3.937 0.0472 3.01 1.01 8.91 LACR 0.315 1.101 9.745 0.0018 1.37 1.12 1.67 (e) Marginal Model for Strokes or CHDs AGE 0.072 0.015 23.555 0.0001 1.07 1.04 1.11 SEX 0.692 0.239 8.362 0.0038 2.00 1.25 3.20LTG 0.665 0.200 11.095 0.0009 1.94 1.32 2.88 that significant risk factors differ for different types of CVD events. Age is the only factor common to all the CVD events. Thus, competing risks models provide an opportunity to separate any one ormore specifictypes ofevent orcauseof deathfrom allothertypes orcauses.In practice, it is not necessary to fit a model to every type or cause. Suppose that ‘‘C: /p33EX13d3d1.DAT’’ is a text data file that contains 14 columns similar to Table 12.4 and the successive columns give T, CENS, DG,AGEA, AGEB, SEX, SMOKE, BMI, LACR, SBP, LTG, AGE, HTN, andDM. The following SAS code can be used to obtain the model for stroke inTable 13.5. These codes can easily be modified to obtain the results for CHD,   355 other CVD, and stroke/CHD. data w1; infile ‘c: /p33ex13d3d1.dat’ missover; input t cens dg agea ageb sex smoke bmi lacr sbp ltg age htn dm; run; title ‘‘Model for stroke event times’’; proc phreg data /p58w1; model t*dg (0, 2, 3 )/p58age sex smoke bmi lacr sbpltg htn dm / rl selection /p58b ties /p58efron; run; The following SPSS code can be used. data list file /p58‘‘c:/p33ex13d3d1.dat’’ free / t cens dg agea ageb sex smoke bmi lacr sbpltg age htn dm. coxreg t with age sex bmi lacr sbpltg smoke htn dm /status /p58dg event (1) /method /p58bstepage sex bmi lacr sbpltg smoke htn dm /criteria pin (0.05)pout (0.05) /print /p58all The following code is for the BMDP 2L procedure. /input file /p58‘c:/p33ex13d3d1.dat’ . variables /p5814. format /p58free. /print cova. Survival. /variable names /p58t,cens, dg, agea, ageb, sex, smoke, bmi, lacr, sbp, ltg, age, htn, dm. /form time /p58t. status /p58dg. response /p581. /regress covariates /p58age, sex, smoke, bmi, lacr, sbp, ltg, htn, dm. step/p58phh. 13.4 RECURRENTEVENTSMODELS So far we have considered events or failures that are allowed to occur only once. Even in competing risks models, the occurrence of one type of eventremoves a person from the risk set thereafter. However, in practice the failureson an individual may be recurrences of essentially the same event, such astumor recurrences after surgeries, or may be successive events of entirelydifferent types, such as strokes and heart attacks. When data include recurrentevents, regression models such as the proportional hazards model becomemuch more mathematically complicated and often involve counting process356         theory,whichisbeyondthe scopeof thisbook.Anumber ofregressionmodels have been proposed in the literature. In this section we introduce three modelsthat can be considered as extensions of the Cox proportional hazards model.We keepthe mathematics to a minimum and use examp les to show how thesemodels can be used to identify important prognostic or risk factors with theaid of available computer software. The three models are based on Prentice etal.(1981 ),AndersenandGill (1982 ),andWeietal. (1989 ).Allthreemodelsare proportional hazards models, and the likelihood functions of these models are constructeddifferently,primarilyintherisksetattheuncensoredobservations.Readers interested in details are referred to the papers cited above. Prentice et al. Model In their 1981 paper, Prentice, Williams, and Peterson (PWP )proposed two modelsforrecurrentevents.BothPWPmodelscanbeconsideredasextensionsof the stratified proportionalhazardsmodelwith strata defined by the numberand time of the recurrent events. The hazard function is extended beyond theperson’s first event to cover subsequent events. In the first PWP model,follow-uptimestarts atthebeginningofthestudy (truetime0 )andthehazard function of the ith person can be written as h(t/p34b/p81,x/p71(t))/p58h/p15/p81(t) exp[b/p30/p81x/p71(t)] (13.4.1 ) wherethe subscript srepresents thestratumthat thepersonis inat time t. The first stratum includes people who have at least one recurrence or are censoredwithout recurrence, the second stratum includes people who have at least tworecurrences or are censored after the first recurrence, and so on. A personmoves from stratum 1 (s/p581)to stratum 2 (s/p582)following his or her first recurrenteventandremainsinstratum2 untilthesecondrecurrenteventtakesplaceor becomesa censorobservation (no more recurrentevent ).Theh/p15/p81(t)in (13.4.1 )is the stratum-specific underlying hazard. Notice that in (13.4.1 ), the coefficients are stratum-specific also. Lett/p81/p16/p58/p37/p58t/p81/p66/p81denote thed/p81ordered distinct failure times in stratum s, x/p81/p71(t/p81/p71)thecovariatevectorofasubjectinstratum swhofailsattime t/p81/p71,x/p81/p74(t/p81/p71) the covariate vector of subject lin stratumsat timet/p81/p71, andR(t,s) the set of persons at risk in stratum sjust prior to time t. Note that the risk set R(t,s) includes only those persons who have experienced the first s/p571 recurrent events. Then the partial likelihood for the first model in (13.4.1 )is L(b)/p58/p147 s/p461/p66/p81/p147 /p71/p14/p16exp[b/p30/p81x/p81/p71(t/p81/p71)] /p26l/p43R(t/p81/p71,s)exp[b/p30/p81x/p81/p74(t/p81/p71)](13.4.2) The following example illustrates the construction of the likelihood function and the necessary data arrangements for using SAS, SPSS, or BMDP to carryout the analysis.   357 Example 13.6 We use the tumor recurrence data from bladder cancer patients (Andrews and Herzberg, 1985; Wei et al., 1989 )in a clinical trial to compare three treatments, which was conducted by the Veterans Administra-tion Cooperative Urological Research Group (Byar, 1980 ). All patients had superficial bladder tumors when they entered the study. These tumors wereremoved and the patients were randomized into three treatment groups:placebo, thiotepa, and pyridoxine. During the follow-up period many patientshad one or more recurrences of tumors and new tumors were removed when discovered.In this example we use the tumor recurrence data from 86 patientswho received either placebo or thiotepa. Only the first four recurrence timesare considered. The data set, reproduced in Table 13.6, includes treatment(1, placebo; 2, thiotepa )follow-uptime, initial number of tumors (N), initial tumor size (S)in centimeters, and recurrent time. Each recurrent time of a patient was measured from the date of first treatment. In this case, the event of interest is tumor recurrence and the strata are defined by the number of recurrences (NRs ).To use SAS and other software to fitdatawiththemodel (13.4.1 ),thedatamustberearrangedinacertainformat by stratum. To facilitate illustration, we selectsix patients from Table 13.6 andplacethe data of thesesix patientsin Table13.7. The follow-upand recurrencetimes are also shown in Figure 13.1. From the figure we see that stratum 1includes patients 1 (censored at 9 months ),2(censored at 59 months ),3(first recurrent at 3 months ),4(first recurrent at 12 months ),5(first occurrence at 6 months ), and 6 (first occurrence at 3 months ). The time intervals, (TL, TR], are(0,9], (0,59], (0,3], (0,12], (0,6], and (0,3], respectively. These intervals are used to determine the risk set in the stratum-specific likelihood function in(13.4.2 ),andthepatientsinthestratumwereatriskonlyinthesetimeintervals. To use software packages such as SAS, BMDP, and SPSS, we need torearrangethe data by stratum. Table 13.8 gives the rearranged data. Note thatthe six patients in stratum 1 are arranged in ascending order according to therightendofthetimeinterval.AlsointroducedinthistableareT1 —T4,N1—N4, and S1—S4, giving the treatment received, initial tumor number, and initial tumor size of the patients for the four strata, respectively. These variables areset to be zero in the other strata except the stratum they are in. For example,forpatientsinstratum1,T2 —T4,N2—N4,andS2—S4aresettobezerobecause thesesix patients are in stratum1, not in stratum2, 3, or 4. Stratum 2 includesthose patients who had one recurrence and had either another recurrence orwere censored at end of follow-up. Therefore, stratum 2 has patients 3(censored at 14 months after the first recurrence ),4(second recurrence at 16 months ),5(second recurrence at 12 months ), and 6 (second recurrence at 15 months ). The time intervals between successive recurrences for these four patients are (3, 14], (12,16], (6,12], and (3,15], respectively. The rearranged data in order of the right end of the intervals are given in Table 13.8. Strata 3and 4 are constructed in a similar way. Once the data are rearranged exactlyas in Table 13.8, SAS and other software can be used to perform the analysis. This dataarrangementalsofacilitatesexplanationofthe likelihoodfunction358         Table13.6 TumorRecurrenceDataforPatientswithBladderCancer /p63 Recurrence Time Treatment Follow-upInitial Initial GroupTime Number Size 1234 10 1 1 11 1 3 14 2 1 17 1 111 0 5 1 11 0 4 1 6 11 4 1 111 8 1 1 11 8 1 3 5 11 8 1 1 1 2 1 612 3 3 3 12 3 1 3 1 0 1 5 1 23 1 1 3 16 2312 3 3 1 3 9 2 1 1 24 2 3 7 10 16 24 1 25 1 1 3 15 2512 6 1 2 12 6 8 1 1 12 6 1 4 2 2 612 8 1 2 2 5 12 9 1 4 12 9 1 212 9 4 1 13 0 1 6 2 8 3 0 1 30 1 5 2 17 221 3 0 2 1 368 1 2 1 3 1 1 3 1 21 52 4 13 2 1 213 4 2 1 13 6 2 1 13 6 3 1 2 913 7 1 2 1 40 4 1 9 17 22 24 1 4 0 5 1 1 61 92 32 914 1 1 2 14 3 1 1 3 14 3 2 6 61 4 4 2 1 369 1 45 1 1 9 11 20 26 14 8 1 1 1 814 9 1 3 15 1 3 1 3 5 15 3 1 7 1 71 53 3 1 3 15 46 51 15 9 1 1 1 61 3 2 2 15 24 301 64 1 3 5 14 19 27 (Continued overleaf )   359 Table13.6 Continued Recurrence Time Treatment Follow-upInitial Initial GroupTime Number Size 1234 16 4 2 3 2 8 1 2 1 3 21 1 321 1 125 8 1 5 29 1 2 21 0 1 121 3 1 121 4 2 6 2 1 7 5 3 3135 21 8 5 121 8 1 3 1 721 9 5 1 2 22 1 1 1 1 7 1 9 22 2 1 122 5 1 322 5 1 5 22 5 1 1 2 26 1 1 6 12 1322 7 1 1 622 9 2 1 2 23 6 8 3 2 6 3 5 23 8 1 12 3 9 1 1 2 22 32 73 22 39 6 1 4 16 23 27 2 4 0 3 1 2 42 62 94 0 24 1 3 224 1 1 124 3 1 1 1 2 7 24 4 1 1 2 44 6 1 2 20 23 2724 5 1 224 6 1 4 2 24 6 1 4 24 9 3 325 0 1 12 50 4 1 4 24 47 25 4 3 4 25 4 2 1 3 825 9 1 3 Source:Wei et al (1989 )and StatLib web site: http//lib.stat.cmu.edu/datasets/tumor. /p63Treatment group: 1, placebo; 2, thioteps. Follow-up time and recurrence time are measured in months. Initial size is measured in centimeters. Initial number of 8 denotes eight or more initialtumors.360         Figure13.1 GraphicalpresentationofrecurrencetimesofthesixpatientsinTable13.7 (numbers in circle indicate the number of recurrences ).Table13.7 Sixof86BladderCancerPatientsfromthe TumorRecurrenceData /p63 Recurrence Time Patient Treatment Follow-upInitial Initial ID GroupTime Number Size 1 2 3 4 119 1 2 20 5 9 1 1 3 1 14 2 6 3 4 0 18 1 1 12 16 5 1 26 1 1 6 12 136 0 5 3 3 1 31 54 65 1 /p63Treatment group: 0, placebo; 1, thiotepa. Following-up time and recurrence time are measured in months. Initial size is measured in centimeters for the largest initial tumor. in(13.4.2 ).Weusestratum2toshowthesecondproductin (13.4.2 ).Instratum 2(s/p582),d/p81/p583(there are three uncensored observations: patients 5, 6 and 4, according to the ordered recurrent times, 12, 15, and 16 months ). Therefore, the second product is the product of three terms, one for each of these threepatients. Using the notations in (13.4.2 ), we renumber them as patient i/p581, 2, and 3, respectively. The risk set at the first uncensored time t/p17/p16in stratum 2   361 Table13.8 RearrangedDatafromTable13.7forFittingPWPModelwith NR-IndexedCoefficients /p63 ID NR TL TR CS T1 T2 T3 T4 N1 N2 N3 N4 S1 S2 S3 S4 3 1031100020006000 6 10310000300010005 1061100010001000 1 1090100010002000 4 10 1 21000010001000 2 10 5 90000010001000————————————————————————————————————————————5 26 1 21010001000100 3 23 1 40010002000600 6 23 1 510000030001004 21 2 1 61000001000100————————————————————————————————————————————5 31 2 1 31001000100010 4 31 6 1 800000001000106 31 5 4 61000000300010————————————————————————————————————————————5 41 3 2 60000100010001 6 44 6 5 11000000030001 /p63ID, patient ID number; NR, number of recurrence, where 1 /p58first recurrence, 2 /p58second recurrence, and so on; TL and TR, left and right ends of time interval (TL, TR )defined by the successive rcurrence times and the follow-uptime, where TR denotes either the successive recurrence time or the follow-uptime; CS, censoring status, where 0 /p58censored, 1 /p58uncensored; T1 to T4, treatment group; N1 to N4, initial number of tumors; S1 to S4, initial size. (observed from patient 5 ),o rR(t/p17/p16,2) includes patients in stratum 2, whose recurrent times, censored or not, are at least 12 ( t/p17/p16) months. Therefore, R(t/p17/p16,2) includes all four patients in stratum 2. Similarly, the risk set at the second uncensored time t/p17/p17in stratum 2, R(t/p17/p17,2), includes two patients (patients 6 and 4 ), andR(t/p17/p18,2) includes only one patient (patient 4 ). Thus, usingtheIDinTable13.7,let x/p17/p18—x/p17/p21denotethecovariatevectorsforpatients 3—6 in stratum 2, the second product in (13.4.2 )is /p66/p17/p147 /p71/p14/p16exp[b/p30/p17x/p17/p71(t/p17/p71)] /p26l/p43R(t/p17/p71,2)exp[b/p30/p17x/p74(t/p17/p71)] /p58exp(b/p30/p17x/p17/p20) exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21) /p59exp(b/p30/p17x/p17/p21) exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p21)/p59exp(b/p30/p17x/p17/p19) exp(b/p30/p17x/p17/p19)(13.4.3 ) where thex’s represent the covariate vector (T1, T2, T3, T4, N1, N2, N3, N4, S1, S2, S3, S4 ). For example, x/p17/p20/p58(0,1,0,0,0,1,0,0,0,1,0,0 ). It is clear that362         Table13.9 RearrangedDatafromTable13.7forFittingPWPModelwithCommon Coefficients /p63 ID NR TL TR CS TRT N S 31031126 610310315106111111090112 41 0 1 21011 21 0 5 90011————————————————————————————————————————————52 6 1 21111 32 3 1 40126 62 3 1 510314 2 12 16 1 0 1 1————————————————————————————————————————————5 3 12 13 1 1 1 1 4 3 16 18 0 0 1 16 3 15 46 1 0 3 1————————————————————————————————————————————5 4 13 26 0 1 1 1 6 4 46 51 1 0 3 1 /p63TRT, treatment group; N, initial number; S, initial size.inthis modeltheregressioncoefficientsarestratumspecific.Theyrepresentthe importanceofthecoefficientforpatientsindifferentstrataorpatientswhohaddifferent numbers of recurrent events. If the primary interest is the overallimportance of the covariates, regardless of the number of recurrences or if itcanbeassumedthattheimportanceofcovariatesisindependentofthenumberof recurrences, T1 —T4, N1—N4, and S1—S4 can be combined into a single variable. As shown in Table 13.9, the three covariates are named TRT, N, andS for the six patients, and coefficients common to all strata can be estimated. Data sets that have been so rearranged are ready for SAS and other software. To use SAS and other software, the entire data set in Table 13.6 must first be rearranged as in Table 13.8 or 13.9. This can also be accomplished using acomputer. Table 13.10 gives the results from fitting the PWP model to the bladder tumor data in Table 13.6 with stratum-specific coefficients and commoncoefficients.Noneofthestratum-specificcovariatesissignificantexceptN1,theinitial number of tumors in stratum 1 patients (p/p580.0017 ). There is no significant difference between the two treatments in any stratum, and the sizeof the initial tumor has no significant effect on tumor recurrence. Whenstratificationisignored,the resultsare similar (thesecondpart ofTable13.10 ). The number of initial tumors is the only significant prognostic factor, and therisk of recurrence increase would increase almost 13% for every one-tumorincrease in the number of initial tumors.   363 Table13.10 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom FittedPWPModelswithStratum-specificor CommonCoefficients 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper Model with Stratum-Specific Coefficients T1 /p570.526 0.316 2.774 0.0958 0.591 0.318 1.097 T2 /p570.504 0.406 1.539 0.2148 0.604 0.273 1.339 T3 0.141 0.673 0.044 0.8345 1.151 0.308 4.305 T4 0.050 0.792 0.004 0.9493 1.052 0.223 4.963N1 0.238 0.076 9.851 0.0017 1.269 1.094 1.472 N2 /p570.025 0.090 0.075 0.7840 0.976 0.818 1.164 N3 0.050 0.185 0.072 0.7887 1.051 0.731 1.511N4 0.204 0.242 0.712 0.3987 1.227 0.763 1.971 S1 0.070 0.102 0.470 0.4931 1.072 0.879 1.308 S2 /p570.161 0.122 1.722 0.1894 0.852 0.670 1.083 S3 0.168 0.269 0.390 0.5321 1.183 0.698 2.005 S4 0.009 0.339 0.001 0.9786 1.009 0.519 1.961 ModelwithCommonCoefficients TRT /p570.333 0.216 2.380 0.1229 0.716 0.469 1.094 N 0.120 0.053 5.029 0.0249 1.127 1.015 1.251S /p570.008 0.073 0.014 0.9071 0.992 0.860 1.144 In the second PWP model, the follow-uptime starts from the immediately preceding event or failure time. Analogous to (13.4.1 ), the second PWP model can be written in terms of a hazard function as h(t/p34b/p81,x/p71(t))/p58h/p15/p81(t/p57t/p81/p92/p16)exp[b/p30/p81x/p71(t)] (13.4.4 ) wheret/p81/p92/p16denotes the time of the preceding event. The time period between two consecutive recurrent events or between the last recurrent event time andthe end of follow-upis called the gaptime. For thelth subject, who fails at time t/p81/p74in stratum s, denote the gaptime as u/p81/p74/p58t/p81/p74/p57t/p81/p92/p16/p74, wheret/p81/p92/p16/p74is the failure time of the lth subject in the stratum s/p571. Letu/p81/p7/p16/p8/p58/p37/p58u/p81/p7/p66/p81/p8denote the ordered observed distinct gaptimes in stratumsandR/p18(u,s)denote the set of subjects at risk in stratum sjust prior togaptimeu. Again,R/p18(u,s)includesonly thosesubjectswho haveexperienced the firsts/p571 strata. Then we have the partial likelihoodfor the second model (13.4.4 ): L(b)/p58/p147 s/p461/p66/p81/p147 i/p581exp[b/p30/p81x/p81/p71(t/p81/p7/p71/p8)] /p26l/p43R/p18(u/p81/p7/p71/p8,s)exp(b/p30/p81x/p81/p74(t/p81/p7/p71/p8)](13.4.5 )364         Table13.11 RearrangedDatafromTable13.9for FittingPWPGapTimeModelwithCommonCoefficients ID NR GT CS TRT N S 31 3 1 1 2 661 3 1 0 3 151 6 1 1 1 111 9 0 1 1 24 11 21011 2 15 90011————————————————————————————42 4 1 0 1 1 52 6 1 1 1 1 3 21 10126 6 21 21031————————————————————————————53 1 1 1 1 1 43 2 0 0 1 1 6 33 11031————————————————————————————64 5 1 0 3 1 5 41 30111Note that risk sets in (13.4.5 )are defined by the ordered distinct gaptimes in the strata rather than by the failure times themselves. Using the notations in Table 13.9, let GT denote the gaptime, then GT/p58TR—TL. Replacing TR and TL in Tables 13.8 and 13.9 by GT, the data are ready for SAS and other software. Table 13.11 is the corresponding tablefor the same six patients in Table 13.9 using gap times. Using the notation ofExample 13.6, the second product in (13.4.5 )for stratum 2 is /p66 /p17/p147 /p71/p14/p16exp[b/p30/p17x/p17/p71(t/p17/p71)] /p26l/p43R/p18(u/p17/p7/p71/p8,2)exp[b/p30/p17x/p17/p74(t/p17/p74)] /p58exp(b/p30/p17x/p17/p19) exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21) /p59exp(b/p30/p17x/p17/p20) exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21)/p59exp(b/p30/p17x/p17/p21) exp(b/p30/p17x/p17/p21) Note that this is different from (13.4.3 ), due to a different definition of the risk set. The results from fitting the PWP gaptime model to all the data in Table 13.6 with stratum-specific coefficients and common coefficients are given inTable 13.12. Again, the number of initial tumors is the only significant   365 Table13.12 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom theFittedPWPGapTimeModelswithStratum-SpecificorCommonCoefficients 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper Model with Stratum-Speci fic Coef ficients T1 /p570.526 0.316 2.774 0.0958 0.591 0.318 1.097 T2 /p570.271 0.405 0.448 0.5034 0.763 0.345 1.687 T3 0.210 0.550 0.146 0.7022 1.234 0.420 3.626 T4 /p570.220 0.639 0.119 0.7301 0.802 0.229 2.807 N1 0.238 0.076 9.851 0.0017 1.269 1.094 1.472 N2 /p570.006 0.096 0.004 0.9469 0.994 0.823 1.200 N3 0.142 0.162 0.774 0.3791 1.153 0.840 1.582N4 0.475 0.203 5.492 0.0191 1.609 1.081 2.394 S1 0.070 0.102 0.470 0.4931 1.072 0.879 1.308 S2 /p570.119 0.119 1.003 0.3166 0.888 0.703 1.121 S3 0.278 0.233 1.425 0.2326 1.321 0.836 2.086 S4 0.043 0.290 0.022 0.8822 1.044 0.592 1.842 Model with Common Coef ficients TRT /p570.279 0.207 1.811 0.1784 0.757 0.504 1.136 N 0.158 0.052 9.258 0.0023 1.171 1.058 1.297S 0.007 0.070 0.011 0.9157 1.007 0.878 1.156 covariates. There are no major differences between the two PWP models for thisset of data. Itis impossibleto compare the coefficients obtainedin the twomodels. The first model defines time from the beginning of the study andtherefore is recommended if the entire course of recurrent events is of interest.Thesecond model is the choice if the primary interest is to model the gap timebetween events. Suppose that the text file ‘‘C: /p33EX13d4d1.DAT’’ contains the successive columns in Table 13.8 for the entire data set in Table 13.6: NR, TL, TR, CS,T1, T2, T3, T4, N1, N2, N3, N4, S1, S2, S3, and S4, and the text file‘‘C:/p33EX13d4d2.DAT’’containsthesevensuccessivecolumnsinTable13.9:NR, TL, TR, CS, TRT, N, and S. The following SAS code can be used to obtainthe PWP models in Table 13.10. data w1; infile ‘c: /p33ex13d4d1.dat’ missover; input nr tl tr cs t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4; run; title ‘‘PWP model with stratified coefficients‘;proc phreg data /p58w1;366         model (tl, tr )*cs(0)/p58t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4 / ties /p58efron; where tl /p58tr; strata nr; run; data w1; infile ‘c: /p33ex13d4d2.dat’ missover; input nr tl tr cs trt n s; run; title ‘‘PWP model with common coefficients‘; proc phreg data /p58w1; model (tl, tr )*cs(0)/p58trt n s / ties /p58efron; where tl /p58tr; strata nr; run; Suppose that the text file ‘‘C: /p33EX13d4d3.DAT’’ contains 15 successive columns similar to Table 13.8 but with gaptime GT. The 15 columns are NR,GT, CS, T1, T2, T3, T4, N1, N2, N3, N4, S1, S2, S3, and S4. The text file‘‘C:qafii0)’07EX13d4d4.DAT’’containsthesuccessivesixcolumnsfromTable13.11:NR,GT, CS, TRT, N, and S. The following SAS, SPSS, and BMDP codes can beused to obtain the PWP gaptime models in Table 13.12. SAS code: data w1; infile ‘c: /p33ex13d4d3.dat’ missover; input nr gt cs t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4; run;title ‘‘PWP gaptime model with stratified coefficients’’;proc phreg data /p58w1; model gt*cs (0)/p58t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4 / ties /p58efron; strata nr; run;data w1; infile ‘c: /p33ex13d4d4.dat’ missover; input nr gt cs trt n s; run; title ‘‘PWP gaptime model with common coefficients‘;proc phreg data /p58w1; model gt*cs (0)/p58trt n s / ties /p58efron; strata nr; run; SPSS code: data list file /p58‘c:/p33ex13d4d3.dat’ free /n rg tc st 1t 2t 3t 4n 1n 2n 3n 4s 1s 2s 3s 4 . coxreg gt with t1 t2 t3 t4 n1 n2 n3 n4 s1 s2 s3 s4 /status /p58cs event (1)   367 /strata /p58nr /print /p58all. data list file /p58‘c:/p33ex13d4d4.dat’ free / nr gt cs trt n s. coxreg gt with trt n s /status /p58cs event (1) /strata /p58nr /print /p58all. BMDP 2L code: /input file /p58‘c:/p33ex13d4d3.dat’ . variables /p5815. format /p58free. /print cova. Survival. /variable names /p58nr, gt, cs, t1, t2, t3, t4, n1, n2, n3, n4, s1, s2, s3, s4. /form time /p58gt. status /p58cs. response /p581. /regress covariates /p58t1, t2, t3, t4, n1, n2, n3, n4, s1, s2, s3, s4. strata /p58nr. /input file /p58‘c:/p33ex13d4d4.dat’ . variables /p586. format /p58free. /print cova. Survival. /variable names /p58nr, gt, cs, trt, n, s. /form time /p58gt. status /p58cs. response /p581. /regress covariates /p58trt, n, s. strata /p58nr. Anderson--Gill Model ThemodelproposedbyAndersenandGill (1982 ),theAGmodel,assumesthat all events are of the same type and are independent. The risk set in thelikelihood function is totally different from that in the PWP models. The riskset of a person at the time of an event would contain all the people who arestill under observation, regardless of how many events they have experiencedbeforethattime.Themultiplicativehazardfunction h(t,x/p71)fortheithpersonis h(t,x/p71)/p58Y/p71(t)h/p15(t) exp[b/p30x/p71(t)] whereY/p71(t), an indicator,equals 1 whenthe ith person isunder observation (at risk)at timetand 0 otherwise and h/p15(t) is an unspecified underlying hazard368         Table13.13 RearrangedDatafromTable13.7for FittingAGModel ID TL TR CS TRT N S 10 9 0 1 1 22 0 5 9001130 3 1 1 2 63 3 1 401264 0 1 2101141 2 1 6 1 0 1 1 41 6 1 8 0 0 1 1 50 6 1 1 1 15 6 1 2111151 2 1 3 1 1 1 151 3 2 6 0 1 1 160 3 1 0 3 1 6 3 1 51031 61 5 4 6 1 0 3 164 6 5 1 1 0 3 165 1 5 3 0 0 3 1function. The partial likelihood for nindependent persons is L(b)/p58/p76/p147 /p71/p14/p16/p147 /p82/p46/p15/p3Y/p71(t) exp (b/p30x/p71) /p26/p76/p72/p14/p16Y/p72(t) exp (b/p30x/p72)/p4/p66/p71/p7/p82/p8(13.4.6 ) where /afii9829/p71(t)/p581 if theith person has an event at tand/p580 otherwise. Details of this likelihood function and the estimation of the coefficients can be foundin Fleming and Harrington (1991 )and Andersen et al. (1993 ). Similar to the PWP models, software packages are available to carry out the computationprovidedthatthedataarearrangedinacertainformat.Thefollowingexampleillustrates the terms in (13.4.6 )and the data format required by SAS. Example 13.7 We use again the data in Table 13.6 to fit the AG model. To explain the terms in the likelihood function, we use the data of the sixpeople in Table 13.7. In this model, every recurrent event is considered to beindependent.Therefore,wecanrearrangethedatabypersonandbyeventtime‘‘within’’ an individual. Table 13.13 shows the rearranged data. For example,the person with ID /p584 had two recurrences, at 12 and 16, and the follow-up timeendedat 18. Thetimeintervals (TL,TR] are (0, 12], (12, 16], and (16,18], and 12 and 16 are uncensored observations and 18 censored, since there wasno tumor recurrence at 18. For patients with ID /p581 and 2 (i/p581,2), the respectivesecond product terms in (13.4.6 )are equal to 1 since /afii9829/p71(t)/p580,i/p581, 2, for allt. For patient 3 (i/p583),/afii9829/p71(t)/p581 only att/p583(the first tumor recurrence time of the patient ). Thus, the respective second product has only   369 Table13.14 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom theFittedAGModel 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper TRT /p570.412 0.200 4.241 0.0395 0.663 0.448 0.980 N 0.164 0.048 11.741 0.0006 1.178 1.073 1.293 S /p570.041 0.070 0.342 0.5590 0.960 0.836 1.102 one term att/p583 and the denominator of this term sums over all the patients who are under observation and at risk at time t/p583. From Figure 13.1 it is easily seen that the sum is over all six patients; that is, the respective secondproduct is exp(b/p30x/p18) /p26/p21/p72/p14/p16exp(b/p30x/p72)(13.4.7 ) For patient 4 (i/p584), the second product in (13.4.6 )contains two terms. One is fort/p5812(the first recurrence time ), and att/p5812, patients 2, 3, 4, 5, and 6 are still under observation, and therefore the denominator of the term sumsover patients 2 to 6. The other term is for t/p5816(the second recurrence time ) and the denominator sums over patients 2, 4, 5, and 6. Patient 3 is no longerunder observation after t/p5814. Thus, the second product term for i/p584i s exp(b/p30x/p19) /p26/p21/p72/p14/p17exp(b/p30x/p72) /p59exp(b/p30x/p19) exp(b/p30x/p17)/p59/p26/p21/p72/p14/p19exp(b/p30x/p72)(13.4.8 ) Similarly, we can construct each term in (13.4.6 )and the partial likelihood function. Using SAS, we obtain the results in Table 13.14. The AG model identifies treatment and number of initial tumor as significant covariates. Comparedwith placebo, thiotepa does slow down tumor recurrence. ReaderscanconstructtheSAScodesfortheAGmodelbyusingTable13.13 and by following the codes given in Example 13.6. Wei et al. Model By using a marginal approach, Wei, Lin, and Weissfeld (1989 )proposed a model,theWLW model,forthe analysisof recurrentfailures. Thefailures maybe recurrences of the same kind of event or events of different natures,depending on how the stratification is defined. If the strata are defined by the370         times of repeated failures of the same type, similar to the strata defined in the PWP models, it can be used to analyze repeatedfailures of the same kind. Thedifference between the PWP models and the WLW model is that the latterconsiders each event as a separate process and treats each stratum-specific(marginal )partial likelihood separately. In the stratum-specific (marginal ) partial likelihood of stratum s, people who have experienced the (s/p571)th failurecontributeeitheroneuncensoredoronecensoredfailuretimedependingon whether or not they experience a recurrence in stratum s, and the other subjects contribute only censored times (forced as censored times ). Therefore, each stratum contains everyone in the study. This is different from the PWPmodels, in which subjects who have not experienced the (s/p571)th failure are not included in stratum s. If the strata are defined by the type of failure, the WLW model acts like the competing risks model defined in Section 13.3, andthe type-specific (marginal )partial likelihood for the jth type simply treats all failures of types other than jin the data as censored. For thekth stratum of the ith person, the hazard function is assumed to have the form h/p73/p71(t)/p58Y/p73/p71(t)h/p73/p15(t) exp (b/p30/p73x/p73/p71),t/p460 (13 .4.9) whereY/p73/p71(t)/p581, if theith person in the kth stratum is under observation, 0, otherwise,h/p73/p15(t) is an unspecified underlying hazard function. Let R/p73(t/p73/p71) denote the risk set with people at risk at the ith distinct uncensored time t/p73/p71in thekth stratum. Then the specific partial likelihood for the kth stratum is L/p73(b/p73)/p58/p76/p147 /p71/p14/p16 /p3exp(b/p30/p73x/p73/p71) /p26l/p43R/p73(t/p73/p71)exp(b/p30/p73x/p73/p74)/p4/p66/p71(13.4.10 ) where /afii9829/p71/p581 if theith observation in the kth stratum is uncensored and 0 otherwise. The coefficients b/p73are stratum specific. In practice, if we are interested in the overall effect of the covariates, we can assume that thecoefficients from different strata are equal (provided that there are no qualitat- ive differences among the strata ), combine the strata and draw conclusions above the ‘‘average effect’’ of the covariates. We again called the coefficients ofthese covariates common coefficients. The event time is from the beginning ofthe study in this model. Similarto the PWP andAG models,the data must be arranged in a certain format in order to use available software to carry out estimation of thecoefficients and tests of significance of the covariates. Using the same data asin Examples 13.6 and 13.7, the following example illustrates the terms in thestratum-specific likelihood function and the use of software. Example 13.8 First, we use the same six patients to illustrate the compo- nents in the stratum-specific likelihood function in (13.4.10 ). The format the data have to be in for the available software, such as SAS, SPSS, and BMDP,   371 Table13.15 RearrangedDatafromTable13.7forFittingWLWModelwith NR-IndexedCoefficients ID NR TR CS T1 T2 T3 T4 N1 N2 N3 N4 S1 S2 S3 S4 3 1311000200060006 1310000300010005 161100010001000 1 190100010002000 4 11 21000010001000 2 15 90000010001000————————————————————————————————————————————1 290010001000200 5 21 21010001000100 3 21 400100020006006 21 510000030001004 21 610000010001002 25 90000001000100————————————————————————————————————————————1 390001000100020 5 31 310010001000103 31 400010002000604 31 80000000100010 6 34 61000000300010 2 35 90000000100010————————————————————————————————————————————1 490000100010002 3 41 40000100020006 4 41 800000000100015 42 600001000100016 45 110000000300012 45 90000000010001 is similar to that in the PWP and AG models except that all six people are in each of the four strata (Table 13.15 ). The first stratum (NR/p581)is exactly the sameasinTable 13.8.Thesix patientsareorderedaccordingto themagnitudeoftheeventtime (censoredornot,TR ).Instratum2(NR /p582),thethreepeople (with ID /p584, 5, and 6 )whose times to the second tumor recurrence are uncensored observations. Patients 1 and 2 had censored time at 9 and 59,respectively. Patient 3, who had no second recurrence and was observed until14 months, is considered censored at 14. The other strata are constructed in asimilarmanner.UsingthedataarrangementinTable13.15,wecanseethatforthesecondstratum,the likelihoodfunctionin (13.4.10 )hasthreeterms,onefor each of persons 5, 6, and 4, whose /afii9829/p71/p581(CS/p581 in the table ). For patient 4, the risk set at time t/p5816 has two individuals (ID/p584 and 2 ); for patient 5, the risk set at time t/p5812 contains five individuals (ID/p582, 3, 4, 5, and 6 ); and for patient 6, the risk set at time t/p5815 has three individuals (ID/p582, 4, and 6). Letx/p17/p72be the covariatevector of the patient with ID /p58jin stratum2; then372         Table13.16 RearrangedDatafromTable13.7for FittingWLWModelwithCommonCoefficients ID NR TR CS TRT N S 31 3 1 1 2 661 3 1 0 3 151 6 1 1 1 1 11 9 0 1 1 2 4 11 21011 2 15 90011————————————————————————————12 9 0 1 1 2 5 21 21111 3 21 401266 21 510314 21 610112 25 90011————————————————————————————13 9 0 1 1 2 5 31 311113 31 401264 31 80011 6 34 61031 2 35 90011————————————————————————————14 9 0 1 1 2 3 41 40126 4 41 800115 42 601116 45 110312 45 90011 the likelihood function in (13.4.10 )is L/p17(b/p17)/p58/p21/p147 /p71/p14/p16/p3exp(b/p30/p17x/p17/p71) /p26l/p43R/p17(t/p17/p71)exp(b/p30/p17x/p17/p74)/p4/p66/p71/p58exp(b/p30/p17x/p17/p19) exp(b/p30/p17x/p17/p17)/p59exp(b/p30/p17x/p17/p19) /p59exp(b/p30/p17x/p17/p20) exp(b/p30/p17x/p17/p17)/p59exp(b/p30/p17x/p17/p18)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p20)/p59exp(b/p30/p17x/p17/p21) /p59exp(b/p30/p17x/p17/p21) exp(b/p30/p17x/p17/p17)/p59exp(b/p30/p17x/p17/p19)/p59exp(b/p30/p17x/p17/p21)(13.4.11 ) Note that (13.4.11 )is different from (13.4.3 ). The likelihood function for the other strata and for the entire data set in Table 13.6 can be constructed in asimilar manner. If we ignore the stratum-specific effect and are interested onlyin the average overall effect of the covariates, we combine T1 —T4, N1—N4, and S1—S4. The rearranged data for the six patients are given in Table 13.16.   373 Table13.17 AsymptoticPartialLikelihoodInferenceontheBladderCancerDatafrom theFittedWLWModelswithStratum-SpecificorCommonCoefficients 95% Confidence Interval Regression Standard Chi-Square Hazards Variable Coefficient Error Statistic pRatio Lower Upper Model with Stratum-Speci fic Coef ficients T1 /p570.526 0.316 2.774 0.0958 0.591 0.318 1.097 T2 /p570.632 0.393 2.588 0.1077 0.531 0.246 1.148 T3 /p570.698 0.460 2.308 0.1278 0.496 0.202 1.225 T4 /p570.635 0.576 1.215 0.2703 0.530 0.171 1.639 N1 0.238 0.076 9.851 0.0017 1.269 1.094 1.472 N2 0.137 0.902 2.229 0.1354 1.147 0.958 1.373 N3 0.174 0.105 2.750 0.0973 1.189 0.969 1.460N4 0.332 0.125 7.112 0.0077 1.394 1.092 1.780 S1 0.070 0.102 0.470 0.4931 1.072 0.879 1.308 S2 /p570.078 0.134 0.337 0.5614 0.925 0.712 1.203 S3 /p570.214 0.183 1.371 0.2416 0.807 0.565 1.155 S4 /p570.206 0.231 0.800 0.3712 0.813 0.517 1.279 Model with Common Coef ficients TRT /p570.585 0.201 8.460 0.0036 0.557 0.376 0.826 N 0.210 0.047 20.230 0.0001 1.234 1.126 1.352S /p570.052 0.070 0.548 0.4592 0.950 0.828 1.089 The results from fitting the WLW models to the entire data set in Table 13.6 are given in Table 13.17. The model with stratum-specific coefficientssuggests that more initial tumors accelerate tumor recurrence and the acceler-ation is particularly faster for the first recurrence and the third and fourthrecurrences. The signs of the coefficients for T1 —T4 suggest that thiotepa may slow down tumor growth, but the evidence is not statistically significant. Themodel with common coefficients suggests that thiotepa is significantly moreeffective in prolonging the recurrence time. The results suggest that whenlooking at each stratum independently, there is no strong evidence thatthiotepais more effectivethan placebo.However,thecombinedestimate of thecommon coefficient provides stronger evidence that thiotepa is more effectiveover the course of the study. 13.5 MODELSFORRELATEDOBSERVATIONS In Cox’s proportional hazards model and other regression methods, a key assumptionisthat observedsurvivalor eventtimes are independent.However, inmanypracticalsituations,failuretimesareobservedfromrelatedindividuals374         orfromsuccessiverecurrenteventsorfailuresofthesameperson.Forexample, in an epidemiological study of heart disease, some of the participants may befrom the same family and therefore are not independent. These families withmultiple participants may be called clusters. In this case, the regression methodsweintroducedearliermaynotbeappropriate.Severaltypesofmodelsintroduced especially for related observations are discussed by Andersen et al.(1993 ), Liang et al. (1995 ), Klein and Moeschberger (1997 ), and Ibrahim et al. (2001 ). Details about these models are beyond the scope of this book. In the following, we introduce briefly the frailty models. Thefrailty models assume that there is an unmeasured random variable (frailty )inthehazardfunction.Thisrandomvariableaccountsforthevariation or heterogeneity among individuals in a cluster. It is also assumed that thefrailtyis independent of censoring.Let nbe the total numberof participants in the study, some of them related and forming clusters. Let v/p71be the unknown random variable, frailty, associated with the ith cluster, 1 /p45i/p58n. The frailty model associated with the proportional hazards model can be written in termsof the log hazard function as log[h/p71/p72(t;x/p71/p72/p34v/p71)]/p58log[h/p15(t)]/p59v/p71/p59b/p30x/p71/p72(13.5.1 ) for 1/p45j/p45m/p71and 1 /p45i/p58n, wherebdenotes the p/p591 column vector of unknown regression coefficients, x/p71/p72is the covariate vector of the jth person in theith cluster,m/p71is the number of individuals in the ith cluster, and h/p15(t)i sa n unknown underlying hazard function. Compared with the Cox proportionalhazards model, the difference here is the random effect v/p71. Becausev/p71remains the same in the ith cluster, the association between failure and covariates within each cluster in this model is assumed to have a symmetric pattern. In afamily study, this model can be used, for example, to model failure timesobserved from siblings by treating each family as a cluster. This model wasproposed by Vaupel et al. (1979 )and developed and discussed by many researchers, including Clayton and Cuzick (1985 ). The main approach to this model is to assume that v/p71follows a parametric distribution. The frailty model in (13.5.1 )can be extended to handle more complicated situations. For example, the frailty can be a time-dependent variable [replacev/p71byv/p71(t)i n (13.5.1 )]. The frailty model with v/p71(t) can be used to model successive or recurrent failure time as an alternative to the models in Section13.4. Another example is that there may be more than one type of frailty ineach cluster, and v/p71in(13.5.1 )can be replaced by v/p71/p59u/p71orv/p71/p59u/p71/p59w/p71, and so on. Inferences of these frailty models are also based on either a likelihood functionorapartiallikelihoodfunction.Sincethemodelsinvolveaparametricdistribution,the likelihood or partial likelihood functions are complicatedandare beyond the level of this book. The frailty models have not been used widely primarily because of the lack of commercially available software. There are some computer programs    375 available; for example, a SAS macro is available for a gamma frailty model at the Web site of Klein and Moeschberger (1997 ), and another program is described by Jenkins (1997 ). BibliographicalRemarks Most of the major references for nonproportional hazards models have been cited in the text of this chapter. Applications of these models include: stratified models: Vasan et al. (1997 ), Aaronson et al. (1997 ), and Yakovlev et al. (1999 ); frailtymodels :Yashin andIachine (1997 ),Kessinget al. (1999 ),Siegmundet al. (1999 ),Albert (2000 ),LeeandYau (2001 ),Wienkeelal. (2001 ),andXue (2001 ); competingrisksmodels : Mackenbach et al. (1995 ), Fish et al. (1998 ), Albertsen et al. (1998 ), Blackstone and Lytle (2000 ), Yan et al. (2000 ), and Tai et al. (2001 ). EXERCISES 13.1Consider the cancer-free times from the participants with IDs 15 to 23 in Table 13.1. Follow Example 13.1 to construct the partial likelihoodfunctionbasedon the observed cancer-free times from these nine partici-pants. 13.2Considerthesurvivaltimesfrom30resectedmelanomapatientsinTable 3.1.LetAGEGdenoteagegroup,AGEG /p581ifage /p5845andAGEG /p582 otherwise. Fit the survival times with an AGEG-stratified Cox propor-tional hazards model with the covariates age, gender, initial stage, andtreatmentreceived.Discusstheassociationofthetreatmentreceivedwiththe survival time. 13.3Using the data in Table 12.4, following Example 13.5 and the sample codes for SAS, SPSS, or BMDP, fit the competing risk model for stroke,CHD,other CVD, or STROKE/CHDseparately, and discuss the resultsobtained. 13.4Using the rearranged data in Tables 13.7 to 13.13 and following Examples 13.6 to 13.8, complete construction of the remaining terms inthe partial likelihood function based on the PWP model (13.4.2 ), PWP gaptimemodel (13.4.3 ),and AG model (13.4.9 ), and the remaining three marginal likelihood functions based on the WLW model (13.4.13 ).376         CHAPTER 14 Identification of Risk Factors Related to Dichotomous andPolychotomous Outcomes In biomedical research we are often interested in whether a certain survival-relatedevent willoccur andthe importantfactorsthat influenceits occurrence.Such events may involve two or more possible outcomes; examples are thedevelopment of a given condition and response to a given treatment. If thegiven condition is diabetes and we are only interested in whether someonedevelops the disease (yes or no ), the outcome is binary or dichotomous. If we are interested in whether the person develops impaired glucose tolerance,diabetes, or remains having normal glucose tolerance, there are three possibleoutcomes,orwesaytheoutcomeis trichotomous .Similarly,responsetoagiven treatment can have dichotomous (response or no response )orpolychotomous outcomes (complete response, partial response, or no response ). To determine whether one is likely to develop a given disease, we need to know the important characteristics (or factors )related to its development. High- and low-risk groups can then be defined accordingly. Factors closelyrelated to the development of a given disease are usually called risk factors or risk variables by epidemiologists. We shall use these terms in a broader sense to mean factors closely related to the occurrence of any event of interest. Forexample, to find out whether a woman will develop breast cancer because oneof her relatives did, we need to know whether a family history of breast canceris an important risk factor. Therefore, we need to know the following: 1. Of age, race, family history of breast cancer, number of pregnancies, experience of breast-feeding, and use of oral contraceptives—which aremost important? 2. Can we predict, on the basis of the important risk factors, whether a woman will develop breast cancer or is more likely to develop breastcancer than another person? 377 In this chapter we introduce several methods for answering these ques- tions. The general approach is to relate various patient characteristics(or independent variables, or covariates )to the occurrence of an event (dependent or response variable )on the basis of data collected from patients in each of the outcome groups. In the case of dichotomous out-comes, there are two outcome groups. For example, to relate variables suchas age, race, and number of pregnancies to the development of breast cancer,we need to collect information about these variables from a group of breast cancerpatientsaswellasfromagroupofhealthynormalwomen.Foraneventwith polychotomous outcomes, we need to collect data from each outcomegroup. Often, a large number of patient characteristics deserve consideration. These characteristics may be demographic variables such as age; geneticvariables such as gene variant or phenotype; behavioral variables such assmoking or drinking behavior and use of estrogen or progesterone medic-ation; environmental variables such as exposure to sun, air pollution, oroccupational dust; or clinical variables such as blood cell counts, weight, andblood pressure. The number of possible risk factors can be reduced throughmedical knowledge of the disease and careful examination of the possible riskfactors individually. In Section 14.1 we present two methods for examination of individual variables. One is to compare the distribution of each possible risk variableamong the outcome groups. The other method is the chi-square test for acontingency table. This test is particularly useful when the risk variables arecategorical: for example, dichotomous or trichotomous. In this case, a 2 /p59cor r/p59ccontingency table can be set up and a chi-square test performed. In Section 14.2 we discuss logistic, conditional logistic, and other regressionmodels for binary responses and for examining the possible risk variablessimultaneously. Models for multiple outcomes are discussed in Section 14.3. 14.1 UNIVARIATE ANALYSIS 14.1.1 Comparing the Distributions of Risk Variables Among Groups When the outcome is binary, it is often convenient to call an observation a success or a failure. Successmay mean that a survival-related event occurred, andfailurethat it failed to occur. Thus, a success may be a responding patient, a patient who survives more than five years after surgery, or aperson who develops a given disease. A failure may be a nonrespond-ing patient, a patient who dies within five years after surgery, or a personwho does not develop a given disease. A preliminary examination of thedata can compare the distribution of the risk variables in the success andfailure groups. This method is especially appropriate if the risk variable is378         Table 14.1 Ages of 71 Leukemia Patients (Years) Responders 20, 25, 26, 26, 27, 28, 28, 31, 33, 33, 36, 40, 40, 45, 45, 50, 50, 53 56, 62, 71, 74, 75, 77, 18, 19, 22, 26, 27, 28, 28, 28, 34, 37, 47, 56, 19 Nonresponders 27, 33, 34, 37, 43, 45, 45, 47, 48, 51, 52, 53, 57, 59, 59, 60, 60, 61, 61, 61, 63, 65, 71, 73, 73, 74, 80, 21, 28, 36, 55, 59, 62, 83 Source:Hart et al. (1977 ). Data used by permission of the author. continuous. If, for example, the risk factor xis weight and the dependent variable yis having cardiovascular disease, we may compare the weight distribution of patients who have developed disease to that of disease-freepatients. If the disease group has significantly higher weights than those of thedisease-freegroup,wemayconsiderweightanimportantriskfactor.Common-ly used statistical methods for comparing two distributions are the t-test for two independent samples if the assumption of normality holds and theMann —Whitney U-test if the normality assumption is violated and a non- parametric test is preferred. Similarly,if there are more than two possibleoutcomes, we canuse analysis of varianceor the Kruskal —Wallisnonparametrictest to compare the multiple distributionsofacontinuousvariable.Thefollowingexamplecomparestheagedistribution of responders with that of nonrespondersin a cancer clinical trial. Example 14.1 Consider the ages of 71 leukemia patients—37 responders and 34 nonresponders (response is defined as a complete response only )— given in Table 14.1. Figure 14.1 gives us the estimated age distributions of the two groups. By using the Mann —Whitney U-test (or Gehan’s generalized Wilcoxon test ), we find that the difference in age between responders and nonresponders is statistically significant (p/p580.01). In consequence, a question may arise as to what age is critical. Can we say that patients under 50 mayhave a better chance of responding than do patients over 50? To answer thisquestion, one can dichotomize the age data and use the chi-square test,discussed next. 14.1.2 Chi-Square Test and Odds Ratio The chi-square test and the odds ratio are most appropriate when the independentvariableiscategorical.Iftheindependentvariableisdichotomous,a2/p592 table can be used to represent the data. Any variables that are not dichotomous can be made so (with a loss of some information )by choos- ing a cutoff point: for example, age less than 50 years. For multiple-outcome events, 2 /p59corr/p59ctables can be constructed. The independent  379 Figure 14.1 Age distribution of responders and nonresponders. variables are then examined to find which ones (in some sense )provide the best risk associations with the dependent variable. We first consider binaryoutcomes and independent variables that have two categories; that is, we setupa2/p592contingencytablesimilartoTable14.2foreachindependentvariable and look for a high degree of proportionality. The first step is to calculate the sample proportion of successes in the two risk groups, a/C/p16andb/C/p17. Further analysis of the table is concernedwith the precision of these proportions. A standard chi-square test can be used.380         Table 14.2 General Setup of a 2 /p592 Contingency Table Risk Factor Present ( E) Absent (E/p16)Total Dependent variable Success ab R/p16Failure cd R/p17Total C/p16C/p17N Proportion of successes (success rate ) a/C/p16b/C/p17 If the rates of success for the two groups EandE/p16are exactly equal, the expectednumber of patients in the ijth cell (ith row and jth column )is E/p71/p72/p58N/p59R/p71N/p59C/p72N/p58R/p71/p59C/p72N(14.1.1 ) For example, in the top left cell, the expected number is E/p16/p16/p58R/p16/p59C/p16N since the overall success rate is R/p16/Nand there are C/p16individuals in the E group.Similarexpectednumberscanbe obtainedfor eachof thefourcells. LetO/p71/p72be the number of patients observed in the ijth cell. Then the discrepancies canbemeasuredbythedifferences( O/p71/p72/p57E/p71/p72). Inaroughsense,thegreaterthe discrepancies, the more evidence we have against the null hypothesis that thesuccess rates are the same for the two groups. The chi-square test is based on these discrepancies. Let X/p17/p58/p17/p26 /p71/p14/p16/p17/p26 /p72/p14/p16(O/p71/p72/p57E/p71/p72)/p17 E/p71/p72 (14.1.2) Underthenullhypothesis, X/p17followsthechi-squaredistributionwith1 degree of freedom (df). The hypothesis of equal success rates for groups EandE/p16is rejected if X/p17/p57/afii9851/p17/p16/p11/p63, where /afii9851/p17/p16/p11/p63is the 100 /afii9825percentage point of the chi-square distribution with 1 degree of freedom. An alternative way to compute X/p17is X/p17/p58(ad/p57bc)/p17N R/p16R/p17C/p16C/p17(14.1.3)  381 Theodds ratio (Cornfield,1951 )is a commonlyusedmeasureofassociation in2/p592 tables.The odds ratio (OR)is theratiooftwoodds:theoddsof success when the risk factor is present and the odds of success when the risk factor isabsent. In terms of probabilities, OR/p58P(success /p34E)/P(failure /p34E) P(success /p34E/p16)/P(failure /p34E/p16)(14.1.4 ) Using the notation in Table 14.2, P(success /p34E)andP(failure /p34E)may be estimated by a/C/p16andc/C/p16, respectively. Similarly, P(success /p34E/p16)and P(failure /p34E/p16)may be estimated, respectively, by b/C/p17andd/C/p17. Therefore, the numerator and denominator of (14.1.4 )may be estimated, respectively, by a/C/p16c/C/p16/p58a c and b/C/p17d/C/p17/p58b d Consequently, the OR may be estimated by OR/p19/p58a/c b/d/p58ad bc(14.1.5 ) which is also referred to as the cross-product ratio . Several methods are available for an interval estimate of OR: for example, Cornfield (1956 )and Woolf (1955 ). Cornfield’s method, which requires an iterative procedure, is considered more accurate but more complicated thanWoolf’smethod.WoolfsuggestsusingthelogarithmofOR.Thestandarderrorof log OR /p19may be estimated by SE/p19(logOR/p19)/p58/p11 a/p591 b/p591 c/p591 d/p2/p16/p30/p17(14.1.6) Then a 100 (1/p57/afii9825)%confidence interval (CI)for log OR is logOR/p19/p60Z/p63/p30/p17SE/p19(logOR/p19) The confidence interval for OR can be obtained by taking the antilog of the confidence limits for log OR. If logOR/p51and logOR/p42are the upper and lower382         confidence limits for logOR, elogOR/p51andelogOR/p42are the upper and lower confidence limits for OR. Notice that in (14.1.5 ),i fborcis zero, OR /p19is undefined. If any one of the four cell frequencies is zero, the estimated standard error in (14.1.6 )is also undefined. Should this occur, some statisticians (Haldane, 1956; Fleiss, 1979, 1981 )suggest that 0.5 be added to each cell before using (14.1.5 )and (14.1.6 ) to solve the computational problem. However, if the cell frequencies are assmall as zero, the addition of 0.5 to each cell will substantially affect the resulting estimate of OR and its standard error (Mantel, 1977; Miettinen, 1979 ). The estimates so obtained must be interpreted with caution. An odds ratio of 1 indicates that the odds of success are the same whether or not the risk factor is present. An odds ratio greater than 1 means that theodds in favor of success is higher when the risk factor is present, and thereforethereis a positiveassociationbetween the riskfactor andsuccess. Similarly,anodds ratio of less than 1 signifies a negative associationbetween the risk factorand success. The interpretation should not be based totally on the pointestimate. A confidence interval is always more meaningful, just as in any otherestimation procedure. The chi-square statistic in (14.1.2 )may be used to test the null hypothesis thatthereis no associationbetweentheriskfactorandsuccess,or H/p15:OR /p581. The following example illustrates the chi-square test and odds ratio. Example 14.2 In the study of the response rate of 71 leukemia patients (Example 14.1 ), age is considered one of the possible risk variables. The following 2 /p592 table is constructed. Age/p5850 Age /p4650 Total Response 27 10 37 Nonresponse 12 22 34 Total 39 32 71 The question is whether the response rates in the two age groups differ significantly or whether age is associated with response. TheX/p17value according to (14.1.3 )is X/p17/p58(594/p57120)/p17(71) (37)(34)(39)(32)/p5810.16 with 1 degree of freedom. Referenceto Table B-2 shows that the probability of  383 gettinga X/p17valueof10.16ifthetworesponseratesareequalinthepopulation is less than 0.01. Hence the difference between the two response rates issignificant at the 1%level. The estimate odds ratio, according to (14.1.5 ),i s OR/p58(27)(22) (10)(12)/p584.95 The data show that the odds in favor of response are almost five times higher in patients under 50 years of age than in patients at least 50 years old. Thedifference is significantly different, as indicated by the chi-square test above. To obtain a confidence interval for OR, we first compute log OR/p19/p581.60. The estimated standard error of log OR /p19following (14.1.6 )is SE/p19(logOR/p19)/p58/p11 27/p591 10/p591 12/p591 22/p2/p16/p30/p17/p580.515 A95%confidenceinterval for log OR is 1.60 /p601.96(0.515 ),or(0.59, 2.61 ), and a 95% confidence interval for OR is ( e/p15/p13/p20/p24,e/p17/p13/p21/p16),o r (1.80, 13.60 ). The wide interval may be due to the small cell frequencies. Note that the standard errorof log OR/p19is inversely related to the cell frequencies. In this example, the cutoff point, 50, was chosen arbitrarily. It is often of interesttotrymorethan onecutoff pointif thenumber ofobservationsineachcell is not too small. There are cases where the independent variable has c/p572 classes. The chi-squaretest can be extendedto 2 /p59ctables.The odds ratiomethod canalso be extended to handle polychotomous independent variables. It is done byselecting one of the classes as the reference class (theE/p16group )and calculating the measure of association of each of the other classes relative to the referenceclass. Formultiple-outcomeevents,the chi-square testcan be extended to r/p59c tables. The expected frequencies are computed just as in (14.1.1 ), and compu- tation of X/p17[chi-square distributed with (r/p571)(c/p571)degrees of freedom] is thesame asin (14.1.2 )exceptthat the sum is over all r/p59ccells.Fordetails, see Snedecor and Cocharan (1967, Sec. 9.7 ). The following example illustrates the procedures. Example 14.3 Suppose that in the study of response rates of leukemia patients, another possible risk variable is the marrow absolute leukemic infil- trate,whichis definedas the percentageofthe total marrowthat is eitherblast cellsorpromyelocytes.Itisbelievedthatpatientsshouldbeclassifiedintothreeclasses: /p4545%,46 —90%,and /p5790%.The 2 /p593 table is given below. Numbers in parentheses are expected frequencies. For example, 18.68 /p58(39)(34)/71.384         Marrow Absolute Infiltrate /p4545% 46 —90% /p5790% Total Response 4 (8.34)20(20.32 )13(8.34) 37 Nonresponse 12 (7.66)19(18.68 )3(7.66) 34 Total 16 39 16 71 Response rate, (%)25 51 81 OR/p19 1 3.16 13.0 95%CI for OR (0.86, 11.52 )(2.40, 70.46 ) Thequestioniswhetherthedifferencein marrowabsoluteleukemicinfiltrateis related to response. The value of X/p17is X/p17/p58(4/p578.34)/p17 8.34/p59(12/p577.66)/p17 7.66/p59/p37/p59(3/p577.66)/p17 7.66/p5810.17 Thenumberofdegreesoffreedomis3 /p571/p582.With X/p17/p5810.17and2degrees of freedom, the probability that the three absolute infiltrate groups have thesame responserate is less than0.01. The datasuggest that patientswith a highpercentage of marrow absolute infiltrate tend to have a high response rate.Marrow absolute infiltrate may be an important factor in predicting response. The OR /p19s given in the table above are calculated using the /p4545% class as thereference class (or group ). Forexample,for the /p5790%class, the oddsratio is 13/p5912/4/p593/p5813. The 95% confidence intervals for the ORs are obtained using (14.1.6 ). Although the odds ratio for the 46 —90%group is larger than 1, the 95% confidence intervals covers 1. Therefore, the point estimate, 3.16,cannot be taken too seriously. It appears that the major difference is betweenthe/p5790%and /p4545%groups. Individual examination of each independent variable can provide only a preliminary idea of how important each variable is by itself. The relativeimportance of all the variables has to be examined simultaneously usingmultivariate methods. In the following section we discuss the linear logisticregression analysis. 14.2 LOGISTIC AND CONDITIONAL LOGISTIC REGRESSION MODELS FOR DICHOTOMOUS RESPONSES 14.2.1 Logistic Regression Model for Prospective Studies In a typical prospective study, a random sample of subjects is taken and the valuesoftheindependentvariablesaremeasuredatagiventime (usuallycalled baseline measurements ). The subjects are then followed for a given period of      385 time and the outcome (dependent )variable is measured at the end of the follow-up. Therefore, for a prospective study, the independent variables areregardedasfixedquantitiesduringthefollow-up,buttheoutcomesarerandomand unknown. The purpose of a prospective study is to examine the outcomesandrelate them to the baselinemeasurements.Examples of prospectivestudiesare cohort epidemiologic studies and clinical trials. Suppose that there are nsubjects andto some of whomthe event of interest occurred. They are called successes; the others are failures. Let y/p71/p581 if the ith subject is a success and y/p71/p580 if the ith subject is a failure. Suppose that for each of the nsubjects, pindependent variables x/p71/p16,x/p71/p17,...,x/p71/p78are measured. These variables can be either qualitative, such as gender and race, or quanti-tative, such as blood pressure and white blood cell count. The problem is torelate the independent variables, x/p71/p16,...,x/p71/p78, to the dichotomous dependent variable y/p71. LetP/p71be the probability of success, P/p71/p58P(y/p71/p581/p34x/p71/p16,...,x/p71/p78), for the ith subject. The logistic regression model, proposed by Cox (1970 )assumes that the dependence of the probability of success on independent variables is P/p71/p58P(y/p71/p581/p34x/p71)/p58exp(/p26/p78/p72/p14/p15b/p72x/p71/p72) 1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72) (14.2.1) and 1/p57P/p71/p58P(y/p71/p580/p34x/p71)/p581 1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)(14.2.2) wherex/p71/p58(x/p71/p15/p37x/p71/p78),x/p71/p15/p891, and b/p72are unknowncoefficients.The logarithm of the ratio of P/p71and 1 /p57P/p71is a simple linear function of the x/p71/p72’s. Let /afii9838/p71/p58logP/p711/p57P/p71/p58/p78/p26 /p72/p14/p15b/p72x/p71/p72(14.2.3) /afii9838/p71/p58log[P/p71/(1/p57P/p71)] iscalledthe logistic transform ofP/p71and(14.2.3 )isalinear logistic model. Another name for /afii9838/p71islog odds. Thus, the model relates the independent variables to the logistic transform of P/p71, or log odds. The probability of success P/p71can then be found from (14.2.3 )or(14.2.1 ). In many ways (14.2.3 )is the most useful analog for dichotomous response data of the ordinary regression model for normally distributed data. To estimate the coefficients b/p72’s, Cox suggests the maximum likelihood method. Let y/p16,y/p17,...,y/p76be observations with dichotomous values on n subjects.Thelikelihoodfunctionbasedon the binomialdistributioncontains afactor (14.2.1 )whenever y/p71/p581 and (14.2.2 )whenever y/p71/p580. Thus, the likeli-386         hood function is L(b/p15,b/p16,...,b/p78)/p58/p76/p147 /p71/p14/p16P/p87/p71/p71(1/p57P/p71)/p16/p92/p87/p71 /p58/p147/p76/p71/p14/p16exp(y/p71/p26/p78/p72/p14/p15b/p72x/p71/p72) /p147/p76/p71/p14/p16[1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)]/p58exp(/p26/p78/p72/p14/p15b/p72t/p72) /p147/p76/p71/p14/p16[1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)] (14.2.4) where t/p72/p58/p26/p76/p71/p14/p16x/p71/p72y/p71. The log-likelihood function is l(b/p15,b/p16,...,b/p78)/p58logL/p58/p78/p26 /p72/p14/p15b/p72t/p72/p57/p76/p26 /p71/p14/p16log/p31/p59exp/p1/p78/p26 /p72/p14/p15b/p72x/p71/p72/p2/p4(14.2.5) The maximum likelihood estimates of b/p72’s that maximize the log-likelihood function in (14.2.5 )can be obtained by solving the following pequations simultaneously: t/p72/p57/p76/p26 /p71/p14/p16x/p71/p72exp(/p26/p78/p72/p14/p15b/p72x/p71/p72) 1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72)/p580j/p580, 1,...,p(14.2.6) This can be done by an iterative procedure such as the Newton —Raphson procedure. The second derivative of lin(14.2.5 )is I*/p72/p129/p72/p130/p58/p42/p17l /p42b/p72/p129/p42b/p72/p130/p58/p57/p76/p26 /p71/p14/p16x/p71/p72/p129x/p71/p72/p130exp(/p26/p78/p72/p14/p15b/p72x/p71/p72) 1/p59exp(/p26/p78/p72/p14/p15b/p72x/p71/p72) j/p16/p580,...,p;j/p17/p580,...,p (14.2.7 ) LetI/p72/p129/p72/p130/p58(/p571)I*/p72/p129/p72/p130. Then the estimated inverse of the Imatrix, I/p19/p92/p16, is the asymptotic covariance matrix of the b/p72’s. If we use the notation in Section 7.1 and letb/p19/p58(b/p19/p15,b/p19/p16,...,b/p19/p78)/p30denote the MLE of b, the estimated covariance matrix of the MLE b/p19isV/p19(b/p19)/p58(v/p71/p72)/p58(/p57/p42/p17l(b/p19)//p42b/p42b/p30)/p92/p16 /p58I/p19/p92/p16, where v/p71/p72denotes the ijth element of V/p19(b/p19)or the ijth element of I/p19/p92/p16. The coefficients so obtained indicate the relationships between the variables andthelogoddsinfavorofsuccess.Foracontinuousvariable,thecorrespond-ing coefficient gives the change in the log odds for an increase of 1 unit in thevariable.Foracategoricalvariable,thecoefficientisequaltothelogoddsratio(see Section 14.1 ). An approximate 100 (1/p57/afii9825)%confidence interval for b/p72is b/p19/p72/p60Z/p63/p30/p17/p40v/p72/p72 (14.2.8 ) where Z/p63/p30/p17is the 100 (1/p57/afii9825/2)percentile of the standard normal distribution.      387 To test the hypothesis that some of the b/p72’s are zero, a likelihood ratio test can be used. For example, to test H/p15:b/p72/p580, the log-likelihood ratio test statistic is X/p42/p58/p572[l(b/p19/p15,b/p19/p16,...,b/p19/p72/p92/p16,0 ,b/p19/p72/p62/p16,...,b/p19/p78)/p57l(b/p19/p15,b/p19/p16,b/p19/p17,...,b/p19/p78)] (14.2.9 ) where the first term is the maximized log-likelihood subject to the constraint b/p72/p580. If the hypothesis is true, X/p42is distributed asymptotically as chi-square with 1 degree of freedom. An alternative test for the significance of the coefficients is the Wald test, which can be written as X/p53/p58b/p19/p17/p72v/p72/p72(14.2.10 ) Under the null hypothesis that b/p72/p580,X/p53has an asymptotic chi-square distributionwith1degreeoffreedom.AlthoughtheWaldtestis usedbymany,it is less powerful than the likelihood ratio test (Hauck and Donner, 1977; Jennings, 1986 ). In other words, the Wald test often leads the user to conclude that the coefficient (consequently, the respective risk factor )is not significant when, in fact, it is significant. Similar to earlier discussion of model selection, forward, backward, and stepwise variable selection methods can be used to select the risk factors thatare significantly associated with a dichotomous response. The independentvariables x/p71/p72in this model do not have to be the original variables. They can be any meaningful transforms of the original variables: for example, thelogarithm of the original variable, log x/p71/p72, and the deviation of the variable from its mean, x/p71/p72/p57x/p21/p72. From (14.2.1 )and (14.2.2 ), the logarithm of the odds ratio for ith and kth subjects is logP/p71/(1/p57P/p71) P/p73/(1/p57P/p73) /p58/p78/p26 /p72/p14/p16b/p72(x/p71/p72/p57x/p73/p72)( 14.2.11 ) Thus,an estimate ofthe odds ratio canbe obtainedby replacing b/p72in(14.2.11 ) with its MLE, b/p19/p72. From the estimated regression equation, a predicted probability of success can be computed by substituting the values of the risk factors in the equation.Using these predicted probabilities, a goodness-of-fit test can be performed totest the hypothesis that the model fits the data adequately. Several such testsare available (Lemeshow and Hosmer, 2000 ): for example, the Pearson chi- square test, the Hosmer —Lemeshow (Hosmer and Lemeshow,1980 )test, a test statistic suggested by Tsiatis (1980 ), and the score of Brown (1982 ). In the following, we introduce the Hosmer—L emeshow test.388         Letp/p71be the estimate of P/p71obtained from the fitted logistic regression equation for the ith subject, i/p581,...,n. Thep/p71’s can be arranged in ascending order from smallest to largest. Those probabilities and the correspondingsubjects are then divided into ggroups according to some cutoff points of the probability.For example, let g/p5810 and the cutoff points of the probability be equalto k/10,k/p581,2,...,10. Thus,the firstgroupcontains allsubjectswhose estimatedprobabilitiesare less than or equal to 0.1, the second group containsall subjects whose estimated probabilities are less than or equal to 0.2, and so on. Let n/p73be the number of subjects in the kth group. The estimated expected number of successes for the kth group is E/p73/p58/p76 /p73/p26 /p72/p14/p16p/p72k/p581, 2,...,g The Hosmer —Lemeshow test statistic is defined as C/p58/p69/p26 /p73/p14/p16(O/p73/p57E/p73)/p17 n/p73p/p21/p73(1/p57p/p21/p73)(14.2.12 ) where O/p73is the observed number of successes in the kth group and p/p21/p73is the average estimated probability of the kth group, that is, p/p21/p73/p581 n/p73/p76/p73/p26 /p72/p14/p16p/p72 Under the null hypothesis that the model is adequate, the distribution of Cin (14.2.12 )iswellapproximatedbythechi-squaredistributionwith g/p572degrees of freedom. The test is basically a chi-square test of the discrepancy betweenthe observed and predicted frequencies of success. Thus, a Cvalue larger than the 100 /afii9825percentage point of the chi-square distribution (orpvalue less than /afii9825)indicates that the model is inadequate. Similarto otherchi-squaregoodness-of-fittests,theapproximationdepends ontheestimatedexpectedfrequenciesbeingreasonablylarge.Ifalargenumber(say, far more than 20% )of the expected frequencies are less than 5, the approximation may not be appropriate and the pvalue must be interpreted carefully. If this is the case, adjacent groups may be combined to increase theestimated expected frequencies. However, Hosmer and Lemeshow warn that iffewer than six groups are used to calculate C, the test would be insensitiveand would almost always indicate that the model is adequate. Most statistical software packages provide programs for logistic regression analysis:forexample,SAS (proceduresLOGISTIC,PHREG,andCATMOD ), BMDP (proceduresLRandPR ),andSPSS (proceduresNOMREG,PROBIT, PLUM, and LOGISTIC ). Most of them provide estimates of the coefficients and test statistics, variable selection procedures, and tests of goodness of fit.      389 Example 14.4 In a study of 238 non-insulin-dependent diabetic patients, 10 covariates are considered possible risk factors for proteinuria (the outcome variable ).Thelogisticregressionmethodis usedtoidentifythemostimportant risk factors and to predict the probability of proteinuria on the basis of theserisk factors. The 10 potential risk factors are age, gender (1, male; 2, female ), smoking status (0, no; 1, yes ), percentage of ideal body mass index, hyperten- sion (0, no; 1, yes ), use of insulin (0, no; 1, yes ), glucose control (0, no; 1, yes ), duration of diabetes mellitus (DM)in years, total cholesterol, and total triglyceride. Among the 238 patients, 69 have proteinuria ( y/p71/p581). Using the stepwise procedure in BMDP, it is estimated that at step 1, the model contains only b/p15andb/p19/p15/p58/p570.896 and l(b/p19/p15)in(14.2.5 )is/p57143.292. At step 2, duration of diabetes is added to the model because its maximumlog-likelihood value is the largest among all the covariates. The MLEs of thetwo coefficients are b/p19/p15/p58/p571.467 and b/p19/p16/p58/p570.055, and l(b/p19/p15,b/p19/p16)/p58/p57139.429. Since X/p42/p58/p572[l(b/p19/p15)/p57l(b/p19/p15,b/p19/p16)]/p587.726 which is significant (p/p580.005 ), the duration of DM is related significantly to thechanceofproteinuria.TheHosmer —Lemeshowteststatisticforgoodnessof fit with only duration of DM in the model, C/p589.814 with 8 degrees of freedom, gives a pvalue of 0.278. Atstep3,genderisaddedtothemodelbecauseitsadditionyieldsthelargest maximum log-likelihood value among all the remaining covariates. The maxi-mumlog-likelihoodvalue, l(b/p19/p15,b/p19/p16,b/p19/p17)/p58/p57137.749,b/p19/p15/p58/p571.453,b/p19/p16/p58/p570.060, andb/p19/p17/p58/p570.279. To test if gender is significantly related to proteinuria after duration of DM, we perform the likelihood ratio test X/p42/p58/p572[l(b/p19/p15,b/p19/p16)/p57l(b/p19/p15,b/p19/p16,b/p19/p17)]/p583.360 which is significant at p/p580.067. The stepwise procedure terminates after the third step because no other covariates are significant enough to enter theregression model; that is, none of the other covariates have a pvalue less than 0.15, which is set by the program (BMDP ). If any covariate already in the regressionbecomes insignificantafter some other variables are in, the insignifi-cant variable would be removed. The pvalues for entering and removing a variable can be determined by the user. The default values for entering andremoving a variable are, respectively, 0.10 and 0.15. Thus, the procedureidentifies duration of DM and gender as the two most important risk factorsbasedonthe datagiven. Aquestionthat may be raisedatthis pointiswhetherone should include gender in the equation since its significance level is largerthan the commonly used 0.05. The recommendation is to include it since it iscloseto0.05andsincethe pvalue shouldnotbe theonly basisfordetermining whether a covariate should be included in the model. In addition, theHosmer —Lemeshow goodness of fit test statistic C/p585.036, when gender is390         Table 14.3 Estimated Coefficients for a Linear Logistic Regression Model Using Data from Diabetic Patients Estimated Standard Variable Coefficient Error Coefficient/SE exp (coefficient ) Constant /p571.453 0.264 /p575.504 0.234 Duration of DM 0.060 0.020 2.956 1.062 Sex /p570.279 0.152 /p571.836 0.756 included, yields a pvalue of 0.754. Thus, inclusion of this covariate improves considerably the adequacy of the model. Thus, the final regression equationwith the two significant risk factors is logP/p711/p57P/p71/p58/p571.453/p590.060(duration of DM )/p570.279 (gender ) Table 14.3 gives the details for the estimated coefficients. The signs of the coefficients indicate that male patients and patients with a longer duration of diabetes have a higher chance of proteinuria. Furthermore,for each increase of one year in duration of diabetes, the log odds increase by0.060. Probabilities of proteinuria can be estimated following (14.2.1 ). For example, the probability of developing proteinuria for a male patient who hashad diabetes for 15 years is P/p58e/p92/p15/p13/p23/p18/p17 1/p59e/p92/p15/p13/p23/p18/p17 /p580.303 where /p570.832 is obtained by substituting the values of the two covariates in the fitted equation; that is, /p571.453 /p590.060 (15)/p570.279 (1)/p58/p570.832. Similarly, for a female patient who has the same duration of diabetes, the probability is0.248. In addition to individual variables, interaction terms can be included in the logistic regression model. If the association between an independent variablex/p16and the dependent variable yis not the same in different levels of another variable, x/p17, there is interaction between x/p16andx/p17. To check if there is interaction, one can include the product of x/p16andx/p17in the regression model and test the significance of this new variable. The following example illustratesthe procedure. Example 14.5 It is well known that adriamycin is effective for treating certain types of cancer. It is also well known that adriamycin is highly toxic.      391 Some patients develop congestive heart failure (CHF ), but others who receive asimilardose ofadriamycindo not.In anattempttodetectfactorsthatwouldincrease the risk of developing adriamycin cardotoxicity, various patientcharacteristics of 53 cancer patients were studied. Seventeen of these patientsdeveloped CHF and 36 patients did not. After a careful investigation, it wasfound that the total dose ( z/p16) and percentage decrease in electrocardiographic QRS voltage ( z/p17) are most closely related to CHF. Table 14.4 shows the data and some summary statistics. The following linear logistic regression model with transformed variables z/p16/p58z/p16/p57z/p21/p16andx/p17/p58z/p17/p57z/p21/p17is used: /afii9838/p58logp 1/p57p /p58b/p15/p59b/p16x/p16/p59b/p17x/p17/p59b/p18x/p16x/p17 The stepwise procedure selects percentage decrease in QRS as the most important variable, followed by the total dose (TD)and interaction (TD/p59QRS ). The logistic regression analysis results are given in Table 14.5. The stepwise log-likelihood values given in the last column indicate that onlyQRS is significant since 2 (/p5710.185 /p5933.254 )/p5846.138, which yields a pvalue less than 0.001. Neither the total dose nor the interaction is significant. Suppose that the last three columnsof Table 14.4, Y, Z1, and Z2, are stored in a text data file ‘‘C: /p33EX14d2d2.DAT’’, separated by a space. The following SAS, SPSS, or BMDP codes can be used to obtain the results in Table 14.5. SAS code: data w1; infile ‘c: /p33ex14d2d2.dat’ missover; input y z1 z2; x1/p58z1-517.679; x2/p58z2-26.019; x12/p58x1*x2; run;proc logistic data /p58w1 descending; model y /p58x1 x2 x12/ selection /p58s plcl plrl lackfit; run; SPSS code (forward selection method ): data list file /p58‘c:/p33ex14d2d2.dat’ free / y z1 z2. Compute x1 /p58z1-517.679. Compute x2 /p58z2-26.019. Compute x12 /p58x1*x2. Logistic regression y with x1 x2 x12 /method /p58fstep /print /p58all.392         Table14.4 TotalDoseandPercentDecreasein QRSof 53 Patients Receiving Adriamycin Total Percent Decrease Patient CHF /p63,yDose, z/p16in QRS, z/p17 1 1 435 41 2 1 600 71 3 1 600 51 4 1 540 40 5 1 510 636 1 740 797 1 825 618 1 535 449 1 510 53 10 1 483 27 11 1 460 5312 1 460 6013 1 550 6514 1 540 5815 1 310 41 16 1 500 64 17 1 400 4418 0 440 919 0 600 4220 0 510 1921 0 410 24 22 0 540 /p5724 23 0 575 3924 0 564 3525 0 450 1026 0 570 627 0 480 6 28 0 585 21 29 0 420 1430 0 470 131 0 540 3332 0 585 3333 0 600 4 34 0 570 2 35 0 570 536 0 510 1237 0 470 /p571 38 0 405 4439 0 575 14 40 0 540 /p5710 41 0 500 /p5743 42 0 450 2343 0 520 /p571 (Continued overleaf )      393 Table 14.4 Continued Total Percent Decrease Patient CHF /p63,yDose, z/p16in QRS, z/p17 44 0 495 29 45 0 585 4046 0 450 30 47 0 450 23 48 0 500 12 49 0 540 /p5711 50 0 440 751 0 480 /p5722 52 0 550 2053 0 500 19 Source:Minow et al. (1977 ). /p631, yes; 0, no; z/p21/p16/p58517.679,z/p21/p17/p5826.019. Table 14.5 Linear Logistic Regression Analysis Results of Data in Table 14.4 Estimated Standard Variable Coefficient Error Coefficient/SE Log Likelihood Constant /p573.757 1.576 /p572.384 /p5733.254 QRS 0.254 0.102 2.480 /p5710.185 TD /p570.024 0021 /p571.160 /p579.225 TD/p59QRS 0.001 0.001 0.677 /p578.803BMDP codes for procedure LR: /input file /p58‘c:/p33ex14d2d2.dat’ . variables /p583. format /p58free. /variable names /p58y, z1, z2. /transform x1 /p58z1-517.679. x2/p58z2-26.019. x12/p58x1*x2. /regress depend /p58y. Interval /p58x1, x2, x12. Model /p58x1, x2, x12. Start /p58in, in, in. Move /p580, 0, 0.394         Method /p58mlr. /print cell /p58used. /end When the independent variables are dichotomous or polychotomous, the logistic regression coefficients can be linked with odds ratios. Consider thesimplest case, where there is one independent variable, x/p16, which is either 0 or 1. The linear regression model in (14.2.1 )and (14.2.2 )becomes P(y/p581/p34x/p16)/p58e b/p15/p59b/p16x/p16 1/p59eb/p15/p59b/p16x/p16 P(y/p580/p34x/p16)/p581 1/p59eb/p15/p59b/p16x/p16 Values of the model when x/p16/p580, 1 are P(y/p581/p34x/p16/p580)/p58eb/p15 1/p59eb/p15 P(y/p581/p34x/p16/p581)/p58eb/p15/p59b/p16 1/p59eb/p15/p59b/p16 P(y/p580/p34x/p16/p580)/p581 1/p59eb/p15 P(y/p580/p34x/p16/p581)/p581 1/p59eb/p15/p59b/p16 The odds ratio in (14.1.4 )is OR/p58P(y/p581/p34x/p16/p581)/P(y/p580/p34x/p16/p581) P(y/p581/p34x/p16/p580)/P(y/p580/p34x/p16/p580)/p58eb/p15/p59b/p16 eb/p15/p58eb/p16 and the log odds ratio is log (OR)/p58log(eb/p16)/p58b/p16[this can also be derived directly from (14.2.11 )]. Thus, the estimated logistic regression coefficient also provides an estimate of the odds ratio, that is, OR /p19/p58eb/p19/p16.I f(b/p16/p42,b/p16/p51)is the confidence interval for b/p16, the corresponding interval for OR is ( eb/p16/p42,eb/p16/p51). Example 14.6 Consider the age (x)and response (y)data from the 71 leukemia patients presented in Table 14.6 (Examples 14.1 and 14.2 ). The logistic regression analysis results are given in Table 14.7. Notice thatexp(b/p16)/p58exp(1.5994 )/p584.95, which is equal to theestimateof ORobtainedin Example 14.2 using (14.1.5 ), and the standard error of b/p19/p16is the same as that oflogOR/p19exceptforasmallrounding-offerror.TheconfidenceintervalforOR can also be obtained from the logistic regression analysis results.      395 Table 14.6 Age and Response Data of 71 Leukemia Patients Response x/p5850(1) x/p4650(0) Total Yes(1) 27 10 37 No(0) 12 22 34 Total 39 32 71 Table 14.7 Results of Logistic Regression Analysis of Data in Table 14.6 Estimated Standard Variable Coefficient Error Coefficient/SE exp (coefficient ) Constant ( b/p19/p15) /p570.7885 0 .3814 /p572.067 0 .45 Age(b/p19/p16) 1.5994 0.5156 3.102 4.95 The relationship between the logistic regression coefficient and odds ratio can be extended to polychotomous variables by creating dummy variables (or design variables ). The following example illustrates the procedure. Example 14.7 Consider the data in Example 14.3. The variable marrow absolute infiltrate (MAI )has three levels. As in Example 14.3, the /p4545%level is considered as the reference group. In this case, two design variables will beused and their values are assigned as follows: D/p16/p58/p71 if MAI /p5846—90% 0 otherwiseD/p17/p58/p71 if MAI /p5790% 0 otherwise For MAI, the design variable values and the respective number of responders (Event )and total number of patient in each MAI level (N)are listed in the following table. MAI (%)D/p16D/p17Event N /p454 5 00 41 6 46—90 1 0 20 39 /p5790 0 1 13 16 Usingthese design variables, the logistic regression analysisgives the results in Table 14.8.396         Table 14.8 Results of Logistic Regression Analysis of Data in Example 14.7 and Two Design Variables Estimated Standard Variable Coefficient Error Coefficient/SE exp (coefficient ) Constant /p571.0986 0.5774 /p571.903 0.33 MAI ( D/p16) 1.1499 0.6603 1.742 3.16 MAI ( D/p17) 2.5649 0.8623 2.974 13.00 The coefficient corresponding to D/p16, 1.1499, is the log odds ratio between the46 —90%groupandthe /p4545%group.Theoddsratioisexp (1.1499 )/p583.16, which is exactly equal to the estimate obtained in Example 14.3. Similarly, thecoefficient correspondingto D/p17is the log odds ratio betweenthe /p5790%group and the /p4545%group. The odds ratio obtained from the regression coefficient, 13.00, is the same as that obtained in Example 14.3. The estimated standard error for D/p16, 0.6603, is also the standard error of logOR/p19. A 95%confidence interval for the coefficient is 1.1499 /p601.96(0.6603 ), or(/p570.1443, 2.4441 ), and consequently, a 95% confidence interval for OR is (e/p92/p15/p13/p16/p19/p19/p18,e/p17/p13/p19/p19/p19/p16), or (0.86, 11.52 ), which is identical to that obtained in Example 14.3 using Woolf’s estimate of SE (log OR ). Suppose that the datain Example14.6are arrangedin fourcolumnsfor D/p16, D/p17, Event, and Nas in the table above and are saved in a text data file ‘‘C:/p33EX14d2d4.DAT’’. The values of D/p16,D/p17, Event, and Nin each row are separated by a space. The following SAS, SPSS, or BMDP codes can be usedto obtain the results in Table 14.8. SAS code: data w1; infile ‘c: /p33ex14d2d4.dat’ missover; input d1 d2 event n; run;proc logistic data /p58w1 ; model event/n /p58d1 d2 / plcl plrl; run; SPSS code: data list file /p58‘c:/p33ex14d2d4.dat’ free / d1 d2 event n. Probit event OF n WITH d1 d2 /model /p58logit /log/p582.718 /print /p58all.      397 BMDP code for procedure LR: /input file /p58‘c:/p33ex14d2d4.dat’ . variables /p584. format /p58free. /variable names /p58d1, d2, event, n. /regress count is n. fcount is event. Interval /p58d1, d2. /print cell /p58used. /end For a continuous independent variable, the logistic regression coefficient givesthechangeinlogoddsforanincreaseof1unitinthevariable.Ingeneral,for an increaseof munits in the variable,the log odds ratiois equal to mtimes the logistic regression coefficient. The derivation is left to the reader as anexercise. When more than one independent variable is included in the logistic regressionmodel,eachestimatedcoefficientcanbeinterpretedasanestimateofthe log odds ratio statistically adjusting for all the other variables. Forexample, in Example 14.4, the regression coefficient for the gender variable,/p570.279,is an estimate of the log odds ratio for femalesversus males, adjusting for duration of diabetes. Or the adjusted odds ratio for females versus males isestimatedas exp (/p570.279 )/p580.76; suggesting that female diabetic patients have alowerriskofhavingproteinuriathanthatofmalepatients,afteradjustingforduration of diabetes. This interpretation, commonly used by epidemiologists,is appropriate if the linear relationship between the log odds and the indepen-dent variables holds. Press and Wilson (1978 )compare the logistic regression method to the discriminant analysis and find that if the independent variables are normalwith identical covariance matrices, discriminant analysis is preferred. Underabnormality, the logistic regression method is preferred. In particular, if theindependentvariablesaredichotomous,wecannotexpecttopredictaccuratelythe probability of success with a discriminant function, even with a largeamountofdata.Theirexamplesshowthat thelogisticregressiongivesa highercorrect classification rate. 14.2.2 Logistic and Conditional Logistic Regression Model for Retrospective Studies As mentioned earlier, the logistic regression model defined in (14.2.3 )is originally designed for prospective studies, where a set of covariates or independent variables are measured at a baseline examination on a group ofpeople without the disease of interest. These subjects are then followed for aperiodoftimeanddevelopmentofthediseaseamongthemarerecordedduring398         follow-up. The model can be extended to analyze data from retrospective studies, such as a case —control study. In a case—control study , cases (subjects with the diseaseof interest )andcontrols (subjectswithoutthe disease )are first selected and risk factor data such as exposure variables and other covariatesare collected retrospectively. For example, in a case —control study of lung cancer and cigarette smoking, a group of lung cancer patients and a group ofpeople without lung cancer are selected. Their smoking histories are thencollected along with other risk factors. Therefore, in a case —control study, participants are selected first based on their disease status, and their history ofrisk factor exposures is collected later. The purpose of a case —control study is to estimate the association between the risk factors and the disease understudy.Usingprobabilityterms,wearedealingwiththeprobabilitythattheriskfactors take on certain values given that a person is a case or a control. Wedenote this conditional probability by P(x/p34y), wherexdenote the covariates andythe outcome variable. Using the same notation, the probability of interest in a prospective study is P(y/p34x). Based on conditional probability theory, P(x/p34y)can be written as P(x/p34y)/p58P(y/p34x)P(x) P(y) (14.2.13) where y/p581 for cases and 0 for controls. Thus, P(x/p34y) is a function of P(y/p34x), P(x), and P(y). The likelihood function of the logistic regression model for a retrospective study, similar to (14.2.4 ), is the product of terms in the form of P(x/p34y)in (14.2.13 )for the cases and controls selected [ (P(x/p34y/p581)from a case and P(x/p34y/p580)from a control )]. We introduce here two most widely used approachestothislikelihoodfunction.Oneapproachconsiderstheprobabilityof case/control selection. Since in case —control studies, the cases and controls areselectedfromthepopulationandthelikelihoodfunctionisbasedonsubjectselection, we introduce an indicator variable to denote whether a person isselected ( s/p581) or is not selected (s/p580). Letn/p16andn/p15be, respectively, the numbers of selected cases and controls in the study. The likelihood function is L/p2/p2/p58/p76 /p16/p147 /p71/p14/p16P(x/p71/p34y/p71/p581,s/p581)/p76/p15/p147 /p71/p14/p16P(x/p71/p34y/p71/p580,s/p581) (14 .2.14) Define /afii9843/p16/p58P(s/p581/p34y/p581) to be the probability that a diseased person is selected for the study as a case and /afii9843/p15/p58P(s/p581/p34y/p580)      399 to be the probability that a disease-free person is selected for the study as a control. Assume that the sampling probabilities depend only on disease statusand not on the covariates. Using Bayes’ theorem in probability theory and(14.2.1 ),it canbe shown (thederivationis left to thereaderasan exercise )that the probability that a person is diseased given that he or she has risk factorsxand was selected for the study can be written as P(y/p71/p581/p34x/p71,s/p581)/p58exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72) 1/p59exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72) (14.2.15 ) where b*/p15/p58b/p15/p59log/afii9843/p16/afii9843/p15(14.2.16 ) According to the conditional probability in (14.2.13 ), the first term in the likelihood function in (14.2.14 )is P(x/p71/p34y/p71/p581,s/p581)/p58exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72) 1/p59exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)/p3P(x/p71/p34s/p581) P(y/p581/p34s/p581)/p4(14.2.17) Similarly, the second term in (14.2.14 )fory/p71/p580 can be obtained: P(x/p71/p34y/p71/p580,s/p581)/p581 1/p59exp(b*/p15/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)/p3P(x/p71/p34s/p581) P(y/p580/p34s/p581)/p4(14.2.18) Substituting (14.2.17 )and (14.2.18 )into (14.2.14 ), we obtain the likelihood function for a case —control study: L/p2/p2/p58L(b*/p15,b/p16,...,b/p78)/p59/p76/p147 /p71/p14/p16P(x/p71/p34s/p581) P(y/p34s/p581)(14.2.19) where n/p58n/p15/p59n/p16,L(b*/p15,b/p16,...,b/p78)is the likelihood function for prospective studies in (14.2.4 )except that the intercept term b/p15is replaced by b*/p15in (14.2.16 ).If we assume that the probability distributionof the covariates, P(x), contains no information about the parameters of interest, or P(x)is indepen- dent of the coefficients b/p72, and the selection is independent of x, then maximizing L/p2/p2toobtainestimatesof b/p72isequivalenttomaximizingonly L(b*/p15, b/p16,...,b/p78) since P(y/p581/p34s/p581)/p58n/p16/nandP(y/p580/p34s/p581)/p58n/p15/n. This im- plies that we can use the computer program for prospective studies to analyzecase—control study data except that the intercept term cannot be interpreted meaningfully unless /afii9843/p16and/afii9843/p15are known. In most practical situations, the assumption made above about P(x)is reasonable.Historically,in earlyapplicationsof thelogistic regressionmethod,the covariates were assumed to have multivariate normal distribution. Then400         estimationofthecoefficients, b/p72in(14.2.19 ),wouldinvolvethedistribution P(x) and thus became much more complicated. However, in practice, many of thecovariates are categorical or discrete and are therefore distinctly nonnormal.Thus, it is appropriate to allow P(x)to remain completely arbitrary and use simply L(b*/p15,b/p16,...,b/p78)in(14.2.19 )in case —control studies. Another approach to the likelihood function based on (14.2.13 )is to consider a conditional probability instead. Suppose that n/p16cases and n/p15controls were selected in a case —control study and n/p58n/p16/p59n/p15; letx/p16, x/p17,...,x/p76be the risk factor sets of the nsubjects without specifying which of them pertain to the cases and which to the controls. Then the conditionalprobability that the first n/p16x’s are observed from the n/p16cases and the remainder are from the n/p15controls may be written as /p147/p76 /p16/p71/p14/p16P(x/p71/p34y/p581)/p147/p76/p71/p14/p76/p16/p62/p16P(x/p71/p34y/p580) /p26/p43l/p16,...,l/p76/p129/p44(/p147/p76/p16/p71/p14/p16P(x/p74/p71/p34y/p581)/p147/p76/p71/p14/p76/p16/p62/p16P(x/p74/p71/p34y/p580))(14.2.20) where the summation in the denominator is over the n!/(n/p16!n/p15!)possible ways of selecting n/p16individuals as cases from the nsubjects, with the remaining n/p15as controls. By using (14.2.13 )and (14.2.1 ),(14.2.20 )reduces to /p147/p76/p16/p71/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p71/p72) /p26/p43l/p16,...,l/p76/p129/p44/p147/p76/p16/p71/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p74/p71/p72)(14.2.21 ) Comparing with (12.1.17 ),(14.2.21 )can be considered as a special case of (12.1.17 )in which there is only one distinct uncensored failure time, say at t/p581, and all of the n/p16persons failed at t/p581. The remaining n/p15subjects survive longer than 1 and are censored, say at t/p582, while all nsubjects are at risk at t/p581. This interpretation of (14.2.21 )permits us to apply the method andcomputersoftwarefortheCoxproportionalhazardsmodelwithadiscretetime scale and ties to obtain an estimate of bin the logistic regression model and to perform the corresponding inferences. The procedure is illustrated inExample 14.8. When n/p16andn/p15are large enough, it can be shown that an analysisbasedonthisconditionalprobabilitywillproduceresultsequivalenttothose based on the likelihood function defined in (14.2.19 )(Efron, 1975; Farewell, 1979; Breslow and Day, 1980 ). When n/p16andn/p15are large, one may prefer using (14.2.19 )to(14.2.21 )since the former is easier. However, for a case—control study with a matched or stratified design, analysis based on the conditional probability defined in (14.2.21 )is a better choice than (14.2.19 ) (Efron, 1975; Farewell, 1979; Breslow and Day, 1980 ). In the following section we discuss the application of logistic regression analysis for two widelyaccepted matched designs in case —control studies. 1 : R Matched Design A widely used case —control design is to have one or more controls matched      401 for each case based on matching variables such as age and gender. Suppose that for each case there are R(/p461) matched controls. Let x/p71/p72/p73denote the observed value of the jth covariate ( j/p581,...,p)from the kth subject (k/p581 for the case and k/p582,..., R/p591 for matched controls )in the ith matched set (i/p581,...,n). The nmatched sets are considered as the samples from the n different strata defined by the matching variables. Following (14.2.21 )with n/p16/p581 and n/p15/p58R, the conditionalprobability for the matchedset (1 case and Rcontrols )in the ith stratum is exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p16) exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p16)/p59/p26/p48/p62/p16/p73/p14/p17exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p73)/p581 1/p59/p26/p48/p62/p16/p73/p14/p17exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p73/p57x/p71/p72/p16)] (14.2.22 ) and thus the conditional likelihood function for all nstrata is the product of thenterms in (14.2.22 ), that is, /p76/p147 /p71/p14/p161 1/p59/p26/p48/p62/p16/p73/p14/p17exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p73/p57x/p71/p72/p16)](14.2.23) When R/p581, that is, a one-to-one pair matching, the conditional likelihood function obtained from (14.2.23 )reduces to L(b/p16,...,b/p78)/p58/p76/p147 /p71/p14/p161 1/p59exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p17/p57x/p71/p72/p16)] /p58/p76/p147 /p71/p14/p16exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p16/p57x/p71/p72/p17)] 1/p59exp[/p26/p78/p72/p14/p16b/p72(x/p71/p72/p16/p57x/p71/p72/p17)](14.2.24) Compared with the likelihood function for the ordinary logistic regression in (14.2.4 ),theconditionallikelihoodfunction (14.2.24 )canbetreatedasaspecial case of (14.2.4 )withy/p71/p891,b/p15/p890, and x/p71/p72can be replaced by the difference in x/p71/p72between the case and its matched control. This fact permits the use of computer programs for ordinary logistic regression in one-to-one matchedcase—control studies. The procedure is as follows: 1. Let nbe the number of case —control pairs. 2. Use x/p71/p72/p16/p57x/p71/p72/p17, the difference between covariates for the case (x/p71/p72/p16)and its matched control (x/p71/p72/p17), as the independent variable in the model. 3. Let y/p71/p891 for all pairs. 4. Delete the intercept term b/p15from the model.402         n1:n0Matched Design or Stratified Design Suppose that there are n/p16cases and n/p15controls in the ith stratum and n/p58n/p16/p59n/p15. Let x/p71/p72/p73denote the observed value of the jth covariate (j/p581,...,p)from the kth subject (k/p581,2,..., n/p16for the n/p16cases and k/p58n/p16/p591,...,nfor the n/p15controls )in the ith stratum ( i/p581,...,m). From (14.2.21 ), the contribution of the ith stratum to the conditional likelihood function is /p147/p76/p16/p73/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p73) /p26/p43k/p16,...,k/p76/p129/p44/p147/p76/p16/p74/p14/p16exp(/p26/p78/p72/p14/p16b/p72x/p71/p72/p73/p74)(14.2.25 ) where the summation in the denominator is over all the n!/(n/p16!n/p15!)possible waystoselect n/p16outofthe nsubjectsascasesandtheremaining n/p15ascontrols. The term in (14.2.25 )has the same mathematical form as (14.2.21 ). Thus, as discussed earlier, the computer software for the proportional hazards modelcan be used to estimate the coefficients. In both 1: Randn/p16:n/p15matched designs, most of the other features in the ordinary logistic regression model fitting, including the use of design (dummy ) variables and statistical inferences,remain the same. However, the goodness offittestofHosmerandLemeshowisnotapplicabletomatcheddesigns.Readerswho are interested in assessing the logistic regression model in matchedcase—control studies are referred to Pregibon (1984 )and Moolgavkar et al. (1985 ). The following example illustrates the basic procedure for the one-to-one matched design using (14.2.24 ). Example 14.8 Tostudy theeffect ofobesity, familyhistoryof diabetes,and level of physical activity to non-insulin-dependent diabetes (NIDDM ),3 0 nondiabeticpersons arematchedwith30 NIDDMpatientsbyage andgender.Obesity is measured by body mass index (BMI ), which is defined as weight in kilogramsdividedbyheightinmeterssquared.Familyhistoryofdiabetes (FH) and levels of physical activity (PHY )are binary variables. Table 14.9 gives the partially fictitious data. Following the procedure given above, the results offitting the three variables using BMDP are given in Table 14.10. Suppose that text data file ‘‘C: /p33EX14d2d5.DAT’’ contains six successive columns of data: BMIC, FHC, PHYC, BMIN, FHN, and PHYN, as in Table14.9, separated by a space. The following SAS, SPSS, or BMDP code can beused to generate the results in Table 14.10. SAS code: data w1; infile ‘c: /p33ex14d2d5.dat’ missover; input bmic fhc phyc bmin fhn phyn;bmi/p58bmic-bmin;      403 Table 14.9 Data of 30 Matched Case--Control Pairs Case (Diabetic ) Control (Nondiabetic ) Pair BMI FH /p63PHY /p64BMI FH /p63PHY /p64 1 22.1 1 1 26.7 0 1 2 31.3 0 0 24.4 0 1 3 33.8 1 0 29.4 0 0 4 33.7 1 1 26.0 0 0 5 23.1 1 1 24.2 1 06 26.8 1 0 29.7 0 07 32.3 1 0 30.2 0 18 31.4 1 0 23.4 0 19 37.6 1 0 42.4 0 0 10 32.4 1 0 25.8 0 0 11 29.1 0 1 39.8 0 112 28.6 0 1 31.6 0 013 35.9 0 0 21.8 1 114 30.4 0 0 24.2 0 115 39.8 0 0 27.8 1 1 16 43.3 1 0 37.5 1 1 17 32.5 0 0 27.9 1 118 28.7 0 1 25.3 1 019 30.3 0 0 31.3 0 120 32.5 1 0 34.5 1 121 32.5 1 0 25.4 0 1 22 21.6 1 1 27.0 1 1 23 24.4 0 1 31.1 0 024 46.7 1 0 27.3 0 125 28.6 1 1 24.0 0 026 29.7 0 0 33.5 0 027 29.6 0 1 20.7 0 0 28 22.8 0 0 29.2 1 1 29 34.8 1 0 30.0 0 130 37.3 1 0 26.5 0 0 /p631, yes; 0, no. /p641, physically active; 0, sedentary. fh/p58fhc-fhn; phy/p58phyc-phyn; y/p581; run; proc logistic data /p58w1; model y /p58bmi fh phy / noint plcl plrl lackfit; run;404         Table 14.10 Results of a Logistic Regression Analysis of Data in Table 11.13 Estimated Estimated Variable Coefficient Standard Error Coefficient/SE exp (coefficient ) BMI 0.090 0.065 1.381 1.094 FH 0.968 0.588 1.646 2.633PHY /p570.563 0.541 /p571.041 0.569 SPSS code: data list file /p58‘c:/p33ex14d2d5.dat’ free / bmic fhc phyc bmin fhn phyn. Compute bmi /p58bmic-bmin. Compute fh /p58fhc-fhn. Compute phy /p58phyc-phyn. Compute y /p581. Logistic regression y with bmi fh phy /origin/print /p58all. BMDP LR code: /input file /p58‘c:/p33ex14d2d5.dat’ . variables /p586. format /p58free. /variable names /p58bmic, fhc, phyc, bmin, fhn, phyn. /transform bmi /p58bmic-bmin. fh/p58fhc-fhn. phy/p58phyc-phyn. y/p581. /regress depend /p58y. Interval /p58bmi, fh, phy. Model /p58bmi, fh, phy. Start /p58in, in, in. Constant /p58out. Move /p580, 0, 0. Method /p58mlr. /print cell /p58used. /end Thefollowingexample illustratesthe estimatingproceduresfor the1: Rand n/p16:n/p15matched case —control designs. Example 14.9 Table 14.11 lists a subset of simulated data from a case — control diabetes study that is based on a cohort study of heart disease with a      405 Table 14.11 A Subset of Age-Group and Gender-Matched DM Data in Example 14.9 /p63 AGE AGEG SEX SBP DBP LACR HDL LINSUL SMOKE DMS DM SN 51.8 50 1 148 91 1.35 37 2.81 0 1 0 1 50.9 50 1 116 95 1.04 43 2.74 0 1 0 150.9 50 1 114 85 1.26 64 1.98 1 1 0 150.9 50 1 120 80 1.56 52 2.53 1 1 0 1 54.6 50 1 119 71 1.55 30 2.87 0 2 0 1 50.8 50 1 123 78 1.40 33 3.31 0 2 0 1 53.3 50 1 119 75 1.69 40 2.13 1 3 1 172.0 70 0 129 73 1.06 25 2.69 0 1 0 273.1 70 0 120 68 0.87 30 2.76 0 1 0 272.8 70 0 111 66 2.52 73 3.17 0 2 0 270.3 70 0 115 65 3.16 42 2.96 0 2 0 2 72.1 70 0 140 66 3.18 52 3.48 0 2 0 2 72.8 70 0 136 72 3.36 59 3.22 0 2 0 271.1 70 0 133 85 2.95 73 3.25 0 3 1 256.4 55 1 110 74 0.58 43 2.18 0 1 0 355.7 55 1 122 77 1.18 34 2.76 1 1 0 356.9 55 1 114 74 1.10 25 2.62 0 1 0 3 58.5 55 1 104 74 1.24 23 2.58 0 1 0 3 55.2 55 1 128 77 1.25 43 2.77 0 2 0 355.9 55 1 130 83 1.34 44 2.34 1 2 0 357.4 55 1 116 79 2.23 38 2.46 1 3 1 360.7 60 0 136 85 2.04 42 3.66 0 1 0 462.0 60 0 115 74 1.32 33 2.85 0 1 0 4 64.7 60 0 155 89 2.55 46 3.73 1 1 0 4 64.4 60 0 191 107 3.66 34 2.67 1 2 0 460.5 60 0 109 74 0.89 64 2.88 0 2 0 462.4 60 0 106 72 0.88 35 3.44 1 2 0 462.8 60 0 234 91 7.28 49 2.38 1 3 1 473.2 70 0 119 72 1.00 47 2.53 0 1 0 5 73.0 70 0 128 70 2.87 51 2.62 0 1 0 5 72.2 70 0 124 69 1.43 33 2.49 0 1 0 573.7 70 0 128 72 2.12 38 3.30 0 2 0 571.7 70 0 111 68 3.16 64 3.14 0 2 0 571.2 70 0 104 67 3.00 40 3.30 0 2 0 574.5 70 0 140 82 2.84 46 2.95 1 3 1 5 58.2 55 0 112 77 2.84 71 2.23 0 1 0 6 57.3 55 0 111 77 2.23 41 2.57 0 1 0 658.7 55 0 120 76 1.85 60 2.57 0 1 0 657.2 55 0 120 73 0.55 48 3.10 0 2 0 657.2 55 0 112 73 0.49 45 2.86 0 2 0 655.5 55 0 120 73 /p570.53 49 3.12 0 2 0 6 59.5 55 0 156 76 3.22 59 3.14 1 3 1 6 78.4 75 0 119 75 1.28 53 1.72 1 1 0 777.8 75 0 112 74 1.39 44 1.80 1 1 0 775.5 75 0 123 74 1.41 72 1.93 0 1 0 778.3 75 0 149 84 0.53 40 2.84 0 1 0 7406         Table 14.11 Continued AGE AGEG SEX SBP DBP LACR HDL LINSUL SMOKE DMS DM SN 76.9 75 0 153 75 3.78 43 2.95 0 2 0 7 75.4 75 0 144 77 4.57 45 2.54 0 2 0 777.9 75 0 156 86 5.38 39 3.34 0 3 1 768.0 65 0 123 70 1.62 48 2.49 0 1 0 8 66.2 65 0 131 72 1.71 56 2.47 0 1 0 8 65.8 65 0 136 80 3.82 56 2.84 1 1 0 8 68.8 65 0 120 66 2.20 45 2.45 0 1 0 868.0 65 0 162 60 2.35 62 3.86 0 2 0 867.9 65 0 115 54 2.33 39 3.92 0 2 0 867.8 65 0 132 79 2.68 42 2.47 0 3 1 863.1 60 1 123 80 1.81 46 1.93 1 1 0 9 61.3 60 1 122 78 1.41 85 1.88 0 1 0 9 60.8 60 1 131 83 2.43 31 1.90 1 1 0 961.8 60 1 109 69 1.27 61 2.35 1 2 0 963.1 60 1 114 73 1.04 38 2.53 0 2 0 960.7 60 1 130 76 0.84 46 2.59 0 2 0 962.2 60 1 133 85 2.10 37 3.04 0 3 1 9 78.7 75 0 147 85 0.48 58 2.62 0 1 0 10 77.4 75 0 167 84 0.92 71 2.66 0 1 0 1078.0 75 0 165 85 0.44 49 3.02 0 1 0 1075.2 75 0 117 70 2.05 42 2.88 0 2 0 1077.8 75 0 151 75 4.27 41 2.76 0 2 0 1078.5 75 0 137 74 2.13 40 2.86 0 2 0 10 78.5 75 0 156 81 5.33 52 327 0 3 1 10 56.5 55 0 108 71 1.58 41 2.27 0 1 0 1158.8 55 0 104 73 2.55 34 2.53 0 1 0 1155.7 55 0 135 77 2.06 106 2.32 1 1 0 1157.8 55 0 110 74 2.59 49 2.37 0 1 0 1157.3 55 0 153 91 0.54 44 3.13 0 2 0 11 55.4 55 0 141 94 1.47 61 3.15 0 2 0 11 56.9 55 0 123 78 3.72 40 3.18 0 3 1 1153.3 50 0 113 74 1.19 46 2.71 0 1 0 1250.8 50 0 143 89 3.45 78 1.84 0 1 0 1250.3 50 0 136 79 /p571.44 48 2.48 1 2 0 12 55.0 50 0 131 77 0.23 49 2.57 1 2 0 12 54.4 50 0 132 77 /p570.08 34 2.45 0 2 0 12 53.3 50 0 135 77 0.48 37 2.93 1 2 0 1250.7 50 0 114 78 2.52 41 4.37 0 3 1 12 /p63AGEG /p5850 if50 /p45age/p5855,/p5855if55 /p45age/p5860,/p5860if 60 /p45age/p5865,/p5865 if 65 /p45age/p5870,/p5870 if 70/p45age/p5875,/p5875 if 75 /p45age/p5880; SEX /p581 if male and /p580 if female; SMOKE /p581 if current smoker and 0 otherwise; SBP, systolic blood pressure; DBP, diastolic blood pressure; LACR,logarithm of the ratio of urinary albumin and creatinine; HDL, high-density lipoprotein in cholesterol; LINSUL, logarithm of insuline; DM /p581 if fasting glucose /p46126 mg/dL and /p580 otherwise; DMS, diabetic status defined by ADA fasting glucose criterion: DMS /p581 if normal fasting glucose, /p582 if impaired fasting glucose, and /p583 if diabetic; SN, stratum number.      407 Table 14.12 Results from the Conditional Logistic Regression Model for the DM Data in Example 14.9 95% Confidence Interval for Odds Ratio Regression Standard Chi-Square Odds Variable Coefficient Error Statistic pRatio Lower Upper AGE 0.327 0.161 4.127 0.0422 1.39 1.01 1.90 DBP 0.046 0.020 5.639 0.0176 1.05 1.01 1.09LACR 0.395 0.133 8.776 0.0031 1.48 1.14 1.93 LINSUL 0.860 0.288 8.900 0.0029 2.36 1.34 4.16 baseline and second examinations (about five years after the baseline examin- ation ). In this study, 33 persons with diabetes at the second examination are selected, and for each of these cases, six age group (in five-year interval )- and gender-matched diabetes-free controls are selected randomly from all partici-pantswithoutdiabetesinthesecondexamination.Thereare33strataandeachstratum contains one case (DM/p581)and its six matched controls (DM/p580), for a total of 231 participants. The demographic, physical, blood, and urinarydata collected at the baseline examination of the first 12 strata are listed inTable14.11 and arranged by stratum. The conditionallogistic modelbased on(14.2.20 )is used for these 1:6 matched data to identify risk factors for diabetes. The stepwise selection method is used to select the significant risk factors. Theresults are shown in Table 14.12. AGE, DBP, LACR, and LINSUL aresignificantriskfactorsfordiabetes.ThelargerthevaluesofAGE,DBP,LACR,and LINSUL, the higher is the risk of being diabetic. Asnotedearlier,for (14.2.21 ),thecomputersoftwareforaCoxproportional model with discrete time scale can be used to obtain an estimate of parameterb. Suppose that the text data file ‘‘C: /p33EX14d2d6.DAT’’ contains 12 successive columns,separatedbyaspace,withdataasinTable14.11:AGE,AGEG,SEX,SBP, DBP, LACR, HDL, LINSUL, SMOKE, DMS, DM, and SN. Thefollowing SAS code shows how the SAS procedure for the proportionalhazards model with discrete time scale can be used to obtain an estimate ofparameter bfor the conditional logistic regression model in a matched case—controlstudy (theresultsaregivenin Table14.12 ).In theSAScode,first, we define a nominal variable (for survival time ), TIME, and let TIME /p581i fa case (DM/p581), and /p582 if a control (DM/p580). This is accomplished by the statement ‘‘time /p582-dm;’’. Second, DM is also used to indicate censoring status, DM /p580 meaning censored, and /p581 uncensored. Thus,a case will have an uncensored time 1 and a control will have a censored time 2. data w1; infile ‘c: /p33ex14d2d6.dat’ missover;408         Table 14.13 Data from Example 14.9 If Stratified by Age Group Only Number of Number of Stratum AGEG DMs Non-DMs Total 15 0 —54 9 54 63 25 5 —59 8 48 56 36 0 —64 4 24 28 46 5 —69 4 24 28 57 0 —74 5 30 35 67 5 —79 3 18 21—— — ——Total 33 198 231 Table 14.14 Results from the Conditional Logistic Regression Model for the Data in Example 14.9 with Strata Defined by Age Groups 95% Confidence Interval for Odds Ratio Regression Standard Chi-Square Odds Variable Coefficient Error Statistic pRatio Lower Upper DBP 0.050 0.020 6.304 0.0120 1.05 1.01 1.09 LACR 0.391 0.124 9.876 0.0017 1.48 1.16 1.89 HDL /p570.037 0.019 4.034 0.0446 0.96 0.93 1.00 LINSUL 0.694 0.292 5.668 0.0173 2.00 1.13 3.55input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn; time/p582-dm; run;proc phreg data /p58w1 noprint; model time*dm (0)/p58age sbp dbp lacr hdl linsul smoke / ties /p58discrete selection /p58s; strata sn; run; Using the same data, if we stratify by age group only, Table 14.13 lists the number of cases and controls in each age group (stratum ). This can be consideredasanexampleofastratifieddesignwithadifferentnumbersofcasesand controls in each stratum: stratum 1 has a 9:54 match, stratum 2 an 8:48match, and so on. The results from the conditional logistic regression modelbased on this new stratification are given in Table 14.14.      409 The following SAS code can be used to generate the results in Table 14.14. The code can be modified to perform conditional logistic regression analysisfor data from any n/p16:n/p15matched design or a stratified design. data w1; infile ‘c: /p33ex14d2d6.dat’ missover; input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;time/p582-dm; run; proc phreg data /p58w1 noprint; model time*dm (0)/p58age sex sbp dbp lacr hdl linsul smoke / ties /p58discrete selection /p58s; strata ageg; run; 14.2.3 Other Models for Dichotomous Outcomes In the logistic regression model (14.2.3 ), the left side is a function of the probability of success, P/p71, and the right side is the linear combination of covariates.The function,calleda link function ,definesthe relationshipbetween the covariates and P/p71. In general, a link function represents the underlying biological, physical, or epidemiological relationship between the mean of thedependentvariable (probabilityofsuccess )andthecovariates x/p16,x/p17,...,x/p78.I n the logistic regressionmodel the link function,say g, is the logit functionof P/p71, that is, g(P/p71)/p58logit (P/p71)/p58logP/p711/p57P/p71 andg(P/p71)is assumed to be linearly related to the covariates, that is, logP/p711/p57P/p71/p58/p78/p26 /p72/p14/p15b/p72x/p71/p72 or P/p71/p58exp(/p26/p46/p72/p14/p15b/p72x/p71/p72) 1/p59exp(/p26/p46/p72/p14/p15b/p72x/p71/p72)(14.2.26) Twootherformsoflinkfunction g(P) that assumealinearrelationshipwith the covariates have been proposed and used in the literature. In the following,weintroducethesetwolinkfunctionsandthecorrespondingregressionmodel. 1.T he probit (or normit)function. This is the link function defined by the inverse of the cumulative standard normal distribution function, /afii9818/p92/p16(·): g(P/p71)/p58/afii9818/p92/p16(P/p71) (14 .2.27)410         The corresponding model is /afii9818/p92/p16(P/p71)/p58/p78/p26 /p72/p14/p15b/p72x/p71/p72 or P/p71/p58/afii9818/p1/p46/p26 /p72/p14/p16b/p72x/p71/p72/p2(14.2.28 ) 2.The complementary log-log link function . This function is defined by g(P/p71)/p58log[/p57log(1/p57P/p71)] (14.2.29 ) The corresponding model is log[/p57log(1/p57P/p71)]/p58/p78/p26 /p72/p14/p15b/p72x/p71/p72 or P/p71/p581/p57exp/p3/p57exp/p1/p46/p26 /p72/p14/p15b/p72x/p71/p72/p2/p4(14.2.30) The logistic regression model in (14.2.26 )is for binary outcomes such as diseased versus nondiseased. The model (14.2.28 )can be thought of as an alternative model for binary outcome. In addition, it can be used to modelthose binary outcomes that are defined by a cutoff point on the basis of anormallydistributed variable.For example,the cutoff point may be defined bythelastquintileorquartileofacontinuousmeasurementinanepidemiologicalstudy. When the binary outcomes are defined by a cutoff point in anasymmetricdistribution,themodelin (14.2.30 )maybeappropriate.Themodel (14.2.30 )can also be considered as a version of the Cox proportional hazards model for grouped survival times (Kalbfleisch and Prentice, 1973 ). Lety/p16,y/p17,...,y/p76be the observations with dichotomous values on the n subjects: y/p71/p581 for success and y/p71/p580 for failure. Similar to (14.2.4 ), the likelihood functions for the models in (14.2.28 )and (14.2.30 )can be obtained by replacing the corresponding P/p71in the following formula: L(b/p15,b/p16,...,b/p78)/p58/p76/p147 /p71/p14/p16P/p87 /p71/p71(1/p57P/p71)/p16/p92/p87/p71 The MLEs of the coefficients and the asymptotic likelihood inferences are similar to those given in Section 14.2.1 for the ordinary logistic regressionmodel except that interpretation of the odds ratio is not possible for the lattertwo models. The procedure LOGISTIC in SAS provides options for all threemodels.      411 Table 14.15 Asymptotic Partial Likelihood Inference from the Regression Models with Different Link Functions for the Data in Example 14.9 95% Confidence Interval for Odds Ratio Regression Standard Chi-Square Odds Variable Coefficient Error Statistic pRatio Lower Upper Model with L ogit L ink Function INTERCPT /p578.419 1.792 22.061 /p580.0001 DBP 0.044 0.018 5.673 0.0172 1.05 1.01 1.08 LACR 0.343 0117 8.627 0.0033 1.41 1.13 1.79 LINSUL 0.870 0.287 9.191 0.0024 2.39 1.38 4.27 Hosmer —Lemeshow test statistic 18.9460 0.0152 Model with Inverse Normal Link Function INTERCPT /p574.532 0.953 22.597 /p580.0001 DBP 0.023 0.010 5.302 0.0213 LACR 0.186 0.065 8.060 0.0045 LINSUL 0.445 0.156 8.146 0.0043 Hosmer —Lemeshow test statistic 7.386 0.4956 Model with Log-Log Link Function INTERCPT /p577.740 1.530 25.589 /p580.0001 DBP 0.038 0.016 5.919 0.0150 LACR 0.305 0.096 10.153 0.0014 LINSUL 0.785 0.241 10.592 0.0011 Hosmer —Lemeshow test statistic 17.415 0.0261Example 14.10 Consider the data in Example 14.9 as nonstratified data, Table 14.15 gives the results from the regression models defined in (14.2.26 ), (14.2.28 ), and (14.2.30 )by using the stepwise selection method. Based on Hosmer —Lemeshow test statistics, the regression model with the inverse normal link function gives a good fit to the data (p/p580.4956 ), whereas the othertwomodelsdonot (p/p580.0152and p/p580.0261 ).Allthreemodelsidentify DBP, LACR, and LINSUL as significant covariates for the development ofdiabetes. The following SAS, SPSS, and BMDP codes may be used to generate the results in Table 14.15. SAS code: data w1; infile ‘c: /p33ex14d2d6.dat’ missover; input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn; run;412         title ‘‘Regression model with the logit link function-generalized logistic regression’’; proc logistic data /p58w1 descending; model dm /p58age sex sbp dbp lacr hdl linsul smoke / selection /p58s lackfit link /p58logit; run;title ‘‘Regression model with the inverse normal link function‘;proc logistic data /p58w1 descending; model dm /p58age sex sbp dbp lacr hdl linsul smoke / selection /p58s lackfit link /p58probit; run; title ‘‘Regression model with the log-log link funtion‘;proc logistic data /p58w1 descending; model dm /p58age sex sbp dbp lacr hdl linsul smoke / selection /p58s lackfit link /p58cloglog; run; SPSS code for the model in (14.2.28 )with the forward selection method: data list file /p58‘c:/p33ex14d2d6.dat’ free / age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn. Logistic regression dm with age sex sbp dbp lacr hdl linsul smoke htn /method /p58fstep /print /p58all. BMDP code for procedure LR and the model in (14.2.28 ): /input file /p58‘c:/p33ex14d2d6.dat’ . variables /p5812. format /p58free. /variable names /p58age, ageg, sex, sbp, dbp, lacr, hdl, linsul, smoke, dms, dm, sn. Use/p58age, sex to smoke. /regress depend /p58dm. Interval /p58age, sex to smoke. Method /p58mlr. /print cell /p58used. /end 14.3 MODELS FOR POLYCHOTOMOUS OUTCOMES TheregressionmodelsinSection14.2canbeextendedtohandleoutcomesthat have more than two categories. Thesecategoriesmay be nominal, forexample,different types of heart disease or psychological conditions; or ordinal, forexample,differentlevelsofglucoseintoleranceordifferentseverityofcommuni-cationdisorders.Anoutcomevariablewithmorethantwopossibilitiesiscalledpolychotomous orpolytomous . In this section we discuss first the model for    413 nominalpolychotomousoutcomes (generalizedlogisticregressionmodel ),then the model for ordinal polychotomous outcomes (ordinal regression model ). Details regarding these models can be found in Aitchison and Silvey (1957 ), McCullagh (1980 ), Green (1984 ), McCullagh and Nelder (1989 ), Hosmer and Lemeshow (1989, 2000 ), Cox and Snell (1989 ), Afifi and Clark (1990 ), Agresti (1990 ), Collett (1991 ), and Ananth and Kleinbaum (1997 ). 14.3.1 Models for Nominal Polychotomous Outcomes: Generalized Logistic Regression Models LetY/p71denote the outcome for individual i. The outcome can be one of the m nominalcategories,suchasdifferentcelltypesoflungcancer.Let Y/p71/p58kdenote thatY/p71belongs to the kth category and k/p581,2,...,m. Suppose that for each ofnsubjects, pindependent variables x/p71/p58(x/p71/p16,x/p71/p17,...,x/p71/p78)/p30are measured. These variables can be either qualitative or quantitative. Let P(Y/p71/p58k/p34x/p71)be the probability that Y/p71/p58kgiven the pmeasured covariates x/p71; then /p26/p75/p73/p14/p16P(Y/p71/p58k/p34x/p71)/p581. Without loss of generality,using the last catalogas the reference, the generalized logistic regression model logP(Y/p71/p58k/p34x/p71) P(Y/p71/p58m/p34x/p71)/p58a/p73/p59/p78/p26 /p72/p14/p16b/p73/p72x/p71/p72k/p581, 2,...,m/p571 (14 .3.1) can be used to study the association of the covariates xto the outcome. To simplifythe notation,let u/p73/p71/p58a/p73/p59/p26/p78/p72/p14/p16b/p73/p72x/p71/p72. Similarto (14.2.1 )and (14.2.2 ), the model in (14.3.1 )assumes that the dependence on the covariates of the probability of being in the kth category is P(Y/p71/p58k/p34x/p71)/p58/p7exp(u/p73/p71) 1/p59/p26/p75/p92/p16/p72/p14/p16exp(u/p72/p71)k/p581, 2,...,m/p571 1 1/p59/p26/p75/p92/p16/p72/p14/p16exp(u/p72/p71)k/p58m(14.3.2) This model reduces to the logistic regression model in (14.2.1 )and (14.2.2 ) when m/p582. Letk/p16,...,k/p76be the outcomes observed for the nsubjects. Then the log-likelihood function based on the noutcomes observed is the logarithm of the product of all P(Y/p71/p58k/p71/p34x/p71)’s from the nsubjects, that is, l(a/p16,a/p17,...,a/p75/p92/p16,b/p16,b/p17,...,b/p75/p92/p16)/p58logL/p58log/p3/p76/p147 /p71/p14/p16P(Y/p71/p58k/p71/p34x/p71)/p4(14.3.3 ) where P(Y/p71/p58k/p71/p34x/p71)is given in (14.3.2 )andb/p73/p58(b/p73/p16,...,b/p73/p78)/p30,k/p581, 2,...,m/p571. There are a total of (m/p571)(p/p591)unknown coefficients. The estimation and hypothesis testing procedures for the coefficients are similar to414         those in the logistic regression model for dichotomous outcomes. Strictly speaking, the models in (14.3.1 )are not logistic regression models if m/p572. Therefore, the interpretation of the coefficients in these models needs to beclarified. Let us consider modeling the relationship between gender andcardiovascular disease status, NORMAL, STROKE, and CHD (coronary heart disease ). Let the outcome variable Ybe defined as Y/p581 if CHD, /p582i f STROKE, and /p583 if NORMAL, and the covariate SEX defined as SEX /p581 if male and /p580 if female. Then the two models according to (14.3.1 )are logP(Y/p71/p581/p34SEX/p71) P(Y/p71/p583/p34SEX/p71) /p58a/p16/p59b/p16·SEX/p71 logP(Y/p71/p582/p34SEX/p71) P(Y/p71/p583/p34SEX/p71)/p58a/p17/p59b/p17·SEX/p71 It is clear that neither of them is a logistic regression model. In the following, we show how to interpret the coefficients b/p16andb/p17in these models. From the first model, logP(Y/p581/p34SEX /p581)/P(Y/p583/p34SEX /p581) P(Y/p581/p34SEX /p580)/P(Y/p583/p34SEX /p580) /p58logP(Y/p581/p34SEX /p581) P(Y/p583/p34SEX /p581)/p57logP(Y/p581/p34SEX /p580) P(Y/p583/p34SEX /p580) /p58(a/p16/p59b/p16)/p57a/p16 /p58b/p16 and thus P(Y/p581/p34SEX /p581)/P(Y/p583/p34SEX /p581) P(Y/p581/p34SEX /p580)/P(Y/p583/p34SEX /p580)/p58exp(b/p16) (14 .3.4) Now let us cast the data into a 3 /p592 contingency table as in Table 14.16. The left side of (14.3.4 )can be estimated by (f/n/p16)/(b/n/p16) (e/n/p15)/(a/n/p15)/p58fa be However, if only the data from the normal and CHD participants are used, fa be/p58[f/(b/p59f)]/[b/(b/p59f)] [e/(e/p59a)]/[a/(e/p59a)]    415 Table 14.16 Nominal Cross-Classification of Cardiovascular (CVD) Status by Gender SEX CVD Status (Y) Female (0) Male (1) NORMAL (3) ab STROKE (2) cd CHD (1) ef——Total n/p15n/p16 which is an estimate of P(CHD /p34Male )/[1/p57P(CHD /p34male )] P(CHD /p34female )/[1/p57P(CHD /p34female )] or the ratio of the odds of a male having CHD to the odds of a female having CHD. Therefore, the exp( b/p19/p16)obtained from the first model can be interpreted as an estimate of the ratio of the odds of a male having CHD to the odds ofa female having CHD if only the data from the normal and CHD participantsareused. Similarly,exp (b/p19/p17)obtainedfromthe secondmodel canbe interpreted as an estimate of the ratio of the odds of a male having STROKE to the oddsof a female having STROKE if only the data from the normal and STROKEparticipants are used. The same interpretation also holds for coefficients ofcontinuous covariates in the models of (14.3.1 ); that is, an exponentiated coefficient for a continuous covariate is the odds ratio of a 1-unit increase inthe covariate assuming that other covariates are the same. Example 14.11 We use the data in Example 14.9 and assume that DM (Y/p581), IFG (Y/p582), and NFG (Y/p583)are three nominal categories. Let the referent category be NFG. For simplicity, only two covariates, systolic bloodpressure (SBP )and log insulin (LINSUL ), are included. Table 14.17 gives the results from fitting these covariates to the model (14.3.1 ). logP(ith participant is DM ) P(ith participant is NFG ) /p58logP(Y/p71/p581/p34x/p71) P(Y/p71/p583/p34x/p71) /p58/p577.648 /p590.026SBP/p71/p591.047LINSUL/p71 logP(ith participant is IFG ) P(ith participant is NFG )/p58logP(Y/p71/p582/p34x/p71) P(Y/p71/p583/p34x/p71) /p58/p574.949 /p590.011SBP/p71/p590.876LINSUL/p71416         Table 14.17 Asymptotic Partial Likelihood Inference from the Generalized Logistic Regression Model for Example 14.11 95% Confidence Interval Order of for Odds Ratio Coefficients Regression Standard Chi-Square Odds byS A S kVariable Coefficient Error Statistic pRatio Lower Upper DM vs. NFG b/p161 INTERCP /p577.648 1.648 21.530 /p580.0001 b/p181 SBP 0.026 0.010 6.300 0.0121 1.03 1.01 1.05 b/p201 LINSUL 1.047 0.304 11.870 00006 2.85 1.57 5.17 ---------------------------------------------------------------------------------------------------------- IFG vs. NFG b/p172 INTERCP /p574.949 1.427 12.020 00005 b/p192 SBP 0.011 0.010 1.410 0.2346 1.01 0.99 1.03 b/p212 LINSUL 0.876 0.262 11.150 0.0008 2.40 1.44 4.01 ---------------------------------------------------------------------------------------------------------- DM vs. IFG INTERCP /p572.699 1.850 2.130 0.1445 SBP 0.015 0.012 0.480 0.2239 1.02 0.99 1.04LINSUL 0.171 0.333 1.270 0.6066 1.19 0.62 2.28 ---------------------------------------------------------------------------------------------------------- H/p15:b/p18/p58b/p191.48 0 .2239 H/p15:b/p20/p58b/p210.27 0 .6066 417 Consequently, logP(ith participant is DM ) P(ith participant is IFG )/p58logP(Y/p71/p581/p34x/p71) P(Y/p71/p582/p34x/p71) /p58logP(Y/p71/p581/p34x/p71) P(Y/p71/p583/p34x/p71)/p57logP(Y/p71/p582/p34x/p71) P(Y/p71/p583/p34x/p71) /p58(/p577.648 /p594.949 )/p59(0.026 /p570.011 )SBP/p71 /p59(1.047 /p570.876 )LINSUL/p71 /p58/p572.699 /p590.015SBP/p71/p590.171LINSUL/p71 Thus, the odds ratio is 1.03 [exp (0.026 )] times (or 3% higher )for a 1-unit increase in SBP, and 2.85 [exp (1.047 )] times (or 185% higher )for a 1-unit increase in LINSUL from the model for DM vs. NFG. The odds ratio is 2.40[exp (0.876 )] times (or 140%higher )for a 1-unit increase in LINSULfrom the modelforIFGversusNFG.SBPis notsignificantinthemodelforIFGversusNFG (p/p580.2346 ). Neither SBP nor LINSUL is significant in the model for DMversus IFG (p/p580.2239and p/p580.6066, respectively ).One canalso follow the examples in Chapter 7, 9, 11, and 12 to perform additional statisticalinferences.Forinstance,wecantestwhetherthecoefficientsforSBPinthefirsttwo models are equal (whether the odds ratio for a 1-unit increase of SBP in the model for DM versus NFG is equal to that in the model for IFG versusNFG ), that is, H/p15:b/p18/p57b/p19/p580(where the subscripts 3 and 4 are the orders of the coefficients given by SAS ). From (11.2.13 ), under H/p15, Wald’s statistic, X/p53/p58(b/p19/p18/p57b/p19/p19)/p17/(v/p18/p18/p59v/p19/p19/p572v/p18/p19), has an asymptotic chi-square distribution with 1 degree of freedom, where v/p18/p18andv/p19/p19are the estimated variance of b/p18andb/p19, respectively, and v/p18/p19is the estimated covariance of b/p18andb/p19. From Table 14.17, the hypothesis is not rejected (p/p580.2239 ). Similarly, the hypoth- esisH/p15:b/p20/p57b/p21/p580 is not rejected (p/p580.6066 ); that is, there is insufficient evidence to say that the change in odds ratio for a 1-unit increase in LINSULin the model for DM versus NFG is not equal to that in the model for IFGversus NFG. The following SAS, SPSS, and BMDP codes can be used to obtain the results in Table 14.17. SAS code: data w1; infile ‘c: /p33ex14d2d6.dat’ missover; input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn;y/p584-dms; run; title ‘‘Generalized logistic regression model’’; proc catmod data /p58w1; direct sbp linsul;418         model y /p58sbp linsul / ml covb; contrast ‘Equal coefficients for SBP’ all —parms 0 0 1 /p57100 ; contrast ‘Equal coefficients for LINSUL’ all —parms00001 /p571; run; SPSS code: data list file /p58‘c:/p33ex14d2d6.dat’ free / age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn. Compute y /p584-dms. nomreg y with sbp linsul /print /p58fit history parameter lrt. BMDP PR code: /input file /p58‘c:/p33ex14d2d6.dat’ . variables /p5812. format /p58free. /variable names /p58age, ageg, sex, sbp, dbp, lacr, hdl, linsul, smoke, dms, dm, sn. Use/p58age, sex to smoke. /transform y /p584-dms. /group codes (y)/p581, 2, 3. Names (y)/p58DM, IFG, NFG. /regress depend /p58y. Level /p583. Type /p58nom. Interval /p58age, sex to smoke. enter /p58.05, .05. remove /p58/p58.05, .05. /print cell /p58model. /end 14.3.2 Model for Ordinal Polychotomous Outcomes: Ordinal Regression Models If the outcomes involve a rank ordering, that is, the outcome variable is ordinal,severalmultivaluedregressionmodelsareavailable.Readersinterestedin these models are referred to McCullagh and Nelder (1989 ), Agresti (1990 ), Ananth and Kleinbaum (1997 ), and Hosmer and Lemeshow (2000 ). In the followingdiscussion,weintroducethemostfrequentlyusedmodel,thepropor-tionaloddsmodel.Inthismodel,theprobabilityofanoutcomebeloworequalto a given ordinal level, P(Y/p45k), is compared to the probability that it is higher than the level given, P(Y/p57k). LetY/p71betheoutcomeofthe ithsubject.Assumethat Y/p71canbeclassifiedinto mordinal levels. Let Y/p71/p58kifY/p71is classified into the kth level and    419 k/p581,2,...,m. Suppose that for each of nsubjects, pindependent variables x/p71/p58(x/p71/p16,x/p71/p17,...,x/p71/p78)/p30are measured. These variables can be either qualitative orquantitative.Ifthe logit link function definedinSection14.2.3isused,similar to the logistic regression model (14.2.3 ), we consider the following models: logit (P(Y/p71/p45k/p34x/p71))/p58logP(Y/p71/p45k/p34x/p71) 1/p57P(Y/p71/p45k/p34x/p71)/p58a/p73/p59/p78/p26 /p72/p14/p16b/p72x/p71/p72 k/p581, 2,...,m/p571 (14.3.5 ) or, equivalently, let u/p73/p71/p58a/p73/p59/p26/p78/p72/p14/p16b/p72x/p71/p72, P(Y/p71/p45k/p34x/p71)/p58exp(a/p73/p59/p26/p78/p72/p14/p16b/p72x/p71/p72) 1/p59exp(a/p73/p59/p26/p78/p72/p14/p16b/p72x/p71/p72)/p58exp(u/p73/p71) 1/p59exp(u/p73/p71) k/p581, 2,...,m/p571 (14 .3.6) Therefore, P(Y/p71/p58k/p34x/p71)/p58P(Y/p71/p45k/p34x/p71)/p57P(Y/p71/p45k/p571/p34x/p71) /p58/p7exp(u/p16/p71) 1/p59exp(u/p16/p71)k/p581 exp(u/p73/p71) 1/p59exp(u/p73/p71)/p57exp(u/p73/p92/p16/p71) 1/p59exp(u/p73/p92/p16/p71)k/p582,...,m/p571 1/p57exp(u/p75/p92/p16/p71) 1/p59exp(u/p75/p92/p16/p71)k/p58m(14.3.7) Ifm/p582, that is, there are only two outcome levels, (14.3.7 )reduces to the logistic regression model in (14.2.3 ). The models in (14.3.5 )can be thought of as having only two outcomes [( Y/p45k) versus ( Y/p57k)] and therefore are logistic regression models. Thus, interpretation of the coefficients, b/p72, such as the exponentiated coefficient [exp (b/p72)] for a discrete or a continuous covariate is similar to that in a logistic regression model. Letk/p16,...,k/p76beobservedoutcomesfrom nsubjects.Thenthelog-likelihood function based on the noutcomes observed is the logarithm of the product of allP(Y/p71/p58k/p71/p34x/p71)’s from the nsubjects, that is, l(a/p16,a/p17,...,a/p75/p92/p16,b/p16,b/p17,...,b/p78)/p58logL/p58log/p3/p76/p147 /p71/p14/p16P(Y/p71/p58k/p71/p34x/p71)/p4(14.3.8 ) where P(Y/p71/p58k/p71/p34x/p71)is as givenin (14.3.7 ). Themaximumlikelihoodestimation and hypothesis-testing procedures for the coefficients are similar to thosediscussed previously. If the probit link function in(14.2.27 )is used, the models420         and formula corresponding to (14.3.5 )—(14.3.7 )are /afii9818/p92/p16(P(Y/p71/p45k/p34x/p71))/p58a/p73/p59/p78/p26 /p72/p14/p16b/p72x/p71/p72k/p581, 2,...,m/p571 P(Y/p71/p45k/p34x/p71)/p58/afii9818(u/p73/p71)k/p581, 2,...,m/p571 P(Y/p71/p58k/p34x/p71)/p58P(Y/p71/p45k/p34x/p71)/p57P(Y/p71/p45k/p571/p34x/p71) /p58/p7/afii9818(u/p16/p71) k/p581 /afii9818(u/p73/p71)/p57/afii9818(u/p73/p92/p16/p71)k/p582,...,m/p571 1/p57/afii9818(u/p75/p92/p16/p71) k/p58m If the complementary log-log link function in(14.2.29 )is used, the models and formula corresponding to (14.3.5 )—(14.3.7 )are log[/p57log(1/p57P(Y/p71/p45k/p34x/p71))]/p58a/p73/p59/p78/p26 /p72/p14/p16b/p72x/p71/p72k/p581, 2,...,m/p571 P(Y/p71/p45k/p34x/p71)/p581/p57exp[/p57exp(u/p73/p71)] k/p581, 2,...,m/p571 P(Y/p71/p58k/p34x/p71)/p58P(Y/p71/p45k/p34x/p71)/p57P(Y/p71/p45k/p571/p34x/p71) /p58/p71/p57exp[/p57exp(u/p16/p71)] k/p581 exp[/p57exp(u/p73/p92/p16/p71)]/p57exp[/p57exp(u/p73/p71)] k/p582,...,m/p571 exp[/p57exp(u/p75/p92/p16/p71)] k/p58m The log-likelihood function based on these two models can be obtained by replacing P(Y/p71/p58k/p71/p34x/p71)in(14.3.8 )with the respective expressions above. Example 14.12 Now consider the NFG, IFG, and DM categories in Example14.9 that represent three levels of severity in glucoseintolerance. DM(diabetes )is defined as fasting plasma glucose (FPG )/p46126 mg/dL, IFG (impaired fasting glucose )as FPG between 110 and 125 mg/dL, and NFG (normal fasting glucose )as FPG /p58110 mg/dL. Thus, it is reasonable to consider the outcome variable as ordinal. Let the outcome variable Y/p581i f DM, 2 if IFG, and 3 if NFG. We fit the models in (14.3.5 )using the SAS procedure LOGISTIC with all the covariates. The SAS program allows usersto use a variable selection method (forward, backward, and stepwise ). In this case,weuse thestepwiseselectionmethod, andthe resultsare given in the firstpart of Table 14.18. The stepwise method identifies SBP and LINSUL assignificant independent variables. For k/p581 [i.e., we compare diabetes with    421 Table 14.18 Asymptotic Partial Likelihood Inference from the Ordinal Regression Model with Different Link Functions for the Diabetic Status Data in Example 14.9 95%Confidence Interval for Odds Ratio Regression Standard Chi-Square Odds k Variable Coefficient Error Statistic p Ratio Lower Upper Model with Logit Link Function 1 INTERCP1 /p576.753 1.183 32.571 0.0001 2 INTERCP2 /p575.485 1.151 22.708 0.0001 SBP 0.019 0.007 6.114 0.0134 1.02 1.00 1.03LINSUL 0.925 0.213 18.803 0.0001 2.52 1.67 3.90 Log-likelihood ratio statistic for H/p15:b/p16/p58b/p17/p580/p6326.831 0 .0001 Model with Inverse Normal Link Function 1 INTERCP1 /p573.971 0.677 34.415 0.0001 2 INTERCP2 /p573.240 0.664 23.790 0.0001 SBP 0.011 0.004 6.311 0.0120 LINSUL 0.530 0.123 18.674 0.0001 Log-likelihood ratio statistic for H/p15:b/p16/p58b/p17/p5802 6 .261 0 .0001 Model from Complementary Log-Log Link Function 1 INTERCP1 /p575.626 0.915 37.813 0.0001 2 INTERCP2 /p574.562 0.894 26.025 0.0001 SBP 0.014 0.006 5.721 0.0168 LINSUL 0.715 0.162 19.534 0.0001 Log-likelihood ratio statistic for H/p15:b/p16/p58b/p17/p5802 5 .835 0 .0001 /p63b/p16andb/p17are coefficients for SBP and LINSUL, respectively. 422 nondiabetes (NFG /p59IFG)] the estimated model in (14.3.5 )is logP(Y/p71/p451/p34x/p71) 1/p57P(Y/p71/p451/p34x/p71)/p58logP(participant iis diabetic ) P(participant iis nondiabetic ) /p58/p576.753 /p590.019SBP/p71/p590.925LINSUL/p71 Fork/p582, the estimated model in (14.3.5 )is logP(Y/p71/p452/p34x/p71) 1/p57P(Y/p71/p452/p34x/p71)/p58logP(participant iis either DM or IFG ) P(participant iis NFG ) /p58/p575.485 /p590.019SBP/p71/p590.925LINSUL/p71 Accordingto (14.3.7 ), we can estimatetheprobabilityof developingDM, IFG, or remaining NFG. For example, the probability of developing IFG is P(Y/p71/p582/p34x/p71)/p58P(participant iis IFG ) /p58exp(/p575.485 /p590.019SBP/p71/p590.925LINSUL/p71) 1/p59exp(/p575.485 /p590.019SBP/p71/p590.925LINSUL/p71) /p57exp(/p576.753 /p590.019SBP/p71/p590.925LINSUL/p71) 1/p59exp(/p576.753 /p590.019SBP/p71/p590.925LINSUL/p71) Thus, for a person whose systolic blood pressure is 140 mmHg and whose log insulin is 3, the probability of developing IFG can be obtained by pluggingthese values into the preceding equation. The result is P(participant is IFG )/p580.951 1/p590.951/p570.268 1/p590.268 /p580.276 As noted earlier, the coefficients in these models can be interpreted as those in the ordinary logistic regressionmodel for binary outcomes. In this example,the higher SBP and LINSUL are, the higher the odds of having DM than ofnot having DM, or the higher the odds of having either DM or IFG than ofbeing NFG. The odds ratio is 1.02 [exp (0.019 )] times (or 2% higher )for a 1-unit increase in SBP assuming that LINSUL is the same, and 2.52 times (or 152%higher )for a 1-unit increase in LINSULassuming that SBP is the same. From the table, SBP and LINSUL are related significantly to the diabeticstatus in all models with different link functions. SAS and SPSS can also be used for the other two link functions:the inverse ofthecumulativestandardnormaldistributionandthecomplementarylog-log    423 link functions introduced in Section 14.2.3. Table 14.18 includes the results frommodelswiththesetwolinkfunctions.Theresultsareverysimilartothoseobtained using the logit link function. The following SAS, SPSS, and BMDP codes can be used to obtain the results in Table 14.18. SAS code: data w1; infile ‘c: /p33ex14d2d6.dat’ missover; input age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn; run;title ‘‘Ordinal regression model with logic link function’’;proc logistic data /p58w1 descending; model dms /p58age sex sbp dbp lacr hdl linsul smoke / selection /p58s lackfit link /p58logit; run;title ‘‘Ordinal regression model with inverse normal link function‘;proc logistic data /p58w1 descending; model dms /p58age sex sbp dbp lacr hdl linsul smoke / selection /p58s lackfit link /p58probit; run;title ‘‘Ordinal regression model with complementary log-log link function’’;proc logistic data /p58w1 descending; model dms /p58age sex sbp dbp lacr hdl linsul smoke / selection /p58s lackfit link /p58cloglog; run; SPSS code: data list file /p58‘c:/p33ex14d2d6.dat’ free / age ageg sex sbp dbp lacr hdl linsul smoke dms dm sn. Compute y /p584-dms. plum y with sbp linsul /link /p58logit /print /p58fit history parameter. plum y with sbp linsul /link /p58probit /print /p58fit history parameter. plum y with sbp linsul /link /p58cloglog /print /p58fit history parameter. BMDP PR code for the logit link function only: /input file /p58‘c:/p33ex14d2d6.dat’ . variables /p5812. format /p58free.424         /variable names /p58age, ageg, sex, sbp, dbp, lacr, hdl, linsul, smoke, dms, dm, sn.Use/p58age, sex to smoke. /transform y /p584-dms. /group codes (y)/p581, 2, 3. Names (y)/p58DM, IFG, NFG. /regress depend /p58y. Level /p583. Type /p58ord. Interval /p58age, sex to smoke. enter /p58.05, .05. remove /p58/p58.05, .05. /print cell /p58used. /end Note that the model for ordinal polychotomous outcomes in BMDP PR is defined as logP(Y/p71/p57k/p34x/p71) 1/p57P(Y/p71/p57k/p34x/p71)/p58/afii9825/p23/p73/p59/p78/p26 /p72/p14/p16/afii9826/p18/p72x/p71/p72/p58u/p23/p73/p71k/p581, 2,...,m/p571 Compared with (14.3.5 ),/afii9825/p23/p73/p58/p57a/p73,k/p581, 2,...,m/p571;/afii9826/p18/p72/p58/p57b/p72,j/p581, 2,...,p. Bibliographical Remarks The linear logistic regression method is discussed extensively in Cox (1970 ), Cox and Snell (1989 ), Collett (1991 ), Kleinbaum (1994 ), and Hosmer and Lemeshow (2000 ). Cox’s book provides the theoretical background, and Hosmer and Lemeshow discuss broad application of the method, includingmodel-building strategies and interpretation and presentation of analysisresults. In addition to the papers and books cited in this chapter, other workson the subject include Anderson (1972 ), Mantel (1973 ), Prentice (1976 ), Prentice and Pyke (1979 ), Holford et al. (1978 ), and Breslow and Day (1980 ). Applications of the logistic regression model can easily be found in variousbiomedical journals. EXERCISES 14.1Consider the study presented in Example 3.5 and the data for the 40 patients in Table 3.10.(a)Construct a summary table similar to Table 3.11. (b)Construct a table similar to Table 3.12. (c)Usethechi-squaretesttodetectanydifferencesinretinopathyrates among the subgroups obtained in part (b). 425 (d)On the basis of these 40 patients, identify the most important risk factors using a linear logistic regression method. 14.2Consider the data for the 33 hypernephroma patients given in Exercise Table 3.1. Let ‘‘response’’ be defined as stable, partial response, orcomplete response.(a)Compare each of the five skin test results of the responders with those of the nonresponders. (b)Use alinearlogisticregressionmethodtoidentifythemostimport- ant risk factors related to response. (i)Consider the five skin tests only. (ii)Consider age, gender, and the five skin tests. 14.3Consider all nine risk variables (age, gender, family history of melanoma, and six skin tests )in Exercise 3.3 and Exercise Table 3.3. Identify the most important prognostic factors that are related toremission. Use both univariate and multivariate methods. 14.4Consider the data of 58 hypernephroma patients given in Exercise Table 3.2. Apply the logistic regression method to response (defined as complete response, partial response, or stable disease ). Include gender, age, nephrectomy treatment, lung metastasis, and bone metastasis asindependent variables. (a)Identify the most significant independent variables. (b)Obtain estimates of odds ratios and confidence intervals when applicable. 14.5Consider the case where there is one continuous independent variable X/p16. Show that the log odds ratio for X/p16/p58x/p16/p59mversus X/p16/p58x/p16is mb/p16, where b/p16is the logistic regression coefficient. 14.6Using the data in Table 12.4, define the index function CVD as CVD /p581i fd g /p461, and CVD /p580 otherwise, and fit a logistic re- gression model for CVD by using the stepwise selection method toselectrisk factorsamongthe samefactorsas thosenotedatthe bottomof Table 12.7. Compare the results obtained with those in Table 12.7. 14.7Assumingthat P(apersonissampled /p34y,x)/p58P(apersonissampled /p34y), that is, the sampling probability is independent of the risk factors x, derive (14.2.15 ). 14.8By using (14.2.14 )and (14.2.1 ), show that (14.2.20 )reducesto (14.2.21 ). 14.9Derive (14.3.2 ).426         14.10Consider the data in Table 12.4. Fit the generalized logistic regression modelin (14.3.1 )for DG with covariates AGE, SEX, LACR, and LTG by using the SAS CATMOD, SPSS NOMREG, or BMDP PRprocedure. Select risk factors among those noted at the bottom ofTable 12.7 using the stepwise selection method in the BMDP PRprocedure. Compare the results with those given in Table 13.5. 14.11UsingthesamenotationanddataasinTable14.11, (1)fit theoutcome variable Ywiththegeneralizedlogisticregressionmodelin (14.3.1 )with SEXasthecovariate; (2)fitalogisticregressionforthebinaryoutcome DM versus NFG, with SEX as the covariate, by using the data fromDM and NFG participants only; (3)fit a logistic regression for the binary outcome IFG versus NFG, with SEX as the covariate, by usingthe data from IFG and NFG participants only; (4)compare the coefficients obtained from (2)and (3)with the coefficients obtained from (1), and (5)report what you have found. 14.12Perform the same analyses as in Exercise 14.11 but use SBP as the covariate, and discuss your findings. 427 APPENDIX A Newton- -Raphson Method The Newton —Raphson method (Ralston and Wilf, 1967;Carnahan et al., 1969 ) is a numerical iterative procedure that can be used to solve nonlinearequations. An iterative procedure is a technique of successive approximations,and each approximation is called an iteration. If the successive approximations approach the solution very closely, we say that the iterations converge. The maximum likelihood estimates of various parameters and coefficients discussedin Chapters 7, 9, and 11 to 14 can be obtained by using the Newton —Raphson method. In this appendix we discuss and illustrate the use of this method, firstconsidering a single nonlinear equation and then a set of nonlinear equations. Let f(x)/p580 be the equation to be solved for x. The Newton —Raphson method requires an initial estimate of x, say x/p24/p15, such that f(x/p24/p15) is close to zero preferably, and then the first approximate iteration is given by x/p24/p16/p58x/p24/p15/p57f(x/p24/p15) f/p30(x/p24/p15) (A.1) where f/p30(x/p24/p15) is the first derivative of f(x) evaluated at x/p58x/p24/p15. In general, the (k/p591)th iteration or approximation is given by x/p24/p73/p62/p16/p58x/p24/p73/p57f(x/p73) f/p30(x/p73)(A.2) where f/p30(x/p24/p73)is the first derivative of f(x) evaluated at x/p58x/p24/p73. The iteration terminates at the kth iteration if f(x/p24/p73)is close enough to zero or the difference between x/p24/p73and x/p24/p73/p92/p16is negligible. The stopping rule is rather subjective. Acceptable rules are that f(x/p24/p73)o r d/p58x/p24/p73/p57x/p24/p73/p92/p16is in the neighborhood of 10/p92/p21or 10 /p92/p22. Example A.1 Consider the function f(x)/p58x/p18/p57 x/p592 428 Figure A.1 Graphical presentation of the Newton —Raphson method for Example A.1. We wish to find the value of xsuch that f(x)/p580 by the Newton —Raphson method. The first derivative of f(x)i s f/p30(x)/p583x/p17/p57 1 Since f(/p571)/p582 and f(/p572)/p58/p57 4, graphically (Figure A.1 ), we see that the curve cuts through the xaxis [ f(x)/p580] between /p571 and /p572. This gives us a good hint of an initial value of x. Suppose that we begin with x/p24/p15/p58/p57 1; f(x/p24/p15)/p582 and f/p30(x/p24/p15)/p582. Thus, the first iteration, following (A.1), gives x/p24/p16/p58/p57 1/p572 2/p58/p57 2 and f(x/p24/p16)/p58/p57 4 and f/p30(x/p24/p16)/p5811. Following (A.2), we obtain the following: Second iteration: x/p24/p17/p58/p57 2/p594 11/p58/p57 1.6364 f(x/p24/p17)/p58/p57 0.7456 f/p30(x/p24/p17)/p587.0334 -  429 Third iteration: x/p24/p18/p58/p57 1.6364/p590.7456 7.0334/p58/p57 1.5304 f(x/p24/p18)/p58/p57 0.054 f/p30(x/p24/p18)/p586.0264 Fourth iteration: x/p24/p19/p58/p57 1.5304/p590.054 6.0264/p58/p57 1.52144 f(x/p24/p19)/p58/p57 0.00036 f/p30(x/p24/p19)/p585.9443 Fifth iteration: x/p24/p20/p58/p57 1.52144 /p590.00036 5.9443/p58/p57 1.52138 f(x/p24/p20)/p580.0000017 At the fifth iteration, for x/p58/p57 1.52138, f(x) is very close to zero. If the stopping rule is that f(x)/p4510/p92/p21, the iterative procedure would terminate after the fifth iteration and x/p58/p57 1.52138 is the root of the equation x/p18/p57 x/p592/p580. Figure A.1 gives the graphical presentation of f(x) and the iteration. It should be noted that the Newton —Raphson method can only find the real roots of an equation. The equation x/p18/p57 x/p592/p580 has only one real root, as shown in Figure A.1;the other two are complex roots. The Newton —Raphson method can be extended to solve a system of equations with more than one unknown. Suppose that we wish to find valuesofx/p16,x/p17,...,x/p78such that f/p16(x/p16,...,x/p78)/p580 f/p17(x/p16,...,x/p78)/p580 /p36 f/p78(x/p16,...,x/p78)/p580 Let a/p71/p72be the partial derivative of f/p71with respect to x/p72;that is, a/p71/p72/p58/p42f/p71//p42x/p72.430  -  The matrix J/p58a/p16/p16/p37 a/p16/p78 a/p17/p16/p37 a/p17/p78 /p36/p36 a/p78/p16/p37 a/p78/p78 is called the Jacobian matrix . Let the inverse of J, denoted by J/p92/p16,b e J/p92/p16 /p58b/p16/p16/p37 b/p16/p78 b/p17/p16/p37 b/p17/p78 /p36/p36 b/p78/p16/p37 b/p78/p78 Let x/p73/p16,x/p73/p17,...,x/p73/p78be the approximate root at the kth iteration;let f/p73/p16,...,f/p73/p78be the corresponding values of the functions f/p16,...,f/p78, that is, f/p73/p16/p58f/p16(x/p73/p16,...,x/p73/p78) /p36 f/p73/p78/p58f/p78(x/p73/p16,...,x/p73/p78) and let b/p73/p71/p72be the ijth element of J/p92/p16evaluated at x/p73/p16,...,x/p73/p78. Then the next approximation is given by x/p73/p62/p16/p16/p58x/p73/p16/p57(b/p73/p16/p16f/p73/p16/p59b/p73/p16/p17f/p73/p17/p59/p37/p59b/p73/p16/p78f/p73/p78) x/p73/p62/p16/p17/p58x/p73/p17/p57(b/p73/p17/p16f/p73/p16/p59b/p73/p17/p17f/p73/p17/p59/p37/p59b/p73/p17/p78f/p73/p78)( A.3) /p36 x/p73/p62/p16/p78/p58x/p73/p78/p57(b/p73/p78/p16f/p73/p16/p59b/p73/p78/p17f/p73/p17/p59/p37/p59b/p73/p78/p78f/p73/p78) The iterative procedure begins with a preselected initial approximate x/p15/p16, x/p15/p17,...,x/p15/p78, proceeds following (A.3), and terminates either when f/p16,f/p17,...,f/p78are close enough to zero or when differences in the xvalues at two consecutive iterations are negligible. Example A.2 Suppose that we wish to find the value of x/p16and x/p17such that x/p17/p16/p59x/p16x/p17/p572x/p16/p571/p580 x/p18/p16/p57x/p16/p59x/p17/p572/p580 -  431 In this case, p/p582: f/p16/p58x/p17/p16/p59x/p16x/p17/p572x/p16/p571 f/p17/p58x/p18/p16/p57x/p16/p59x/p17/p572 Since /p42f/p16//p42x/p16/p582x/p16/p59x/p17/p572,/p42f/p16//p42x/p17/p58x/p16,/p42f/p17//p42x/p16/p583x/p17/p16/p571, and /p42f/p17//p42x/p17/p581, the Jacobian matrix is J/p58/p32x/p16/p59x/p17/p572 3x/p17/p16/p571x/p161/p4(A.4) Let the initial estimates be x/p15/p16/p580,x/p15/p17/p581,f/p15/p16/p58/p57 1, and f/p15/p17/p58/p57 1: J/p58/p3/p5710 /p5711/p4J/p92/p16 /p58 /p3/p5710 /p5711/p4 Iteration 1. Following (A.3), we obtain x/p16/p16/p580/p57[(/p571)(/p571)/p590(/p571)]/p58/p57 1 x/p16/p17/p581/p57[(/p571)(/p571)/p591(/p571)]/p581 With these values, f/p16/p16/p581,f/p16/p17/p58/p57 1, and J/p58/p3/p573/p571 21 /p4J/p92/p16 /p58 /p3/p571/p571 23 /p4 Iteration 2. From (A.3)we obtain x/p17/p16/p58/p57 1/p57[(/p571)(1)/p59(/p571)(/p571)]/p58/p57 1 x/p17/p17/p581/p57[(2)(1) /p59(3)(/p571)]/p582 With these values, f/p17/p16/p580 and f/p17/p17/p580. Therefore, the iteration procedure terminates and the solution of the two simultaneous equations is x/p16/p58/p57 1, x/p17/p582. The number of iterations required depends strongly on the initial values chosen. In Example A.2, if we use x/p15/p16/p580,x/p15/p17/p580, it requires about 11 iterations to find the solution. Interested readers may try it as an exercise.432  -  APPENDIX B Statistical Tables 433 Table B-1 Normal Curve Areas Source:Abridgedfrom Table 1 of Statistical Tables and Formulas ,by A. Hald, JohnWiley &Son s, 1952. Reproduced by permissionof JohnWiley & Son s. 434 Table B-2 Percentage Points of the /afii98512-Distribution Source:‘‘Tables of the Percentage Points of the /afii9851/p17-Distribution,’’ by Catherine M. Thompson, Biometrika , Vol. 32, pp. 188 —189(1941 ). Reproduced by permissionof the editor of Biometrika. 435 Table B-3 5 %Points of the F-Distribution 436 oa RT|giabee eaa aT Lee 437 Table B-3 2.5 %Points of the F-Distribution 438 a)835 E2285 S385 SES82 SEESE ERERE EREE 3)-283 S822 22552 25558 RESEE S8ELE SEg83 §8868 C2555 SERAE RSAES Skee SERRE E5523 ¢|<S8e SEEES ZESRE £2298 BORES BECES Bezeg| |S RSE StF SERS GERI ENING ESkae GEES Z|5]£38SE52§ S208 S658 FREES SgnzR ROLES5[8]gageSec8SSekE22R28CERESESEEEGezee £|.)829SESEE G2E23 £LG82 SESE Loans foggy E[k| Bpre Coc8S Bass GEnkS SREL8 ALES GEER i puueib ae |gueFUG8ESEGREELERELEGGEZEEEBRESTBee SEE05 SSEK3 SS555 S8083 o|s8o2 258% EESES SELIG REESE SESE EuEas gers SS598 BEE55 FLERE RESES SERSR EASES VATerwmenusS2unySEESRANEERARERASSESTamang Fo moped foPaaog 439 Table B-3 1 %Points of the F-Distribution 440 |sez se8s2 BEES ©S288 S28E8 FE 2 i] Sonse SEQMsees ay 22seoes . || BRE SRESS FREES GENRE Skee ded ARES oldEe ageSP2352gHESGEFRESESLVRPSSz! |HaeESSEUEEEEE282EDEEESERREES 3 evs RESSZ EEGZE EPSES SRESE SEREY SESRE : |~82% 22522 SEES SELES ZEERE TEES! iB [7S SReeee 441 Table B-3 0.5 %Points of the F-Distribution 442 Source:‘‘Tables of Percentage Points of the Inverted Beta (F)Distribution,’’ by Maxine Merrington and Catheri ne M. Thompson, Biometrika,V o l .3 3 ,p p .7 3 —88(1943 ). Reproduced by permissionof the editor of Biometrika. 443 Table B-4 Upper Tail Probabilities for the Null Distribution of the Kruskal--Wallis H Statistic: k/p583,n1/p581(1)5, n2/p58n1(1)5, 2/p45n3/p58n2(1)5 444 Table B-4 ( continued ) 445 Table B-4 ( continued ) 446 Table B-4 ( continued ) 447 Table B-4 ( continued ) 448 Table B-4 ( continued ) 449 Table B-4 ( continued ) 450 Table B-4 ( continued ) 451 Table B-4 ( continued ) 452 Table B-4 ( continued ) 453 Table B-4 ( continued ) 454 Table B-4 ( continued ) 455 Table B-4 ( continued ) 456 Table B-4 ( continued ) 457 Table B-4 ( continued ) Source:Table F of A Nonparametric Introduction to Statistics , by C. H. Kraft and C van Eedan, Macmillan, New York, 1968. Reproduced by permission of the Macmillan Publishing Company. 458 Table B-5 Selected Critical Values for All Treatments: Multiple Comparisons Based on Kruskal--Wallis Rank Sums Source:‘‘Rank Sum Multiple Comparisons in One- and Two-Way Classification,’’ by B. J. McDonald and W. A. Thompson, Biometrika , Vol. 54, pp. 487 —497 (1967 ). Reproduced by permissionof the editor of Biometrika . The starred values are from ‘‘Distribution-Free Multiple Comparisons,’’ Ph.D. thesis (1963 ), P. Nemenyi, Princeton University, with permission of the author. 459 Table B-6 Selected Critical Values for the Range of kIndependent N(0, 1) Variables: k/p582(1)20(2)40(10)100 For a given kand/afii9825, the tabled entry is q(/afii9825,k,/p45). Source:‘‘TableofRange and StudentizedRange,’’by H.L. Harter, Ann. Math. Statist. , Vol.31,pp. 1122—1147 (1960 ). Reproduced by permissionof the editor of the Annals of Mathematical Statistics. 460 Table B-7Percentage Points of the t-Distribution Source:‘‘Table of Percentage Points of the t-Distribution,’’ by Maxine Merrington, Biometrika , Vol. 32, p. 300 (1941 ). Reproduced by permissionof the editor of Biometrika . 461 Table B-8 Coefficients ( aiandbi) of the Best Estimates of the Mean ( /afii9839) and Standard Deviation ( /afii9846) in Censored Samples Up to n/p5820 from A Normal Population 462 Table B-8 ( continued ) 463 Table B-8 ( continued ) 464 Table B-8 ( continued ) 465 Table B-8 ( continued ) 466 2|8 eae iuHaE =[88 53 58 RS 3jiga88588 iB HEEE |:HERB EHHs\RR EH ;URERNREHH s(QHHRREEHES PEER GEEEESs ;HHBRHHEER AE SHRHNEERHE GEE SHRERHEHRHBUEE sHHEGHERHREERH ES PEEeee SeeREae /HERERHRERHEEHAS FESReeeeenaeas i]tleaes 467 Table B-8 ( continued ) 468 469 Table B-8 ( continued ) 470 =le285? 5(88232 aleg atgeceefce22as PRS EGEEER 471 Table B-8 ( continued ) Source:‘‘Estimation of Location and Scale Parameters by Order St atistics from Singly and Doubly Censored Samples, Parts I a nd II,’’ by A. E. S a r h a na n dB .G .G r e e n b e r g ,Ann. Math. Statist. , Vol. 27, pp. 427 —451 (1956 ). Reproduced by permissionof the editor of the Annals of Mathematical Statistics. 472 Table B-9 Variances and Covariances of the Best Linear Est imates of the Mean ( /afii9839/p24) and Standard Deviation ( /afii9846/p24)f o r Censored Samples Up to Size 20 from a Normal Population 473 Table B-9 ( continued ) Source:Up to n/p5815 of this table is reproduced from A. E. Sarhanan d B. G. G reen berg, ‘‘Estimationof Locationan d Scale Parameters by O rder Statistics from Singly and Censored Samples, Parts I and II,’’ Ann. Math. Statist ., Vol. 27, pp. 427 —451(1956),a ndV o l .2 9 ,p p .7 9 —105(1958), with permissionof the editor of the Annals of Mathematical Statistics . The rest of the table is produced from A. E. Sarhan and B.G. Greenberg, ‘‘Estimation of Location and Scale Parameters by Order Statistics from Singly and D oubly Censored Samples, Part III,’’ Tech. Rep. 4-OOR, Proje ct 1597, U.S. Army Research Office. 474 Table B-10 1 /(1/p57R) and /afii9828/p24for the Estimation of the Parameters of the Gamma Distribution When There Are No Censored Observations Source:‘‘Estimationof Parameters of the Gamma DistributionUsin g Order Statistics,’’ by M. B. Wilk, R. Gnanadesikan, and Marilyn J. Huyett, Biometrika , Vol. 49, pp. 525 —545 (1962 ). Reproduced by permissionof the editor of Biometrika . 475 Table B-11 /afii9828/p24(P,S) and /afii9839/p24(P,S) for Various Values of n/r:n/r/p581.0 ForP/p450.52 read Sfrom the left-hand margin, and for P/p460.56 read Sfrom the right-hand margin. Note that the figures in region 2 are printed in bold roman type and those in region 3 inbold italic type; the remainder of the table (outside of regions 2 and 3 )is region1. 476 Table B-11 ( continued ) 477 Table B-11 ( continued ) 478 Table B-11 ( continued ) 479 Table B-11 ( continued ) 480 Table B-11 ( continued ) 481 Table B-11 ( continued ) 482 Table B-11 ( continued ) 483 Table B-11 ( continued ) 484 Table B-11 ( continued ) Source:‘‘Estimationof Parameters of the Gamma DistributionUsin g Order Statistics,’’ by M. B. Wilk, R. Gnanadesikan, and Marilyn J. Huyett, Biometrika , Vol. 49, pp. 525 —545 (1962 ). Reproduced by permissionof the editor of Biometrika. 485 Table B-12 Percentage Points l/afii9825Such That P(/afii9828/p241//afii9828/p242/p58l/afii9825)/p581/p57/afii9825 Source:‘‘Two Sample Test in the Weibull Distribution,’’ by D. R. Thoman and L. J. Bain, Technometrics , Vol. 11, pp. 805 —815 (1969 ). Reproduced by permissionof the editor of Techno- metrics. 486 Table B-13 Percentage Points z/afii9825Such That P(G/p58z/afii9825)/p581/p57/afii9825 Source:‘‘Two Sample Test in the Weibull Distribution,’’ by D. R. Thoman and L. J. Bain, Technometrics , Vol. 11, pp. 805 —815 (1969 ). Reproduced by permissionof the editor of Techno- metrics. 487 References Aaronson, K. D., Schwartz, J. S., Chen, T. M., Wong, K. L., Goin, J. E, and Mancini,D. M.(1997 ). Development and Prospective Validation of a Clinical Index to Predict Survival in Ambulatory Patients Referred for Cardiac Transplant Evaluation.Circulation ,95, 2660—2667. Abramowitz, M., and Stegun, I. A. (1964 ).Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables. Applied Mathematics Series 55. National Bureau of Standards, Washington, DC. Afifi, A. A., and Clark, V. (1990 ).Computer-Aided Multivariate Analysis , 2nd ed. Lifetime Learning Publications, Belmont, CA. Agresti, A. (1990 ).Categorical Data Analysis. Wiley, New York. Aitchison, J. (1970 ). Statistical Problems of Treatment Allocation. Journal of the Royal Statistical Society, Series A ,133, 206—238 Aitchison, J., and Brown, J. A. C. (1957 ).The Lognormal Distribution. Cambridge University Press, Cambridge. Aitchison,J., and Silvey, S. D. (1957 ). The Generalizationof ProbitAnalysisto the Case of Multiple Responses. Biometrika ,44, 131—140. Aitkin, M., Laird, N., and Francis, B. (1983 ). A Reanalysis of the Stanford Heart Transplant Data (with discussion ).Journal of the American Statistical Association , 78, 264—292. Akaike, H. (1969 ). Fitting Autoregressive Models for Prediction. Annals of the Institute of Statistical Mathematics ,21, 243—247. Akaike, H. (1974 ). A New Look at the Statistical Model Identification. IEEE Transac- tions on Automatic Control, AC-19, 716—723. Albert, I. J. (2000 ). The Use of Frailty Models in Genetic Studies: Application to the Relationship between End-Stage Renal Failure and Mutation Type in AlportSyndrome. European Community Alport Syndrome Concerted Action Group(ECASCA ).Journal of Epidemiology and Biostatistics ,5(3), 169—175. Albertson, P. C., Hanley, J. A., Gleason, D. F., and Barry, M. J. (1998 ). Competing Risk Analysis of Men Aged 55 to 74 Years at Diagnosis Managed Conservatively for ClinicallyLocalizedProstate Cancer. Journal of the American Medical Association , 280(11),97 5—980. 488 Alioum, A., and Commenges, D. (1996 ). A Proportional Hazards Model for Arbitrarily Censored and Truncated Data. Biometrics ,52, 512—524. Altshuler, B. (1970 ). Theory for Measurement of Competing Risks in Animal Experi- ments. Mathematical Biosciences ,6,1—11. Ananth, C. V., and Kleinbaum, D. G. (1977 ). Regression Models for Ordinal Data: A Review of Methods and Application. InternationalJournalof Epidemiology ,26, 1323—1333. Andersen,P. K. (1982 ). Testing Goodness of Fit of Cox’sRegression Model. Biometrics , 38,6 7—77. Andersen, P. K. (1992 ). Repeated Assessment of Risk Factors in Survival Analysis. Statistical Methods in Medical Research ,1,2 97—315 Andersen, P. K., Borgan, O., Gill, R. D., and Keiding, N. (1993 ).Statistical Models Based on Counting Processes. Springer-Verlag, New York. Andersen, P. K., and Gill, R. D. (1982 ). Cox’s Regression Model Counting Process: A Large Sample Study. Annals of Statistics ,10, 1100—1120 Anderson,J. A. (1972 ). SeparateSampleLogisticDiscrimination, Biometrika ,59,1 9—35. Andrews,D. F., and Herzberg,A. M. (1985 ).Data: A Collection of Problems from Many Fields for the Student and Research Worker . Springer-Verlag, New York. ARIC Investigators. (1989 ). The Atherosclerosis Risk in Communities (ARIC )Study: Design and Objectives. American Journal of Epidemiology ,129, 687—702. Arjas, E. (1988 ). A Graphical Method for Assessing Goodness of Fit in Cox’s Proportional Hazards Model. Journal of the American Statistical Association ,83, 204—212. Armitage, P. (1959 ). The Comparison of Survival Curves. Journal of the Royal Statistical Society, Series A ,122, 279—300. Armitage, P. (1971 ).Statistical Methods in Medical Research. Blackwell Scientific Publications, Oxford. Armitage, P. (1981 ). Importance of Prognostic Factors in the Analysis of Data from Clinical Trials. Controlled Clinical Trials ,1, 347—353. Armitage, P., and Gehan, E. A. (1974 ). Statistical Methods for the Identification and Use of Prognostic Factors. International Journal of Cancer ,13,1 6—35. Asal, N. R., Geyer, J. R., Risser, D. R., Lee, E. T., Kadamani, S., and Cherng, N. (1988a ). Risk Factors in Renal Cell Carcinoma, Part I. Methodology, Demographics,Tobacco, Beverage and Obesity. Cancer Detection and Prevention ,11, 359—377. Barnard, G. A. (1963 ). Some Aspects of the Fiducial Argument. Journal of the Royal Statistical Society, Series B ,34, 216—217. Bartholomew, D. J. (1957 ). A Problem in Life Testing. Journal of the American Statistical Association ,52, 350—355. Bartholomew, D. J. (1963 ). The Sampling Distribution of an Estimate Arising in Life Testing. Technometrics ,5, 361—374. Baumgartner, R. N., Roche, A. F., et al. (1987 ). Fatness and Fat Patterns: Associations withPlasma Lipidsand Blood Pressure in Adults,18 to 57 Years of Age. American Journal of Epidemiology ,126, 614—628. 489 Beale, E. M. L., Kendall, M. G., and Mann, D. W. (1976 ). The Discarding of Variable in Multivariate Analysis. Biometrika ,54, 357—366. Berkson, J. (1942 ). The Calculation of Survival Rates, in Carcinoma and Other Malignant Lesions of the Stomach , edited by W. Walters, H. K. Gray, and J. T. Priestley. W.B. Saunders, Philadelphia. Berkson, J., and Gage, R. R. (1950 ). Calculation of Survival Rates for Cancer. Proceedings of Staff Meetings, Mayo Clinic ,25, 250. Birnbaum, Z. W., and Saunders, S. C. (1958 ). A Statistical Model for Life-Length of Materials. Journal of the American Statistical Association ,53, 151—160. Blackstone, E. H., and Lytle, B. W. (2000 ). Competing Risks after Coronary Bypass Surgery: The Influence of Death on Reintervention. Journal of Thoracic and Cardiovascular Surgery ,119(6), 1221—1230. Bliwise, D. L., Kutner, N. G., Zhang, R., and Parker, K. P. (2002 ). Survival by Time of Day of Hemodialysis in an Elderly Cohort. Journal of the American Medical Association ,286(21), 2690—2694. Boag, J. W. (1949 ). Maximum Likelihood Estimates of Proportion of Patients Cured by Cancer Therapy. Journal of the Royal Statistical Society, Series B ,11, 15. Bolard, P., Quantin, C. P., Esteve, J., Faivre, J., and Abrahamowicz, M. (2001 ). Modeling Time-Dependent Hazard Ratios in Relative Survival: Application toColon Cancer. Journal of Clinical Epidemiology ,54(10)986—996. Bonadonna, G., et al. (1976 ). Combination Chemotherapy as an Adjuvant Treatment in Operable Breast Cancer. New England Journal of Medicine ,294, 405—410. Brancato, G., Pezzotti, P., Rapiti, E., Perucci, C. A., Abeni, D., Babbalacchio, A., and Rezza, G. (1997 ). Multiple Imputation Method for Estimating Incidence of HIV Infection: The Multicenter Prospective HIV Study. International Journal of Epi- demiology ,26(5), 1107—1114. Breslow, N. (1970 ). A Generalized Kruskal —Wallis Test for Comparing KSamples Subject to Unequal Pattern of Censorship. Biometrika ,57, 579—594. Breslow, N. (1974 ). Covariance Analysis of Survival Data under the Proportional Hazards Model. International Statistical Review ,43,4 3—54. Breslow, N. E. (1975 ). Analysis of Survival Data under the Proportional Hazards Model. International Statistical Review ,43,4 5—48. Breslow, N. E., and Crowley, J. (1974 ). A Large Sample Study of the Life Table and Product Limit Estimates under Random Censoring. Annals of Statistics ,2, 437—453. Breslow, N. E., and Day, N. E. (1980 ).Statistical Methods in Cancer Research , Vol. 1, The Analysis of Case-Control Studies . International Agency for Research on Cancer, Lyon, France. Breslow,N., and Powers, W. (1978 ). Are ThereTwo LogisticRegressionsfor Retrospec- tive Studies? Biometrics ,34, 100—105. Breslow, N. E., Day, N. E., Halvorsen, K. T., Prentice, R. L., and Sabai, C. (1978 ). Estimationof Multiple Relative Risk Functions in Matched Case-Control Studies. American Journal of Epidemiology ,108,2 99—307. Broadbent, S. (1958 ). Simple Mortality Rates. Journal of Applied Statistics ,7, 86.490  Broderick, A., Mori, M., Nettleman, M. D., Streed, S. A., and Wenzel, R. P. (1990 ). Nosocomial Infections: Validation of Surveillance and Computer Modeling toIdentify Patients at Risk. American Journal of Epidemiology ,131, 734—742. Brookmeyer, R., and Goedert, J. J. (1989 ). Censoring in an Epidemic with an Application to Hemophilia-Associated AIDS. Biometrics ,45, 325—335. Brown, B. W., and Hollander, M. (1977 ).Statistics: A Biomedical Introduction . Wiley, New York. Brown, C. C. (1982 ). On a Goodness-of-FitTest for the Logistic Model Based on Score Statistics. Communications in Statistics ,11, 1087—1105. Brown, G. W., and Flood, M. M. (1947 ). Tumbler Mortality. Journal of the American Statistical Association ,42, 562—574. Burdette, W. J., and Gehan, E. A. (1970 ).PlanningandAnalysisof ClinicalStudies. Charles C. Thomas, Springfield, IL. Buzdar, A. U., Gutterman, J. U., Blumehscein, G. R., Hortobagiji, G. H., Tashima, C. K., Smith, T. L, Hersh, E. M., Freiriech, E. J., and Gehan, E. A. (1978 ). Intensive Postoperative Chemoimmunotherapy for Patients with Stage II and Stage IIIBreast Cancer. Cancer,41, 1064—1075. Byar, D. P. (1974 ). Selecting Optimum Treatment in Clinical Trials Using Covariate Information. Presented at the 1974 Annual Meeting of the American Statistical Association, August 28. Byar, D. P. (1980 ). The Veterans Administration Study of Chemoprophylaxis for Recurrent Stage I Bladder Tumors: Comparisons of Placebo, Pyridoxine, andTopical Thiotepa, In Bladder Tumors and Other Topics in Urological Oncology , edited by M. Pavone-Macaluso, P. H. Smith, and F. Edsmyn. Plenum Press, NewYork, pp. 363—370. Byar,D.P.,Huse,R., andBailar,J. C.III, andthe VeteransAdministrationCooperative Urological Research Group (1974 ). An Exponential Model Relating Censored Survival Data and Concomitant Information for Prostatic Cancer Patients.Journal of the National Cancer Institute ,52, 321—326. Byers, R. H. Jr., Morgan, W. M., Darrow, W. W., Doll, L., Jaffe, H. W., Rutherford, G., Hessol, N., and O’Malley, P. M., (1988 ). Estimating AIDS Infection Rates in the San Francisco Cohort. AIDS,2(3), 207—210. Carbone, P., Kellerhouse, L., and Gehan, E. (1967 ). Plasmacytic Myeloma: A Study of the Relationship of Survival to Various Clinical Manifestations and Anomalous Protein Type in 112 Patients. American Journal of Medicine ,42,93 7—948. Carnahan, B., Luther, H. A., and Wilkes, J. O. (1969 ).Applied Numerical Methods. Wiley, New York. Carter, S. K., Oleg, S., and Slavik, M. (1977 ). Phase I Clinical Trials, in Methods of Development of New Anticancer Drugs . National Cancer Institute Monograph 45. U.S. Department of Health, Education, and Welfare Publication (NIH )76—1037. National Cancer Institute, Bethesda, MD. Chernoff, H., and Leiberman, G. J. (1954 ). Use of Normal Probability Paper. Journal of the American Statistical Association ,49, 778—785. Chiang, C. L. (1961 ). Standard Error of the Age-Adjusted Death Rate. Vital Statistics: Special Reports, Selected Studies ,47, 9. U.S. Department of Health, Education, and Welfare, Washington, DC. 491 Chiang, C. L. (1968 ).Introduction to Stochastic Processes in Biostatistics. Wiley, New York. Chiasson, M. A., Stoneburner, R. L., et al. (1990 ). Risk Factors for Human Immunode- ficiency Virus Type 1 (HIV-1 )Infection in Patients at a Sexually Transmitted DiseaseClinicin NewYorkCity. American Journal of Epidemiology ,131,208—220. Clayton, D., and Cuzick, J. (1985 ). The Em algorithm for Cox’s regression model using GLIM.AppliedStatistics ,34, 148—156. Cochran, W. G., and Cox, G. M. (1957 ).Experimental Designs , 2nd ed. Wiley, New York. Cohen, A. C., Jr. (1951 ). Estimating Parameters of Logarithmic-Normal Distributions by Maximum Likelihood. Journal of the American Statistical Association ,46, 206—212. Cohen, A. C., Jr. (1959 ). Simplified Estimators for the Normal Distribution When Samples Are Singly Censored or Truncated. Technometrics ,1(3), 217—237. Cohen, A. C., Jr. (1961 ). Table for Maximum Likelihood Estimates: Singly Truncated and Singly Censored Samples. Technometrics ,3, 535—541. Cohen, A. C., Jr. (1963 ). Progressively Censored Sample in Life Testing. Technometrics , 5, 327—339. Cohen, A. C., Jr. (1976 ). Progressively Censored Sampling in the Three Parameter Log-Normal Distribution. Technometrics ,18. Cohen, J., and Cohen, P. (1975 ).Applied Multiple Regression/Correlation Analysis for the Behavioral Sciences. Lawrence Erlbaum Associates, Hillsdale, NJ. Collett, D. (1991 ).Modelling Binary Data . Chapman & Hall, London. Collins, J. A., Garner, J. B., Wilson, E. H., Wrixon, W., and Casper, R. F. (1984 ).A Proportional Hazards Analysis of the Clinical Characteristics of Infertile Couples.American Journal of Obstetrics and Gynecology ,148, 527—532. Connelly, R. R., Cutler, S. J., and Baylis, P. (1966 ). End Result in Cancer of the Lung: ComparisonofMaleand FemalePatients. Journal of the National Cancer Institute , 36, 277—287. Cornfield, J. (1951 ). A Method of Estimating Comparative Rates from Clinical Data: Applications to Cancer of the Lung, Breast and Cervix. Journal of the National Cancer Institute ,11, 1269—1275. Cornfield, J. (1956 ). A Statistical Problem Arising from Retrospective Studies, in Proceedings of the 3rd Berkeley Symposium on Mathematical Statistics andProbability , Vol. 4, edited by J. Neyman. University of California Press, Berkeley, CA, 135—148. Cornfield, J. (1962 ). Joint Dependence of Risk of Coronary Heart Disease in Serum Cholesterol and Systolic Blood Pressure: A Discriminant Function Analysis.Federation Proceedings ,21,5 8—61. Correa, P., Pickle, L. W., Fortham, E., et al. (1983 ). Passive Smoking and Lung Cancer. Lancet,2,5 95—597. Cox, D. R. (1961 ). Tests of Separate Families of Hypotheses. Proc.FourthBerkeley SymposiuminMathematicalStatistics , I, Berkeley: University of California Press, 105—123.492  Cox, D. R. (1962 ). Further Results on Tests of Separate Families of Hypotheses. J.R. Stat.Soc.B ,24, 406—424. Cox, D. R. (1953 ). Some Simple Tests for Poisson Variates. Biometrika ,40, 354—360. Cox, D. R. (1959 ). The Analysis of Exponentially Distributed Life-Times with Two Types of Failures. Journal of the Royal Statistical Society, Series B ,21, 411—421. Cox, D. R. (1962 ).Renewal Theory . Methuen, London. Cox, D. R. (1964 ). Some Applications of Exponentially Distributed Life-Times with Two Types of Failures. Journal of the Royal Statistical Society, Series B ,26, 103—110. Cox, D. R. (1970 ).Analysis of Binary Data. Methuen, London. Cox, D. R. (1972 ). Regression Models and Life Tables. Journal of the Royal Statistical Society, Series B ,34, 187—220. Cox, D.R., and Hinkley,D. V. (1974 ).TheoreticStatistics ,Chapmanand Hall,London. Cox, D. R., and Oakes, D. (1984 ).Analysis of Survival Data . Chapman & Hall, New York. Cox, D. R., and Snell, E. J. (1968 ). A General Definition of Residuals. Journal of the Royal Statistical Society, Series B ,30, 248—275. Cox, D. R., and Snell, E. J. (1989 ).The Analysis of Binary Data, 2nd ed . Chapman & Hall, London. Crawford, S. L., Tennstedt,S. L., and McKinlay, J. B. (1995 ). A Comparison of Analytic Methods for Non-random Missingness of Outcome Data. J.ClinEpidemiol ,48, 209—219. Crist, W., Boyett, J., and Jackson, J., et al. (1989 ). Prognostic Importance of the Pre-B-CellImmunophenotypeand Other PresentingFeatures in B-Lineage Child-hood Acute Lymphoblastic Leukemia: A Pediatric Oncology Group Study. Blood, 74, 1252—1259. Crowley,J.,and Hu,M. (1977 ).CovarianceAnalysisofHeartTransplantSurvivalData. Journal of the American Statistical Association ,72,2 7—36. Crowley, J., and Thomas, D. R. (1975 ). Large Sample Theory for the Log Rank Test. Technical Report 415 . Department of Statistics, University of Wisconsin, Madison, WI. Cutler, S. J., and Ederer, F. (1958 ). Maximum Utilization of the Life Table Method in Analyzing Survival. Journal of Chronic Diseases ,8,6 99—712. Cutler, S. J., Griswold, M. H., and Eisenberg, H. (1957 ). An Interpretation of Survival Rates: Cancer of the Breast. Journal of the National Cancer Institute ,19, 1107— 1117. Cutler, S. J., Ederer, F., Griswold, M. H., and Greenberg, R. A. (1959 ). Survival of Breast-Cancer Patients in Connecticut, 1935 —54.Journal of the National Cancer Institute,23, 1137—1156. Cutler, S. J., Ederer, F., Griswold, M. H., and Greenberg, R. A. (1960a ). Survival of Patients with Uterine Cancer, Connecticut, 1935 —54.Journal of the National Cancer Institute ,24, 519—539. Cutler, S. J., Ederer, F., Griswold, M. H., and Greenberg, R. A. (1960b ). Survival of Patients with Ovarian Cancer, Connecticut, 1935 —54.Journal of the National Cancer Institute ,24, 541—549. 493 Cutler, S. J., Axtell, L., and Heise, H. (1967 ). Ten Thousand Cases of Leukemia: 1940—62.Journal of the National Cancer Institute ,39,993—1026. Daniel, C. (1959 ). Use of Half-Normal Plots in Interpreting Factorial Two-Level Experiments. Technometrics ,1, 311—341. Daniel, W. W. (1987 ).Biostatistics: A Foundation for Analysis in the Health Sciences. Wiley, New York. Davis, D. J. (1952 ). An Analysis of Some Failure Data. Journal of the American Statistical Association ,47, 113—150. Davis, H. T., and Feldstein, M. L. (1979 ). The Generalized Pareto Law as a Model for Progressively Censored Survival Data. Biometrika ,66,2 99—306. Dawber, T. R. (1980 ).The Framingham Study. Harvard University Press, Cambridge, MA. Dawber, T. R., Meadors, G. F., and Moore, F. E. Jr. (1951 ). Epidemiological Ap- proaches to Heart Disease: The Framingham Study. American Journal of Public Health,41, 279—286. Delong, D. M., Guirguis, G. H., and So, Y. C. (1994 ). Efficient Computation of Subset Selection Probablilities with Application to Cox Regression. Biometrica. 81 607—611. Dharmalingam, A., Pool, I., and Dickson, J. (2000 ). Biosocial Determinants of Hyster- ectomy in New Zealand. American Journal of Public Health ,90(9), 1455—1458. Dixon,W. J.,Brown, M.B., Engelman,L.,Hill,M. A., and Jennrich,R.I. (1990 ).BMDP Statistical Software Manual . University of California Press, Berkeley, CA. Draper, N. R., and Smith, H. (1966 ).Applied Regression Analysis . Wiley, New York. Drenick, R. F. (1960 ). The Failure Law of Complex Equipment. Journal of Social and Industrial Applied Mathematics ,8, 680. Dunn, O. J. (1964 ). New Table for Multiple Comparisons with a Control. Biometrics , 20, 482—491. Ederer, F., Axtell, L. M., and Cutler, S. J. (1961 ). The Relative Survival Rate: A Statistical Methodology. National Cancer Institute Monographs ,6, 101—121. Efron, B. (1975 ). The Efficiency of Logistic Regression Compared to Normal Dis- criminant Analysis. Journal of the American Statistical Association ,70,8 92—898. Efron, B. (1977 ). The Efficiency of Cox’s Likelihood Function for Censored Data. Journal of the American Statistical Association ,72, 557—565. Efron, B. (1994 ). Missing Data, Imputation, and the Bootstrap. JournaloftheAmerican StatisticalAssociation ,89, 463—475. Eisenberger, M., Krasnow, S., Ellenberg, S., et al. (1989 ). A Comparison of Carboplatin Plus Methotrexate versus Methotrexate Alone in Patients with Recurrent and Metastatic Head and Neck Cancer. Journal of Clinical Oncology ,7, 1341—1345. Elaad,E., andBen-Shakhar,G. (1989 ).Effects of MotivationandVerbalResponseType on Psychophysiological Detection of Information. Psychophysiology ,26, 442—451. Elandt-Johnson, R. C., and Johnson, N. L. (1980 ).Survival Models and Data Analysis . Wiley, New York. Enas, G. G., Dornseit, B. E., Sampson, C. B., Rockhold, F. W., and Wuu, J. (1989 ). Monitoring versus Interim Analysis of Clinical Trials: A Perspective from thePharmaceutical Industry. Controlled Clinical Trials ,10,5 7—70.494  Epstein, B. (1958 ). The Exponential Distribution and Its Role in Life Testing. Industrial Quality Control ,15,2—7. Epstein, B. (1960a ). Estimation of the Parameters of Two Parameter Exponential Distribution from Censored Samples. Technometrics ,2, 403—406. Epstein, B. (1960b ). Estimation from Life Test Data. Technometrics ,2, 447—454. Epstein, B., and Sobel, M. (1953 ). Life Testing. Journal of the American Statistical Association ,48, 486—502. Farewell, V. T. (1979 ). Some Results on the Estimation of Logistic Models Based on Retrospective Data. Biometrika ,66,2 7—32. Farrington, C. P. (2000 ). Residuals for Proportional Hazards Models with Interval- Censored Survival Data. Biometrics ,56(2), 473—482. Feigl, P., and Zelen, M. (1965 ). Estimation of Exponential Survival Probabilities with Concomitant Information. Biometrics ,21, 826—838. Feinleib, M. (1960 ). A Method of Analyzing Log-Normally Distributed Survival Data with Incomplete Follow-up. Journal of the American Statistical Association ,55, 534—545. Feinleib, M., and MacMahon, B. (1960 ). Variation in the Duration of Survival of Patients with Chronic Leukemias. Blood,17, 332—349. Feskanich, D., Singh, V., Willett, W. C., and Colditz, G. A. (2002 ). Vitamin A Intake and Hip Fractures among Postmenopausal Women. Journal of the American Medical Association ,287(1)47—54. Fish, E. B., Chapman, J. A. and Link, M. A. (1998 ). Competing Causes of Death for Primary Breast Cancer. Annals of Surgical Oncology ,5(4), 368—375. Fisher, R. A. (1922 ). On the Mathematical Foundation of Theoretical Statistics. Philosophical Transactions of the Royal Society of London, Series A ,222. Fisher, R. A. (1936 ). The Use of Multiple Measurements in Toxonomic Problems. Annals of Eugenics ,7, 312—330. Fleiss, J. L. (1979 ). Confidence Intervals for the Odds Ratio in Case-Control Studies: The State of the Art. Journal of Chronic Diseases ,32,6 9—82. Fleiss, J. L. (1981 ).Statistical Methods for Rates and Proportions. Wiley, New York. Fleming, T. R., and Harrington, D. P. (1979 ). Non-parametric Estimation of the Survival Distribution in Censored Data. Unpublished manuscript. Fleming, T. R., and Harrington, D. P. (1991 ).Counting Processes and Survival Analysis . Wiley, New York. Fleming, T. R., O’Fallon, J. R., O’Brian, P. C., and Harrington, D. P. (1980 ). Modified Kolmogorov—Smirnov Test Procedures with Application to Arbitrarily Right Censored Data. Biometrics ,36, 607—626. Fleming, T. R., Harrington, D. P., and O’Brien, P. C. (1984 ). Designs for Group Sequential Tests. Controlled Clinical Trials ,5, 348—361. Florin, V., and Ronghui, X. (2000 ). Proportional Hazards Model with Random Effects. Statistics in Medicine ,19(24), 3309—3324. Fraser, D. A. S. (1968 ).The Structure of Inference. Wiley, New York. Freedman, L. S. (1982 ). Tables of the Number of Patients Required in Clinical Trials Using the Log Rank Test. Statistics in Medicine ,1, 121—129. 495 Frei, E., et al. (1961 ). Studies of Sequential and Combination Antimetabolite Therapy in Acute Leukemia: 6 —Mercaptopurine and Methotrexate. Blood,18, 431—454. Freireich, E. J., Gehan, E. A., Frei, E., et al. (1963 ). The Effect of 6-Mercaptopurine on the Duration of Steroid-Induced Remissions in Acute Leukemia: A Model for Evaluation of Other Potential Useful Therapy. Blood,21(6),6 99—716. Freireich, E. J., Gehan, E. A., Rall, D. P., Schmidt, L. H., and Skipper, H. E. (1966 ). Quantitative Comparison of Toxicity of Anticancer Agents in Mouse, Rat,Hamster, Dog, Monkey, and Man. Cancer Chemotherapy Report ,50,4 . Freireich, E. J., Gehan, E. A., Bodey, G. P., Hersh, E. M., Hart, J. S., Gutterman, J. U., and McCredie, K. B. (1974 ). New Prognostic Factors Affecting Response and Survival in Adult Leukemia. Transactions of the Association of American Phys- icians,87,2 98—305. Friedman, L. M., Furberg, C. D., and DeMets, D. L. (1985 ).Fundamentals of Clinical Trials, 2nd ed. PSG Publishing, Littleton, MA. Gaddum, J. H. (1945a ). Log Normal Distributions. Nature, London ,156, 463. Gaddum, J. H. (1945b ). Log Normal Distributions. Nature, London ,156, 747. Gail, M., and Gart, J. J. (1973 ). The Determination of Sample Sizes for Use with the Exact Conditional Test in 2 /p592 Comparative Trials. Biometrics ,29, 441—448. Gail, M. H., Lubin, J. H., and Rubinstein, L. V. (1981 ). Likelihood Calculations for Matched Case-Control Studies and Survival Studies with Tied Death Times. Biometrika ,68, 703—707. Gajjar, A. V., and Khatri, C. G. (1969 ). Progressively Censored Samples from Log- Normal and Logistic Distributions. Technometrics ,11,7 93—803. Garside, M. J. (1965 ). The Best Sub-set in Multiple Regression Analysis. Applied Statistics ,14,1 96—200. Gehan, E. A. (1965a ). A Generalized Wilcoxon Test for Comparing Arbitrarily Singly-Censored Samples. Biometrika ,52, 203—223. Gehan, E. A. (1965b ). A Generalized Two-Sample Wilcoxon Test for Doubly-Censored Data. Biometrika ,52, 650—653. Gehan, E. A. (1970 ). Unpublished notes on survival time studies. The University of Texas M. D. Anderson Cancer Center, Houston, Texas. Gehan, E. A. (1969 ). Estimating Survival Function from the Life Table. Journal of Chronic Diseases ,21, 629—644. Gehan, E. A., and Thomas, D. G. (1969 ). The Performance of Some Two-Sample Tests in Small Samples with and without Censoring. Biometrika ,56, 127—132. Gelenberg, A. J., Kane, J. M., Keller, M. B., et al. (1989 ). Comparison of Standard and Low Serum Levels of Lithium for Maintenance Treatment of Bipolar Disorder. New England Journal of Medicine ,321, 1489—1493. George, S. L., Fernback, D. J., et al. (1973 ). Factors Influencing Survival in Pediatric Acute Leukemia: The SWCCSG Experience, 1959 —1970. Cancer,32, 1542—1553. Gertsbakh, I. B. (1989 ).Statistical Reliability Theory. Marcel Dekker, New York. Gill, R., and Schumacher, M. (1987 ). A Simple Test of the Proportional Hazards Assumption. Biometrika ,74, 289—300.496  Gillum, R. F., Fortmann, S. P., Prineas, R. J., and Kottke, T. E. (1984 ). International Diagnostic Criteria for Acute Myocardial Infarction and Acute Stroke. American Heart Journal ,108, 150—158. Glasser, M. (1967 ). Exponential Survival with Covariance. Journal of the American Statistical Association ,62, 561—568. Gompertz, B. (1825 ). On the Nature of the Function Expressive of the Law of Human Mortality and on the New Mode of Determining the Value of Life Contingencies.Philosophical Transactions ,513. Gore, S. M. (1983 ). Graft Survival after Renal Transplantation: Agenda for Analysis. KidneyInt. ,24, 516—525. Grambsch, P M., Therneau, T. M. (1994 ). Proportional Hazards Tests in Diagnostics Based on Weighted Residuals. Biometrika ,81, 515—526. Gray, R. J. (1990 ). Some Diagnostic Methods for Cox Regression Models through Hazard Smoothing. Biometrics ,46,93—102. Green, P. J. (1984 ). Iteratively Reweighted Least Squares for Maximum Likelihood Estimation, and Some Robust and Resistant Alternatives (with discussion ).Jour- nal of the Royal Statistical Society ,46(2), 149—192. Greenwood, J. A., and Durand, D. (1960 ). Aids for Fitting the Gamma Distribution by Maximum Likelihood. Technometrics ,2,5 5—65. Greenwood, M. (1926 ). The Natural Duration of Cancer. Reports on Public Health and Medical Subjects , Her Majesty’s Stationary Office, London, 33,1—26. Griswold, M. H., and Cutler, S. J. (1956 ). The Connecticut Cancer Register: Seventeen Years of Experience. Connecticut Medical Journal ,20, 366—372. Griswold, M. H., Wilder, C. S., Cutler, S. J., and Pollack, E. S. (1955 ).Cancer in Connecticut, 1935 —1951. Monograph. Connecticut State Department of Health, Hartford, CT. Grizzle, J. E. (1967 ). Continuity Correction in the /afii9851/p17-Test for 2 /p592 Tables. American Statistician ,21,2 8—32. Gross, A. J., and Clark, V. A. (1975 ).Survival Distributions: Reliability Applications in the Biomedical Sciences . Wiley, New York. Grove, R. D., and Hetzel, A. M. (1963 ).Vital Statistics Rates in the United States, 1940—1960.National Center for Health Statistics, Washington, DC. Gupta, A. K. (1952 ). Estimation of the Mean and Standard Deviation of a Normal Population from a Censored Sample. Biometrika ,39, 260—273. Gupta, S. S. (1960 ). Order Statistics from the Gamma Distribution. Technometrics ,2, 243—262. Hagar,H. W.,and Bain,L. J. (1970 ). InferentialProceduresforthe GeneralizedGamma Distribution. Journal of the American Statistical Association ,65, 1601—1609. Hahn, G. J., and Shapiro, S. S. (1967 ).StatisticalModelsinEngineering . Wiley, New York. Haldane, J. B. S. (1956 ). The Estimation and Significance of the Logarithm of a Ratio of Frequencies. Annals of Human Genetics ,20, 309—311. Halperin, M. (1952 ). Maximum Likelihood Estimation in Truncated Samples. Annals of Mathematical Statistics ,23, 226—238. 497 Halperin, M., Blackwelder, W. C., and Verter, J. I. (1971 ). Estimation of the Multivari- ate Logistic Risk Function: A Comparison of the Discriminant Function andMaximum Likelihood Approaches. Journal of Chronic Diseases ,24, 125—158. Hammond, I. W., Lee, E. T., Davis, A. W., and Booze, C. F. (1984 ). Prognostic Factors Related to Survival and Complication-Free Times in Airmen Medically Certified after Coronary Bypass Surgery. Aviation, Space, and Environmental Medicine , April, pp. 321—331. Hannan, E. J. (1979 ). The Determination of the Order of an Autoregression. Journal of the Royal Statistical Society, Series B ,41,1 90—195. Hanson, B. S., Isacsson, S-O., Janzon, L., and Lindell, S. E. (1989 ). Social Network and Social Support Influence Mortality in Elderly Men. American Journal of Epi- demiology ,130, 100—111. Harrison,J.D., Jones,J.A.,and Morris,D.L. (1990 ).TheEffectofthe GastrinReceptor Antagonist Proglumide on Survival in Gastric Carcinoma. Cancer,66, 1449—1452. Hart, J. S., George, S. L., Frei, E., Bodey, G. P., Nickerson, R. C., and Freireich, E. J. (1977 ). Prognostic Significance of Pretreatment Proliferative Activity in Adult Acute Leukemia. Cancer,39, 1603—1617. Harter, H. L., and Moore, A. H. (1965 ). Maximum Likelihood Estimation of the Parameters of Gamma and Weibull Populations from Complete and from Cen-sored Samples. Technometrics ,7, 639—643. Harter, H. L., and Moore, A. H. (1966 ). Local Maximum Likelihood Estimation of the Parameters of Three-Parameter Log-Normal Population from Complete and Censored Sample. Journal of the American Statistical Association ,61, 842—851. Harter, H. L., and Moore, A. H. (1967 ). Asymptotic Variance and Covariances of Maximum Likelihood Estimators, from Censored Samples, of the Parameters ofWeibull and Gamma Parameters. Annals of Mathematical Statistic s,38, 557—570. Hastings, N. A. J., and Peacock, J. B. (1974 ).Statistical Distributions . Butterworth, London. Hauck, W. W., Jr., and Donner, A. (1977 ). Wald’s Test as Applied to Hypotheses in Logit Analysis. Journal of the American Statistical Association ,72, 851—853. Haughton,D. M. A. (1988 ). On the Choice of a Model to Fit Data froman Exponential Family. Annals of Statistics ,16, 342—355. Heitjan, D. F. (1997 ). Annotation: What Can be Done About Missing Data? Ap- proaches to imputation, Am.J.PublicHealth ,87, 548—550. Hemstreet, G. P., Yin, S., Ma, Z., Bonner, R. B., Bi, W., Rao, J. Y., Zang, M., Zheng, Q., Bane, B., Asal, N., Li, G., Feng, P., Hurst, R. E., and Wang, W. (2001 ). Biomarker Risk Assessment and Bladder Cancer Detection in a Cohort Exposedto Benzidine,JournaloftheNationalCancerInstitute ,93, 427—436. Hill, A. B. (1960a ).Controlled Clinical Trials. Blackwell Scientific, Oxford Hill, A. B. (1960b ).Statistical Methods in Clinical and Preventive Medicine. Oxford University Press, Oxford. Hill, A. B. (1971 ).Principles of Medical Statistics. Oxford University Press, New York. Hirayama, T. (1981 ). Non-smoking Wives of Heavy Smokers Have a Higher Risk of Lung Cancer: A Study from Japan. British Medical Journal ,282, 183—185.498  Hoel, D. G., Sobel, M., and Weiss, G. H. (1975 ). A Survey of Adaptive Sampling for Clinical Trials. Perspectives in Biometrics ,1,2 9—61. Holford, T. R., White, C., and Kelsey, J. L. (1978 ). Multivariate Analysis for Matched Case-Control Studies. American Journal of Epidemiology ,107, 245—256. Hollander, M., and Proschan, F. (1979 ). Testing to Determine the Underlying Distribu- tion Using Randomly Censored Data. Biometrics ,35,3 93—401. Hollander, M., and Wolfe,D. A. (1973 ).Nonparametric Statistical Methods. Wiley, New York. Horner, R. D. (1987 ). Age at Onset of Alzheimer’s Disease: Clue to the Relative ImportanceofEtiologicFactors? American Journal of Epidemiology ,126,409—414. Hosmer, D. W., and Lemeshow, S. (1980 ). A Goodness-of-Fit Test for the Multiple Logistic Regression Model. CommunicationsinStatistics ,A10, 1043—1069. Hosmer,D.W., and Lemeshow,S. (1999 ).Applied Survival Analysis , 2nd ed.Wiley, New York. Hosmer, D. W., and Lemeshow, S. (2000 ).Applied Logistic Regression. Wiley, New York. Howell, D. W. (1987 ).Statistical Methods for Psychology. Duxbury Press, Boston. Hung, C. T., Lim, J. K. C., and Zoest, A. R. (1988 ). Optimization of High-Performance Liquid Chromatographic Analysis for Isoxazolye Penicillins Using FactorialDesign. Journal of Chromatography ,425, 331—341. Ibrahim, J. G., Chen, M. H., and Sinha, D. (2001 ).Bayesian Survival Analysis. Springer-Verlag, New York Ingram, D. D., and Kleinman, J. C. (1989 ). Empirical Comparisons of Proportional Hazards and Logistic Regression Models. Statistics in Medicine ,8, 525—538. Irwin, J. O. (1949 ). The Standard Error of an Estimate of Expectational Life. Journal of Hygiene ,47, 188—189. Jenkins, S. P. (1997 ). Discrete Time Proportional Hazards Regression. Stata Technical Bulletin,39,1 7—32. Jennings, D. E. (1986 ). Judging Inference Adequacy in Logistic Regression. Journal of the American Statistical Association ,81, 471—476. Johnson, N. L., and Kotz, S. (1970a ).Distributions in Statistics: Continuous Univariate Distributions (Vol. 1 )Houghton Mifflin, Boston. Johnson, N. L., and Kotz, S. (1970b ).Distributions in Statistics: Continuous Univariate Distributions. (Vol. 2 )Houghton Mifflin, Boston Johnson, P., and Pearce, J. M. (1990 ). Recurrent Spontaneous Abortion and Polycystic Ovarian Disease: Comparison of Two Regimens to Induce Ovulation. British Medical Journal ,300, 154—156. Kahn,H. A. (1983 ).An Introduction to Epidemiologic Methods. OxfordUniversityPress, New York. Kalbfleisch, J. D. (1974 ). Some Extensions and Applications of Cox’s Regression and Life Model. Presented at the joint meeting of the Biometric Society and theAmerican Statistical Association, Tallahassee, FL, March 20 —22. Kalbfleisch, J. D., and Prentice, R. L. (1973 ). Marginal Likelihoods Based on Cox’s Regression and Life Table Model. Biometrika ,60, 267—278. 499 Kalbfleisch, J. D., and Prentice, R. L. (1980 ).The Statistical Analysis of Failure Time Data. Wiley, New York. Kao, J. H. K. (1958 ). Computer Methods for Estimating Weibull Parameters in ReliabilityStudies. I.R.E. Transactions on Reliability and Quality Control ,PGRQC 13,1 5—22. Kaplan, E. L., and Meier, P. (1958 ). Nonparametric Estimation from Incomplete Observations.Journalof theAmericanStatisticalAssociation ,53, 457—481. Kay, R. (1979 ). Proportional Hazard Regression Models and the Analysis of Censored Survival Data. Applied Statistics ,26, 227—237. Kay, R. (1984 ). Goodness of Fit Methods for the Proportional Hazards Model: A Review. Revue Epidemiologie et de Santé Publique ,32, 185—198. Kelsey, J. L., Thompson, W. D., and Evans, A. S. (1986 ).Methods in Observational Epidemiology. Oxford University Press, New York. Kessing, L.V., Olsen, E. W., and Andersen, P. K. (1999 ). Recurrence in Affective Disorder:AnalyseswithFrailty Models. American Journal of Epidemiology ,149(5), 404—411. King, J. R. (1971 ).Probability Charts for Decision Making. Industrial Press, New York. King, M., Bailey, D. M., Gibson, D. G., Pitha, J. V., and McCay, P. B. (1979 ). Incidence andGrowth ofMammaryTumors Induced by7,12-Dimethylbenz (/afii9825)antheaceneas Related to the Dietary Content of Fat and Antioxidant. Journal of the National Cancer Institute ,63, 656—664. Kitagawa, E. M. (1964 ). Standardized Comparisons in Population Research. Demogra- phy,1,2 96—315. Klein, J. P., and Moeschberger, M. L. (1997 )Survival Analysis. Springer-Verlag, New York. Kleinbaum, D. G. (1994 ).LogisticRegression:ASelf-LearningText. Springer-Verlag, New York. Kleinbaum, D. G., Kupper, L. L., and Muller, K. E. (1988 ).Applied Regression Analysis and Other Multivariate Methods, 2nd ed. PWS-Kent, Boston. Kodlin, D. (1967 ). A New Response Time Distribution. Biometrics ,23, 227—239. Krishna, I. P. V. (1951 ). A Non-parametric Method of Testing kSamples. Nature,167, 33. Kruskal, W. H., and Wallis, W. A. (1952 ). Use of Ranks in One-Criterion Variance Analysis. Journal of the American Statistical Association ,47, 583—621. Kuzma, J. W. (1967 ). A Comparison of Two Life Table Methods. Biometrics ,23,5 1—64. Lagakos, S. W. (1980 ). The Graphical Evaluation of Explanatory Variables in Propor- tional Hazard Regression Models. Biometrika ,68,93—98. Lan, K. K. G., and DeMets, D. L. (1983 ). Discrete Sequential Boundaries for Clinical Trials. Biometrika ,70, 659—663. Lan, K. K. G., and DeMets, D. L. (1989 ). Changing Frequency of Interim Analysis in Sequential Monitoring. Biometrics ,45, 1017—1020. Lassare, S. (2001 ). Analysis of Progress in Road Safety in Ten European Countries. Accident Analysis and Prevention ,33(6), 743—751. Lawless, J. F. (1982 ).Statistical Methods and Model for Lifetime Data . Wiley, New York.500  Lawless, J. F. (1983 ). Statistical Methods in Reliability. Technometrics ,25, 305—316. Lee, A. H., and Yau, K. K. (2001 )Determining the Effects of Patient Case Mix on Length of Hospital Stay: A Proportional Hazards Frailty Model Approach.Methods of Information in Medicine ,40(4), 288—292. Lee, E. T. (1980 ).Statistical Methods for Survival Data Analysis. Lifetime Learning Publications, Belmont, CA. Lee, E. T. (1992 ).StatisticalMethodsforSurvivalDataAnalysis , second edition, Wiley, New York. Lee, E. T., and Thomas, D. R. (1980 ). Confidence Interval for Comparing Two Life Distributions. IEEE Transactions on Reliability ,R-29,5 1—56. Lee, E. T., Desu, M. M., and Gehan, E. A. (1975 ). A Monte-Carlo Study of the Power of Some Two-Sample Tests. Biometrika ,62, 425—432. Lee, E. T., Ishmael, D. R., Bottomley, R. H., and Murray, J. L. (1982 ). An Analysis of Skin Tests and Their Relationship to Recurrence and Survival in Stage III andStage IV Melanoma Patients. Cancer,49, 2336—2341. Lee, E. T., Yeh, J. L., Cleves, M. A., and Shafer, D. (1988 ). Vascular Complications in Noninsulin Dependent Diabetic Oklahoma Indians. Diabetes,37(Suppl. 1 ). Lee, E. T., Lee, V. S., Lu, M., et al. (1992 ). Incidence and Risk Factors of Diabetic Retinopathy in Oklahoma Indians with NIDDM. Diabetes Care ,15, 1620—1627. Lee, E. T., Russell, D., Jorge, N., Kenny, S., and Yu, M. (1993 ). A Follow-up Study of Diabetic Oklahoma Indians: Mortality and Causes of Death. Diabetes Care ,16, 300—305. Leenen, F. H. H., Balfe, J. A., Pelech, A. N., et al. (1987 ). Postoperative Hypertension after Repair of Coarctation of Aorta in Children: Protective Effect of Propranolol. American Heart Journal ,113, 1164—1173. Lehmann,E. L. (1953 ). The Power of Rank Tests. Annals of Mathematical Statistics ,24, 23—43. Leitner, L. M., Roumy, M. Ruckebusch, M., Sutra, J. F. (1986 ). Monoamines and Their Catabolites in the Rabbit Carotid Body. Effets of reserpine, sympathectomy andcarotid sinus nerve section, EuropeanJournalofPhysiology ,406, 552—556. Lemeshow, S., and Hosmer, D. W. (1982 ). A Review of Goodness-of-Fit Statistics for Use in the Development of Logistic Regression Models. American Journal of Epidemiology ,115,92—106. Leyland-Jones, B., Donnelly, H., Groshen, S., Myskowski, P., Donner, A. L., Fanucchi, M., Fox, J., and the Memorial Sloan-Kettering Antiviral Working Group (1986 ). 2/p30-Fluror-5-Iodoarabinosylcytosine, A New Potent Antiviral Agent: Efficacy in ImmunosuppressedIndividuals with Herpes Zoster. Journal of Infectious Diseases , 154, 430—436. Liang, K. Y., Self, S. G., and Liu, X. (1990 ). The Cox Proportional Hazards Model with Change Point: An Epidemiologic Application. Biometrics ,46, 783—793. Liang, K. Y., Self, S. G., Bandeen-Roche, K. J., and Zeger, S. L. (1995 ). Some Recent Developments for Regression Analysis of Multivariate Failure Time Data. Life- time Data Analysis, 1, 403—415. Lieblein, J., and Zelen, M. (1956 ). Statistical Investigation of the Fatigue Life of Deep-Grove Ball Bearings. Journal of Research of the National Bureau of Stan- dards,57, 273—316. 501 Lilliefors, H. W. (1971 ). Reducing the Bias of Estimators of Parameters for the Erlang and Gamma Distribution. Unpublished manuscript. Lindley, D. V. (1968 ). The Choice of Variables in Multiple Regression. Journal of the Royal Statistical Society, Series B ,30,3 1—53. Linka, A. Z., Sklenar, J., Wei, K. I., Jayaweera,A. R., Skyba, D. M., and Kaul, S. (1998 ). Assessment of Transmural Distribution of Myocardial Perfusion with ContrastEchocardiography. Circulation 3 ;98(18); 1912—1920. Little, R. J., and Rubin, D. B. (1987 ).StatisticalAnalysiswithMissingData , John Wiley & Sons, New York. Liu, P. Y., and Crowley, J. (1978 ). Large Sample Theory of the MLE Based on Cox’s Regression Model for Survival Data. Technical Report 1 . Wisconsin Clinical Cancer Center (Biostatistics ), University of Wisconsin, Madison, WI. Lubin, J. H. (1981 ). A Computer Program for the Analysis of Matched Case-Control Studies. Computers and Biomedical Research ,14, 138—143. McAlister, D. (1879 ). The Law of the Geometric Mean. Proceedings of the Royal Society,29, 367. McCracken, D. D., and Dorn, W. S. (1964 ).Numerical Methods and Fortran Pro- gramming. Wiley, New York. McCullagh, P. (1980 ). Regression Model for Ordinal Data. Journal of the Royal Statistical Society ,42(2), 109—142. McCullagh, P., and Nelder, J. A. (1989 ).Generalized Linear Models . Chapman & Hall, London. McFadden, D. (1976 ). A Comment on Discriminant Analysis ‘‘versus’’ Logit Analysis. Annals of Economic and Social Measurement ,5, 511—523. Mackenbach, J. P., Kunst, A. E., Lautenbach, H., Bijlsma, F., and Oei, Y.B. (1995 ). Competing Causes of Death: An Analysis Using Multiple-Cause-of-Death Data from The Netherlands. 141(5), 466—475. Mafart, P., Couvert, O., Gaillard, S., and Leguerinel, I. (2002 ). On Calculating Sterility inThermalPreservationMethods:Applicationofthe WeibullFrequencyDistribu-tion Model. International Journal of Food Microbiology ,72(12); 107—113. Mann, H. B., and Whitney, D. R. (1947 ). On a Test of Whether One of Two Random Variables Is Stochastically Larger Than the Other. Annals of Mathematical Statistics ,18,5 0—60. Mann, N. R. (1970 ). Estimators and Exact Confidence Bounds for Weibull Parameters Based on a Few Ordered Observations. Technometrics ,12, 345—361. Mann, N. R., Schafer, R. E., and Singpurwalla, N. D. (1974 ).Methods for Statistical Analysis of Reliability and Life Data. Wiley, New York. Manninen,O. (1988 ). Changes in Hearing, CardiovascularFunctions, Haemodynamics, Upright Body Sway, Urinary Catecholamines and Their Correlates after Pro-longed Successive Exposure to Complex Environmental Conditions. International Archives of Occupational and Environmental Health ,60, 249—272. Mantel, N. (1966 ). Evaluation of Survival Data and Two New Rank Order Statistics Arising in Its Consideration. Cancer Chemotherapy Reports ,50, 163—170. Mantel,N. (1967 ). RankingProceduresfor Arbitrarily RestrictedObservations. Biomet- rics,23,6 5—78.502  Mantel, N. (1970 ). Why Stepdown Procedures in Variable Selection. Technometrics ,12, 621—625. Mantel, N. (1973 ). Synthetic Retrospective Studies and Related Topics. Biometrics ,29, 479—486. Mantel, N. (1977 ). Test and Limits for the Common Odds Ratio of Several 2 /p592 Contingency Tables: Methods in Analogy with the Mantel —Haenszel Procedure. Journal of Statistical Planning Information ,1, 179—189. Mantel, N., and Haenszel, W. (1959 ). Statistical Aspects of the Analysis of Data from Retrospective Studies of Disease. Journal of the National Cancer Institute ,22, 719—748. Mantel,N., and Hankey,B. F. (1978 ). A Logistic RegressionAnalysisof Response-Time Data Where the Hazard Function Is Time Dependent. Communications in Statis- tics A: Theory and Methods ,7, 333—347. Mantel, N., and Myers, M. (1971 ). Problems of Convergence of Maximum Likelihood Iterative Procedures in Multiparameter Situation. Journal of the American Statis- tical Association ,66, 484—491. Mantel, N., and Stark, C. R. (1968 ). Computation of Indirect Adjusted Rates in the Presence of Confounding. Biometrics ,24,997—1005. Marascuilo, L. A., and McSweeney, M. (1977 ).Nonparametric and Distribution-Free Methods for the Social Sciences. Brooks/Cole, Monterey, CA. Marubini, E., and Valsecchi, M. G. (1995 ).Analyzing Survival Data from Clinical Trials and Observational Studies. Wiley, New York. Matthews, D. E., and Farewell, V. (1985 ).Using and Understanding Medical Statistics. S. Karger, New York. Mausner, J. S., and Kramer, S. (1985 ).Epidemiology: An Introductory Text. W.B. Saunders, Philadelphia. Meier, P. (1975a ). Statistics and Medical Experimentation. Biometrics ,31, 511—529. Meier, P. (1975b ). Estimation of a Distribution Function from Incomplete Observa- tions, in Perspectives in Probability and Statistics , edited by J. Gaui. Applied Probability Trust, Sheffield, England. Meisinger, C., Thorand, B., Schneider, A., Stieber, J., Doring, A., and Lowel, H. (2002 ) Sex Differences in Risk Factors for Incident Type 2 Diabetes Mellitus: theMONICA Augsburg Cohort Study. ArchInternMed ,162,8 2—89. Miettinen, O. S. (1979 ). Comments on ‘‘Confidence Intervals for the Odds Ratio in Case-Control Studies: The State of the Art,’’ by J. L. Fleiss. Journal of Chronic Diseases,32,8 0—82. Miller, R. G., Jr. (1966 ).Simultaneous Statistical Inference. McGraw-Hill, New York. Miller, R. G. (1981 ).Survival Analysis. Wiley, New York. Minow, R. A., Benjamin, R. S., Lee, E. T., and Gottlieb, J. A. (1977 ). Adriamycin Cardiomyopathy: Risk Factors. Cancer,39, 1397—1402. Molloy, D. W., Guyatt, G. H., Wilson, D. B., et al. (1991 ). Effect of Tetrahydroaminoac- ridine on Cognition, Function and Behaviour in Alzheimer’s Disease. Canadian Medical Association Journal ,144,2 9—34. Montaner, J. S. G., Lawson, L. M., Levitt, N., et al. (1990 ). Costicorsteroids Prevent Early Deterioration in Patients with Moderately Severe Pneumocystis Carinii 503 Pneumonia and the Acquired Immunodeficiency Syndrome (AIDS ).Annals of Internal Medicine ,113,1 4—20. Moolgavkar, S., Lustbader, E., and Venzon, D. J. (1985 ). Assessing the Adequacy of the Logistic Regression Model for Matched Case-Control Studies. Statistics in Medi- cine,4, 425—435. Moreau, T., O’Quigley, J., and Mesbah, M. (1985 ). A Global Goodness-of-Fit Statistic for the Proportional Hazards Model. Applied Statistics ,34, 212—218. Morrison, D. F. (1967 ).Multivariate Statistical Methods. McGraw-Hill, New York. Morrison,R. S., and Siu, A.L. (2000 ).Survivalin End-StageDementiaFollowingAcute Illness. Journal of the American Medical Association ,284(1),4 7—52. Myers, M., Hankey, B. F., and Mantel, N. (1973 ). A Logistic-Exponential Model for Use with the Response-Time Data Involving Regressor Variables. Biometrics ,29, 257—269. Myers, M. H. (1969 ). A Computing Procedure for a Significance Test of the Difference between Two Survival Curves. Methodological Note 18 in Methodological Notes. End Results Sections, National Cancer Institute, National Institutes of Health,Bethesda, MD. Nadas, A. (1970 ). On Proportional Hazard Functions. Technometrics ,12, 413—416. National Cancer Institute (1970 ).Proceedings of the Symposium on Statistical Aspects of Protocol Design , San Juan, Puerto Rico, December 9 —10. Natrella, M. G. (1963 ).Experimental Statistics . National Bureau of Standards Hand- book 91. U.S. Government Printing Office, Washington, DC, Tables A-25, A-26. Nelson, W. (1972 ). Theory and Applications of Hazard Plotting for Censored Failure Data. Technometrics ,14,94 5—966. Nelson, W. (1982 ).Applied Life Data Analysis. Wiley, New York. Nemenyi, P. (1963 ). Distribution-Free Multiple Comparisons. Ph.D. dissertation, Prin- ceton University. Neter, J., and Wasserman, W. (1974 ).Applied Linear Statistical Models. Richard D. Irwin, Homewood, IL. Nie, N. H., Hull, C. H., Jenkins, J. G., Steinbrenner, K., and Bent, D. H. (1975 ).SPSS: Statistical Package for the Social Sciences. McGraw-Hill, New York. O’Brien, P. C., and Fleming, T. R. (1979 ). A Multiple Testing Procedure for Clinical Trials. Biometrics ,35, 549—556. Osgood, E. W. (1958 ). Methods for Analyzing Survival Data, Illustrated by Hodgkin’s Disease. American Journal of Medicine ,24,4 0—47. Parker, R. L., Dry, T. J., Willius, F. A., and Gage, R. P. (1946 ). Life Expectancy in Angina Pectoris. Journal of the American Medical Association ,131,95—100. Parzan, E. (1974 ). Some Recent Advances in Time Series Modeling. IEEE Transactions on Automatic Control. AC-19, 723—730. Pearson, E. S., and Hartely, N. O. (1958 ).Biometrika Tables for Statisticians , Vol. 1. Cambridge University Press, Cambridge. Pearson, K. (1922, 1957 ).Tables of the Incomplete /afii9772-Function. Cambridge University Press, Cambridge. Pershagen,G. (1986 ). Reviewof Epidemiologyin Relationto PassiveSmoking. Archives of Toxicology ,9(Suppl. ),6 3—73.504  Pershagen, G., Hrubec, Z., and Svensson, C. (1987 ). Passive Smoking and Lung Cancer in Swedish Women. American Journal of Epidemiology ,125,1 7—24. Peto, R., and Lee, P. N. (1973 ). Weibull Distributions for Continuous Carcinogenesis Experiments. Biometrics ,29, 457—470. Peto, R., and Peto, J. (1972 ). Asymptotically Efficient Rank Invariant Procedures. Journal of the Royal Statistical Society, Series A ,135, 185—207. Peto, R., Lee, P. N., and Paige, W. S. (1972 ). Statistical Analysis of the Bioassay of Continuous Carcinogens. British Journal of Cancer ,26, 258—261. Peto, R., Pike, M. C., Armitage, P., Breslow, N. E., Cox, D. R., Howard, S. V., Mantel, N., McPherson, K., Peto, J., and Smith, P. G. (1976, 1977 ). Design and Analysis of Randomized Clinical Trials Requiring Prolonged Observation of Each Patient.British Journal of Cancer , Part I, 34, 585—612, 1976; Part II, 35,1—39, 1977. Pierce,M., Borges, W. H., Heyn,R., Wolfe, J., and Gilbert,E. S. (1969 ). Epidemiological Factors and Survival Experience in 1770 Children with Acute Leukemia. Cancer, 23, 1296—1304. Pike, M. C. (1966 ). A Method of Analysis of a Certain Class of Experiments in Carcinogenesis. Biometrics ,22, 142—161. Piper, J. M., Matanoski, G. M., and Tonascia, J. (1986 ). Bladder Cancer in Young Women. American Journal of Epidemiology ,123, 1033—1042. Pregibon, D. (1984 ). Data Analytic Methods for Matched Case-Control Studies. Biometrics ,40, 639—651. Prentice, R. L. (1973 ). Exponential Survivals with Censoring and Explanatory Vari- ables. Biometrika ,60, 279—288. Prentice, R. L. (1974 ). A Log-Gamma Model and Its Maximum Likelihood Estimation. Biometrica. 61539—544. Prentice, R. L. (1976 ). Use of the Logistic Model in Retrospective Studies. Biometrics , 32,5 99—606. Prentice, R. L., and Gloeckler, L. A. (1978 ). Regression Analysis of Grouped Survival Data with Application to Breast Cancer Data. Biometrics ,34,5 7—67. Prentice, R. L., and Kalbfleisch, J. D. (1979 ). Hazard Rate Models with Covariates. Biometrics ,35,2 5—39. Prentice, R. L., and Marek, P. (1979 ). A Quantitative Discrepancy between Censored Data Rank Tests. Biometrics ,35, 861—867. Prentice, R. L., and Pyke, R. (1979 ). Logistic Disease Incidence Models and Case- control Studies.Biometrica ,73, 403—411. Prentice, R. L., Williams, B. J., and Peterson, A. V. (1981 ). On the Regression Analysis of Multivariate Failure Time Data. Biometrica ,68, 373—379. Press, S. J. (1972 ).Applied Multivariate Analysis. Holt, Rinehart & Winston,New York. Press, S. J., and Wilson, S. (1978 ). Choosing between Logistic Regression and Dis- criminant Analysis. Journal of the American Statistical Association ,73,6 99—705. Ralston, A., and Wilf, H. (1967 ).Mathematical Methods for Digital Computers. Wiley, New York. Rao, C. R. (1952 ).Advanced Statistical Methods in Biometric Research. Wiley, New York. 505 Rao, C. R. (1973 ).Linear Statistical Inference and Its Application , 2nd ed. Wiley, New York. Riffenburgh, R. H., and Johnstone, P. A. (2001 ). Survival Patterns of Cancer Patients. Cancer,91(12), 2469—2475. Rissanen, J. (1986 ). A Predictive Least-SquaresPrinciple. IMA Journal of Mathematical Control of Information ,3, 211—222. Rowe-Jones, D. C., Peel, A. L. G., Kingston, R. D., Shaw, J. F. L., Teasdale, C., and Cole, D. S. (1990 ). Single Dose Cefotaxime Plus Metronidazole versus Three Dose Cefuroxime Plus Metronidazole as Prophylaxis against Wound Infection in Colorectal Surgery: Multicentre Prospective Randomised Study. British Medical Journal,300,1 8—22. Sacher, G. A. (1956 ). On the Statistical Nature of Mortality, with Special Reference to Chronic Radiation Mortality. Radiology ,67, 250—257. Sarhan, A. E., and Greenberg, B. G. (1956 ). Estimation of Location and Scale Parameters by Order Statistics from Singly and Doubly Censored Samples, PartI, The Normal Distribution up to Samples of Size 10. Annals of Mathematical Statistics ,27, 427—451. Sarhan, A. E., and Greenberg, B. G. (1957 ). Estimation of Location and Scale Parameters by Order Statistics from Singly and Doubly Censored Samples, Part III.Technical Report 4-OOR, Project 1597. U.S. Army Research Office. Sarhan, A. E., and Greenberg, B. G. (1958 ). Estimation of Location and Scale Parameters by Order Statistics from Singly and Doubly Censored Samples, PartII.Annals of Mathematical Statistics ,29,7 9—105. Sarhan,A. E., and Greenberg,B. G. (1962 ).Contribution to Order Statistics. Wiley, New York. Sacks,H., Chalmers,T. C., and Smith, H. (1982 ). Randomizedversus HistoricalControl for Clinical Trials. American Journal of Medicine ,72, 233—240. SAS Institute. (2000 ).SAS/STAT User ’s Guide, Version 8.1. SAS Institute, Cary, NC. Sasieni, P. D. (1996 ). Proportional Excess Hazards. Biometrika ,83(1), 127—141. Savage, I. R. (1956 ). Contributions to the Theory of Rank Order Statistics: The Two Sample Case. Annals of Mathematical Statistics ,27,5 90—615. Saw, J. G. (1959 ). Estimation of the Normal Population Parameters Given a Singly Censored Sample. Biometrika ,46, 150—159. Schade,D.S., Mitchell,W. J., and Griego,G. (1987 ).AdditionofSulfonylureato Insulin Treatment in Poorly Controlled Type II Diabetes. Journal of the American Medical Association ,257, 2441—2445. Schafer, J. L. (1999 ). Multiple Imputation: a Primer. StatMethods ,8,3—15. Schlesselman, J. J. (1982 ).Case-Control Studies. Oxford University Press, New York. Schoenfeld, D. (1982 ). Partial Residuals for Proportional Hazards Regression Model. Biometrica. 69, 239—241. Schwarz, G. (1978 ). Estimating the Dimension of a Model. Annals of Statistics 6, 461—222. Seaman, S. R., and Bird, S. M. (2001 ). Proportional Hazards Model for Interval- Censored Failure Times and Time-Dependent Covariates: Application to Hazard of HIV Infection of Injecting Drug Users in Prison. Statistics in Medicine ,20(12), 1855—1870.506  Segal, M. R., and Bloch, D. A. (1989 ). A Comparison of Estimated Proportional Hazards Models and Regression Trees. Statistics in Medicine ,8, 539—550. Sellke, T., and Siegmund, D. (1983 ). Sequential Analysis of the Proportional Hazards Model. Biometrika ,70, 315—326. Shapiro, S. S., and Wilk, M. B. (1965a ). An Analysis of Variance Test for Normality (Complete Samples ).Biometrika ,52, 591. Shapiro, S. S., and Wilk, M. B. (1965b ). Testing for Distributional Assumptions: Exponential and Uniform Distributions. Unpublished manuscript. Shibata, R. (1980 ). Asymptotically Efficient Selection of the Order of the Model for Estimating Parameters of a Linear Process. Annals of Statistics ,8, 147—165. Shipley, W. U., Thames, H. D., Sandler, H. M., Hanks, G. E., Zietman, Perez, C. A., Kuban, D. A., Hancock, S. L., and Smith, C. D. (1999 ). Radiation Therapy for ClinicallyLocalizedProstate Cancer. Journal of the American Medical Association , 281(17), 1598—1604. Shryock, H. S., Sigel, J. S., and Associates (1971 ).The Methods and Materials of Demography ,Vols. I and II. U.S. Departmentof Commerce, Bureau of the Census, U.S. Government Printing Office, Washington, DC. Sichieri, R., Everhart, J. E., and Roth, H. P. (1990 ). Low Incidence of Hospitalization with Gallbladder Disease among Blacks in the United States. American Journal of Epidemiology ,131, 826—835. Siegmund, K. D., Todorov, A. A., and Province, M. A. (1999 ). A Frailty Approach for Modelling Diseases with Variable Age of Onset in Families: The NHLBI Family Heart Study. Statistics in Medicine, 18(12), 1517—1528 Sillitto, G. P. (1949 ). Note on Approximations to the Power Function of the ‘‘2 /p592 Comparative Trial.’’ Biometrika ,36, 347—352. Sirott, M. N., Bajorin, D. F., Wong, G. Y., Tao, Y., Chapman, P. B., Templeton, M. A., and Houghton, A. N. (1993 ). Prognostic Factors in Patients with Metastatic Malignant Melanoma: A Multivariate Analysis. Cancer,72(10), 3091—3098. Slud, E. V., and Wei, L. J. (1982 ). Two-Sample Repeated Significance Tests Based on the Modified Wilcoxon Statistic. Journal of the American Statistical Society ,77, 862—868. Snedecor,G. W., and Cochran,W.G. (1967 ).Statistical Methods. Iowa State University Press, Ames, IA. SPSS (2000 ).SPSS-S User ’s Guide, Version 10.1. SPSS, Chicago. Stacy, E. W. (1962 ). A Generalizationof the Gamma Distribution. Annals of Mathemat- ical Statistics ,33, 1187—1192. Stacy, E. W., and Mihram, G. A. (1965 ). Parameter Estimation for a Generalized Gamma Distribution. Technometrics ,7, 349—358. Statistics and Epidemiology Research Corporation (SERC )(1988 ).EGRET Statistical Software. SERC, Seattle, WA. Steering Committee on the Physicians Health Study Research Group (1989 ). Final Report on the Aspirin Component of the Ongoing Physicians’ Health Study. New England Journal of Medicine ,321, 129—135. Tai, B. C., Peregoudov, A., and Machin, D. (2001 ). A Competing Risk Approach to the Analysis of Trials of Alternative Intra-uterine Devices (IUDs )for Fertility Regu- lation. Statistics in Medicine ,20(23), 3589—3600. 507 Tarone, R. E. (1982 ). The Use of Historical Control Information in Testing a Trend in Poisson Means. Biometrics ,38(2), 457—462. Tarone, R. E., and Ware, J. (1977 ). On Distribution-Free Tests for Equality of Survival Distribution. Biometrics ,64, 156—160. Teitelman, A. M., Welch, L. S., Hellenbrand, K. G., and Bracken, M. B. (1990 ). Effect of Maternal Work Activity on Preterm Birth and Low Birth Weight. American Journal of Epidemiology ,131, 104—113. Therneau, T. M., Grambsch, P. M., and Fleming, T. R. (1990 ). Martingale-Based Residuals and Survival Models. Biometrica ,77, 147—160. Thoman, D. R., and Bain, L. J. (1969 ). Two Sample Tests in the Weibull Distribution. Technometrics ,11, 805—815. Thoman, D. R., Bain, L. J., and Antle, C. E. (1969 ). Inferences on the Parameters of the Weibull Distribution. Technometrics ,11, 445—460. Thoman, D. R., Bain, L. J., and Antle, C. E. (1970 ). Maximum Likelihood Estimation, Exact Confidence Intervals for Reliability and Tolerance Limits in the Weibull Distribution. Technometrics ,12, 363—373. Truett, J., Cornfield, J., and Kannel, W. B. (1967 ). A Multivariate Analysis of the Risk of Coronary Heart Disease in Framingham. Journal of Chronic Diseases ,20, 511—524. Tsiatis, A. A. (1980 ). A Note of a Goodness-of-Fit Test for the Logistic Regression Model. Biometrika ,67, 250—251. Tsiatis, A. A. (1981 ). A Large Sample Study of Cox’s Regression Model. Annals of Statistics ,9,93—108. Tsiatis, A. A. (1982 ). Repeated Significance Testing for a General Class of Statistics Used in Censored Survival Analysis. Journal of the American Statistical Associ- ation,77, 855—861. Tsumagari, K., Yamamoto, H., Suganuma, N., Kato, M., Ikeda, S., Imai, K., Kira, S., and Taketa, K. (2000 ). Epidemiological Studies of Coincidental Outbreaks of Enterohemorrhagic Escherichia Coli O157:H7 Infection and Infectious Gastroen-teritis in Niimi City, ActaMedicaOkayama ,54, 265—273. Upton, G. J. G. (1978 ).The Analysis of Cross-Tablated Data . Wiley New York. Vaida, F., and Xu, R. (2000 ). Proportional Hazards Model with Random Effects. StatisticsinMedicine ,19(22), 339—3324. Vasan, R. S., Larson, M. G., Leip, E. P., Evans, J. C., O’Donnell, C. J., Kannel, W. B., and Levy, D. (2001 ). Impact of High-Normal Blood Pressure on the Risk of Cardiovascular Disease. New England Journal of Medicine .345(18), 1337—1340. Vaupel, J. W., Manton, K. G., and Stallard, E. (1979 ). The Impact of Heterogenity in Individual Frailty on the Dynamics of Mortality. Demography ,16, 439—454. Vasan, R. S., Larson, M. G., Levy, D., Evans, J. C., and Benjamin, E. J. (1997 ). Distribution and Categorization of Echocardiographic Measurements in Relationto Reference Limits: the Framingham Heart Study: Formulation of a Height- and Sex-specific Classification and Its Prospective Validation. Circulation ,96, 1863—1873. Vega, G. L., and Grundy, S. M. (1989 ). Comparison of Lovastatin and Gemfibrozil in Normolipidemic Patients with Hypoalphalipoproteinemia. Journal of the Ameri- can Medical Association ,262, 3148—3153.508  Wald, A. (1947 ).Sequential Analysis. Wiley, New York. Wang, Wenyu (1984 ). The Bayesian Estimater of the Orders of AR (k)and ARMA (p,q) Models of Time Series. Acta Mathematicae Applicatae Sinica ,7(2), 185—195. Wang, Wenyu (1989 ). Statistical Inference on Aggregated Markov Processes. Ph.D. dissertation. Department of Mathematics, University of Maryland. Waters,M.A.,Selvin,S.,and Rappaport,S. M. (1991 ).AMeasureofGoodness-of-fitfor the Lognormal Model Applied to Occupational Exposures, AmericanIndustrial HygieneAssociationJournal ,52,4 93—502. Watson, G. S., and Wells, W. T. (1961 ). On the Possibility of Improving the Mean Useful Life of Items by Eliminating Those with Short Lives. Technometrics ,3, 281—298. Wei, L. J. (1984 ). Testing Goodness of Fit for Proportional Hazards Model with Censored Observations. Journal of the American Statistical Association ,79, 649— 652. Wei, L. J. (1992 ). On Predictive Least Squares Principle. Annals of Statistics ,20,1—42. Wei, L. J., Lin, D. Y., and Weissfeld, L. (1989 ). Regression Analysis of Multivariate Incomplete Failure Time Data by Modeling Marginal Distribution. Journal of the American Statistical Association ,84, 1065—1073. Weibull, W. (1939 ). A Statistical Theory of the Strength of Materials. Ingenioers vetenskaps akakemien Handlingar ,151,2 93—297. Weibull, W. (1951 ). A Statistical Distribution of Wide Applicability. Journal of Applied Mathematics, 18,2 93—297. Weiss,H. (1963 ).A Survey of Some MathematicalMethods in the Theory of Reliability, inStatistical Theory of Reliability , edited by M. Zelen. University of Wisconsin Press, Madison, WI. Well, M. D., Lamborn, K., Edwards, M. S. B., and Wara, W. M. (1998 ). Influence of a Child’s Sex on Medulloblastoma. Journal of the American Medical Association , 279(18). Whayne, T. F., Alaupovic, P., Curry, M. D., Lee, E. T., Anderson, P. S., and Schechter, E.(1981 ). Plasma Apolipoprotein B and VLDL-, LDL-, and HDL-Cholesterol as Risk Factors in the Development of Coronary Heart Disease in Male PatientsExamined by Angiography. Atherosclerosis ,39, 411—424. Wienke, A., Holm, N. V., Skytthe, A., and Yashin, A. L. (2001 ). The Heritability of Mortality Due to Heart Diseases: A Correlated Frailty Model Applied to Danish Twins. Twin Research ,4(4), 266—274. Wilcoxon,F. (1945 ).IndividualComparisonby RankingMethods. Biometrics ,1,80—83. Wilk, M. B., Gnanadesikan, R., and Huyett, M. J. (1962a ). Estimation of Parameters of the Gamma Distribution Using Order Statistics. Biometrika ,49, 525—545. Wilk, M. B., Gnanadesikan, R., and Huyett, M. J. (1962b ). Probability Plots for the Gamma Distribution. Technometrics ,4,1—20. Wilkinson, L. (1987 ).SYSTAT: The System for Statistics. Systat, Inc., Evanston, IL. Wilks, S. S. (1948 ). Order Statistics. Bulletin of the American Mathematical Society ,54, 6—50. Wilks, S. S. (1950 ). Mathematical Statistics. Princeton University Press, Princeton, NJ. 509 Williams, C. A., Jr. (1950 ). On the Choice of the Number and Width of Classes for the Chi-Square Test of Goodness of Fit. Journal of the American Statistical Associ- ation,45,7 7—86. Williams, J. E., Nieto, F. J., Sanford, C. P., and Tyroler, H. A. (2002 ). The Association between Trait Anger and Incident Stroke Risk: The Atherosclerosis Risk in Communities (ARIC )Study. Stroke,33(1),1 3—20. Winkleby, M. A., Ragland, D. R., and Syme, L. (1988 ). Self-Reported Stressors and Hypertension:Evidence of an Inverse Association. American Journal of Epidemiol- ogy,127, 124—134. Winter, F. D., Snell, P. G., and Stray-Gundersen, J. (1989 ). Effects of 100%Oxygen on Performance of Professional Soccer Players. Journal of the American Medical Association ,262, 227—229. Woolf, B. (1955 ). On Estimating the Relation between Blood Group and Disease. Annals of Human Genetics ,19, 251—253. Wurpel, J. N., Dundore, R. L., Barbella, Y. R., Balaban, C. D., Keil, L. C., and Severs, W. B. (1986 ). Barrel Rotation Evoked by Intracerebroventricular Vasopressin Injections in Conscious Rats. I. Description and General Pharmacology. Brain Research ,365,2 1—29. Xue, X. (2001 ). Analysis of Childhood Brain Tumour Data in New York City Using Frailty Models. Statistics in Medicine ,20(22), 3459—3473. Yakovlev, A. Y., Tsodikov, A. D., Boucher, K., and Kerber, R. (1999 ). The Shape of the Hazard Function in Breast Carcinoma: Curability of the Disease Revisited.Cancer,85, 1789—1798. Yan, Y., Moore, R. D., and Hoover, D. R. (2000 ). Competing Risk Adjustment Reduces Overestimation of Opportunistic Infection Rates in AIDS. Journal of Clinical Epidemiology. 53(8), 817—822. Yashin,A. I.,and Iachine,I. A. (1997 ).How FrailtyModelsCan BeUsed forEvaluating Longevity Limits: Taking Advantage of an Interdisciplinary Approach. Demogra- phy,34(1),3 1—48. Young, E. M., and Fors, S. W. (2001 ). Factors Related to the Eating Habits of Students in Grades 9—12.JournalofSchoolHealth ,71, 483—488. Zelen, M. (1966 ). Applications of Exponential Models to Problems in Cancer Research. Journal of the Royal Statistical Society, Series A ,129, 368—398. Zhang, M. J., and Klein, J. P. (2001 ). Confidence Bands for the Difference of Two Survival Curves under Proportional Hazards Model. LifetimeDataAnalysis ,7, 243—254. Zippin, C., and Armitage, P. (1966 ). Use of Concomitant Variables and Incomplete Survival Information in the Estimation of an Exponential Survival Parameter.Biometrics ,22, 665—672.510  Index Accelerated failure time (AFT )model, 259 Age-specific failure rate, 11AIC, 230, 241, 288, 289 Anderson —Gill model, 368 Annual survival ratio, 94 BIC, 230, 241, 288, 289 BMDP, 7, 94, 115, 173, 180, 188, 196, 235, 269, 273, 277, 283, 306, 319, 324, 347, 351, 356, 367, 368, 389, 394, 398, 403,405, 413, 419, 424 Case-control study, 399 Censored observations, 2 progressively censored data, 4 singly censored data, 4 Censoring interval, 4, 260 left, 4, 260 random, 4right, 4, 260 type I, 2 type II, 2type III, 3 Chi-square test, 379, 381 Competing risk, 352Conditional mortality rate, 11 Conditional probability, 399, 401 Corrected survival rate, 94 Cox’s F-test, 116, 249 Cox —Mantel test, 109 Cox —Snell residual, 199, 215, 290, 331 Cross-product ratio, 382 Cumulative hazard function, 13 Cumulative survival rate, 9Density function: definition of, 10types of: exponential, 135 extended generalized gamma, 153gamma, 150 generalized gamma, 152 log-logistic, 154lognormal, 145 Weibull, 139 Deviance residual, 331Dichotomous outcomes, 377 Exponential distribution, 134, 263 unit exponential distribution, 135two-parameter, 136 goodness-of-fit test, 226, 233 testfor equalityof two distributions,246,249 Five-year survival rate, 94 Force of mortality, 11 Frailty model, 375 Gamma and generalized gamma distribution, 148, 277 goodness-of-fit test, 227, 234, 235test for equality of two distributions, 252 Gap time, 364 Gehan’s generalized Wilcoxon test, 107 Gompertz distribution, 157, 242 Guarantee time, 136, 173, 175 Hazard function: definition of, 11 exponential, 135 511 Hazard function (Continued ) gamma, 152 log-logistic, 155lognormal, 145 Weibull, 140 Hazard plotting, 29, 209 exponential, 210 log-logistic, 215 lognormal, 213 Weibull, 212 Hollander —Proschan’s test, 236 Hosmer —Lemeshow test of goodness-of-fit, 388 Incomplete gamma function, 151 Instantaneous failure rate, 11 Kaplan —Meier method, 20, 68, 216, 237 Kruskal —Wallis test, 125 multiple comparison, 128 K-sample test for censored data, 130 Life tables: abridged, 86 clinical, 87 cohort, 77current, 77 population, 77 Likelihood ratio test, 243, 246Link function, 410 logit, 410 probit, 410complementary log-log, 411 Linear exponential distribution, 155 Log-logistic distribution, 154, 280 goodness-of-fit test, 235 Log odds, 386 Logistic regression, 45, 385 conditional, 398 dichotomous outcomes, 377 polychotomous outcomes, 377, 413 nominal, 414 ordinal, 419 Logistic transform, 386Log-likelihood ratio statistic, 223, 224, 388 Lognormal distribution, 143, 274 three-parameter, 146goodness-of-fit test, 227, 234 Logrank test, 111 Mantel —Haenszel method, 28, 121 Maximum likelihood estimate: exponential, 166 Gompertz, 196interval-censored data, 165, 260 left-censored data, 165, 260 log-logistic, 195lognormal, 180 right-censored data, 162, 260 standard and generalized gamma, 188two-parameter exponential, 174 Weibull, 178 Martingale residual, 331 Matched design, 401 1:R, 401 n/p16:n/p15,40 3 Median remaining lifetime, 91 Model selection method, 230, 233, 286, 288, 289, 310 Normit function, 410Odds ratio, 379, 382 Partial likelihood function, 301, 304, 340, 348, 354, 357, 364, 369, 371 Peto and Peto’s generalized Wilcoxon test, 116Polychotomous outcomes, 377, 413 nominal, 414 ordinal, 419 Printice, Williams, and Peterson (PWP )model, 357, 363 Probability density function, 10Probability plotting, 29, 200 exponential, 204 log-logistic, 208lognormal, 206 normal, 203 Weibull, 205 Probit function, 410, 420 Product-limit estimate, 65 estimate of mean survival time, 74variance, 70 variance of estimated mean survival time, 75 Prognostic factors, 32, 256, 339, 377Prognostic homogeneity, 21 Proportional hazard model, 264, 270, 298 assessment of, 326 Proportional odds, 280 Rayleigh distribution, 158 Recurrent events, 356 Related observations, 374Relative mortality, 32 Relative survival rate, 94 Retrospective study, 398512  SAS, 7, 94, 115, 172, 173, 80, 188, 194, 196, 235, 268, 273, 276, 279, 283, 289, 291, 306, 310, 317, 322, 333, 346, 350, 355,366, 389, 392, 397, 403, 408, 410, 412, 418, 424 Schoenfeld residual, 331 weighted, 332 Score statistic, 225 SPSS, 7, 94, 115, 318, 323, 336, 347, 351, 356, 367, 389, 392, 397, 405, 413, 419, 424 Standardized rate and ratio, 97 direct method, 98indirect method, 99 SIR, 97 SMR, 97 Standardized mortality ratio, 97 Stritification, 328, 348 Survival curve, 9Survivorship function: definition of, 8 estimation of, 319types: exponential, 135 gamma, 152log-logistic, 154 lognormal, 145 two-parameter-exponential, 136Weibull, 140 Test of goodness of fit, 221, 222, 226, 227, 233, 234, 235, 330 Tied survival times, 302 Time-dependent covariates, 326, 339 Unconditional failure rate, 11 Wald statistic, 223, 224, 262, 388 Wei, Lin, and Weissfeld (WLW )model, 370 Weibull distribution, 138, 269 three-parameter, 141 goodness-of-fit test, 226, 234 test for equality of two distributions, 251 513 WILEY SERIES IN PROBABILITY AND STATISTICS Established by WALTER A. SHEWHART and SAMUEL S. WILKSEditors: David J. Balding, Peter Bloomfield, Noel A. C. Cressie, Nicholas I. Fisher, Iain M. Johnstone, J. B. Kadane, Louise M. Ryan, David W. Scott, Adrian F. M. Smith, Jozef L. TeugelsEditors Emeriti: Vic Barnett, J. Stuart Hunter, David G. Kendall A complete list of the titles in this series appears at the end of this volume.p&s-cp.qxd 3/25/03 9:47 AM Page 1 WILEY SERIES IN PROBABILITY AND STATISTICS ESTABLISHED BY WALTER A. S HEWHART AND SAMUEL S. W ILKS Editors: David J. Balding, Peter Bloomfield, Noel A. C. Cressie, Nicholas I. Fisher, Iain M. Johnstone, J. B. Kadane, Louise M. Ryan, David W. Scott, Adrian F. M. Smith, Jozef L. Teugels Editors Emeriti: Vic Barnett, J. Stuart Hunter, David G. Kendall The Wiley Series in Probability and Statistics is well established and authoritative. It covers many topics of current research interest in both pure and applied statistics and probabilitytheory. Written by leading statisticians and institutions, the titles span both state-of-the-artdevelopments in the field and classical methods. Reflecting the wide range of current research in statistics, the series encompasses applied, methodological and theoretical statistics, ranging from applications and new techniquesmade possible by advances in computerized practice to rigorous treatment of theoreticalapproaches. This series provides essential and invaluable reading for all statisticians, whether in aca- demia, industry, government, or research. ABRAHAM and LEDOLTER ·Statistical Methods for Forecasting AGRESTI ·Analysis of Ordinal Categorical Data AGRESTI ·An Introduction to Categorical Data Analysis AGRESTI ·Categorical Data Analysis, Second Edition ANDE L ·Mathematics of Chance ANDERSON ·An Introduction to Multivariate Statistical Analysis, Second Edition *ANDERSON ·The Statistical Analysis of Time Series ANDERSON, AUQUIER, HAUCK, OAKES, VANDAELE, and WEISBERG · Statistical Methods for Comparative Studies ANDERSON and LOYNES ·The Teaching of Practical Statistics ARMITAGE and DAVID (editors) ·Advances in Biometry ARNOLD, BALAKRISHNAN, and NAGARAJA ·Records *ARTHANARI and DODGE ·Mathematical Programming in Statistics *BAILEY ·The Elements of Stochastic Processes with Applications to the Natural Sciences BALAKRISHNAN and KOUTRAS ·Runs and Scans with Applications BARNETT ·Comparative Statistical Inference, Third Edition BARNETT and LEWIS ·Outliers in Statistical Data, Third Edition BARTOSZYNSKI and NIEWIADOMSKA-BUGAJ ·Probability and Statistical Inference BASILEVSKY ·Statistical Factor Analysis and Related Methods: Theory and Applications BASU and RIGDON ·Statistical Methods for the Reliability of Repairable Systems BATES and WATTS ·Nonlinear Regression Analysis and Its Applications BECHHOFER, SANTNER, and GOLDSMAN ·Design and Analysis of Experiments for Statistical Selection, Screening, and Multiple Comparisons BELSLEY ·Conditioning Diagnostics: Collinearity and Weak Data in Regression BELSLEY, KUH, and WELSCH ·Regression Diagnostics: Identifying Influential Data and Sources of Collinearity BENDAT and PIERSOL ·Random Data: Analysis and Measurement Procedures, Third Edition *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 2 BERRY, CHALONER, and GEWEKE ·Bayesian Analysis in Statistics and Econometrics: Essays in Honor of Arnold Zellner BERNARDO and SMITH ·Bayesian Theory BHAT and MILLER ·Elements of Applied Stochastic Processes, Third Edition BHATTACHARYA and JOHNSON ·Statistical Concepts and Methods BHATTACHARYA and WAYMIRE ·Stochastic Processes with Applications BILLINGSLEY ·Convergence of Probability Measures, Second Edition BILLINGSLEY ·Probability and Measure, Third Edition BIRKES and DODGE ·Alternative Methods of Regression BLISCHKE AND MURTHY (editors) ·Case Studies in Reliability and Maintenance BLISCHKE AND MURTHY ·Reliability: Modeling, Prediction, and Optimization BLOOMFIELD ·Fourier Analysis of Time Series: An Introduction, Second Edition BOLLEN ·Structural Equations with Latent Variables BOROVKOV ·Ergodicity and Stability of Stochastic Processes BOULEAU ·Numerical Methods for Stochastic Processes BOX ·Bayesian Inference in Statistical Analysis BOX ·R. A. Fisher, the Life of a Scientist BOX and DRAPER ·Empirical Model-Building and Response Surfaces *BOX and DRAPER ·Evolutionary Operation: A Statistical Method for Process Improvement BOX, HUNTER, and HUNTER ·Statistics for Experimenters: An Introduction to Design, Data Analysis, and Model Building BOX and LUCEÑO ·Statistical Control by Monitoring and Feedback Adjustment BRANDIMARTE ·Numerical Methods in Finance: A MATLAB-Based Introduction BROWN and HOLLANDER ·Statistics: A Biomedical Introduction BRUNNER, DOMHOF, and LANGER ·Nonparametric Analysis of Longitudinal Data in Factorial Experiments BUCKLEW ·Large Deviation Techniques in Decision, Simulation, and Estimation CAIROLI and DALANG ·Sequential Stochastic Optimization CHAN ·Time Series: Applications to Finance CHATTERJEE and HADI ·Sensitivity Analysis in Linear Regression CHATTERJEE and PRICE ·Regression Analysis by Example, Third Edition CHERNICK ·Bootstrap Methods: A Practitioner’s Guide CHERNICK and FRIIS ·Introductory Biostatistics for the Health Sciences CHILÈS and DELFINER ·Geostatistics: Modeling Spatial Uncertainty CHOW and LIU ·Design and Analysis of Clinical Trials: Concepts and Methodologies CLARKE and DISNEY ·Probability and Random Processes: A First Course with Applications, Second Edition *COCHRAN and COX ·Experimental Designs, Second Edition CONGDON ·Bayesian Statistical Modelling CONOVER ·Practical Nonparametric Statistics, Second Edition COOK ·Regression Graphics COOK and WEISBERG ·Applied Regression Including Computing and Graphics COOK and WEISBERG ·An Introduction to Regression Graphics CORNELL ·Experiments with Mixtures, Designs, Models, and the Analysis of Mixture Data, Third Edition COVER and THOMAS ·Elements of Information Theory COX ·A Handbook of Introductory Statistical Methods *COX ·Planning of Experiments CRESSIE ·Statistics for Spatial Data, Revised Edition CSÖRGO´´and HORVÁTH ·Limit Theorems in Change Point Analysis DANIEL ·Applications of Statistics to Industrial Experimentation DANIEL ·Biostatistics: A Foundation for Analysis in the Health Sciences, Sixth Edition *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 3 *DANIEL ·Fitting Equations to Data: Computer Analysis of Multifactor Data, Second Edition DASU and JOHNSON ·Exploratory Data Mining and Data Cleaning DAVID ·Order Statistics, Second Edition *DEGROOT, FIENBERG, and KADANE ·Statistics and the Law DEL CASTILLO ·Statistical Process Adjustment for Quality Control DETTE and STUDDEN ·The Theory of Canonical Moments with Applications in Statistics, Probability, and Analysis DEY and MUKERJEE ·Fractional Factorial Plans DILLON and GOLDSTEIN ·Multivariate Analysis: Methods and Applications DODGE ·Alternative Methods of Regression *DODGE and ROMIG ·Sampling Inspection Tables, Second Edition *DOOB ·Stochastic Processes DOWDY and WEARDEN ·Statistics for Research, Second Edition DRAPER and SMITH ·Applied Regression Analysis, Third Edition DRYDEN and MARDIA ·Statistical Shape Analysis DUDEWICZ and MISHRA ·Modern Mathematical Statistics DUNN and CLARK ·Applied Statistics: Analysis of Variance and Regression, Second Edition DUNN and CLARK ·Basic Statistics: A Primer for the Biomedical Sciences, Third Edition DUPUIS and ELLIS ·A Weak Convergence Approach to the Theory of Large Deviations *ELANDT-JOHNSON and JOHNSON ·Survival Models and Data Analysis ENDERS ·Applied Econometric Time Series ETHIER and KURTZ ·Markov Processes: Characterization and Convergence EVANS, HASTINGS, and PEACOCK ·Statistical Distributions, Third Edition FELLER ·An Introduction to Probability Theory and Its Applications, Volume I, Third Edition, Revised; Volume II, Second Edition FISHER and VAN BELLE ·Biostatistics: A Methodology for the Health Sciences *FLEISS ·The Design and Analysis of Clinical Experiments FLEISS ·Statistical Methods for Rates and Proportions, Second Edition FLEMING and HARRINGTON ·Counting Processes and Survival Analysis FULLER ·Introduction to Statistical Time Series, Second Edition FULLER ·Measurement Error Models GALLANT ·Nonlinear Statistical Models GHOSH, MUKHOPADHYAY, and SEN ·Sequential Estimation GIFI ·Nonlinear Multivariate Analysis GLASSERMAN and YAO ·Monotone Structure in Discrete-Event Systems GNANADESIKAN ·Methods for Statistical Data Analysis of Multivariate Observations, Second Edition GOLDSTEIN and LEWIS ·Assessment: Problems, Development, and Statistical Issues GREENWOOD and NIKULIN ·A Guide to Chi-Squared Testing GROSS and HARRIS ·Fundamentals of Queueing Theory, Third Edition *HAHN and SHAPIRO ·Statistical Models in Engineering HAHN and MEEKER ·Statistical Intervals: A Guide for Practitioners HALD ·A History of Probability and Statistics and their Applications Before 1750 HALD ·A History of Mathematical Statistics from 1750 to 1930 HAMPEL ·Robust Statistics: The Approach Based on Influence Functions HANNAN and DEISTLER ·The Statistical Theory of Linear Systems HEIBERGER ·Computation for the Analysis of Designed Experiments HEDAYAT and SINHA ·Design and Inference in Finite Population Sampling HELLER ·MACSYMA for Statisticians HINKELMAN and KEMPTHORNE: ·Design and Analysis of Experiments, Volume 1: Introduction to Experimental Design *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 4 HOAGLIN, MOSTELLER, and TUKEY ·Exploratory Approach to Analysis of Variance HOAGLIN, MOSTELLER, and TUKEY ·Exploring Data Tables, Trends and Shapes *HOAGLIN, MOSTELLER, and TUKEY ·Understanding Robust and Exploratory Data Analysis HOCHBERG and TAMHANE ·Multiple Comparison Procedures HOCKING ·Methods and Applications of Linear Models: Regression and the Analysis of Variance, Second Edition HOEL ·Introduction to Mathematical Statistics, Fifth Edition HOGG and KLUGMAN ·Loss Distributions HOLLANDER and WOLFE ·Nonparametric Statistical Methods, Second Edition HOSMER and LEMESHOW ·Applied Logistic Regression, Second Edition HOSMER and LEMESHOW ·Applied Survival Analysis: Regression Modeling of Time to Event Data HØYLAND and RAUSAND ·System Reliability Theory: Models and Statistical Methods HUBER ·Robust Statistics HUBERTY ·Applied Discriminant Analysis HUNT and KENNEDY ·Financial Derivatives in Theory and Practice HUSKOVA, BERAN, and DUPAC ·Collected Works of Jaroslav Hajek— with Commentary IMAN and CONOVER ·A Modern Approach to Statistics JACKSON ·A User’s Guide to Principle Components JOHN ·Statistical Methods in Engineering and Quality Assurance JOHNSON ·Multivariate Statistical Simulation JOHNSON and BALAKRISHNAN ·Advances in the Theory and Practice of Statistics: A Volume in Honor of Samuel Kotz JUDGE, GRIFFITHS, HILL, LÜTKEPOHL, and LEE ·The Theory and Practice of Econometrics, Second Edition JOHNSON and KOTZ ·Distributions in Statistics JOHNSON and KOTZ (editors) ·Leading Personalities in Statistical Sciences: From the Seventeenth Century to the Present JOHNSON, KOTZ, and BALAKRISHNAN ·Continuous Univariate Distributions, Volume 1, Second Edition JOHNSON, KOTZ, and BALAKRISHNAN ·Continuous Univariate Distributions, Volume 2, Second Edition JOHNSON, KOTZ, and BALAKRISHNAN ·Discrete Multivariate Distributions JOHNSON, KOTZ, and KEMP ·Univariate Discrete Distributions, Second Edition JUREC KOVÁ and SEN ·Robust Statistical Procedures: Aymptotics and Interrelations JUREK and MASON ·Operator-Limit Distributions in Probability Theory KADANE ·Bayesian Methods and Ethics in a Clinical Trial Design KADANE AND SCHUM ·A Probabilistic Analysis of the Sacco and Vanzetti Evidence KALBFLEISCH and PRENTICE ·The Statistical Analysis of Failure Time Data, Second Edition KASS and VOS ·Geometrical Foundations of Asymptotic Inference KAUFMAN and ROUSSEEUW ·Finding Groups in Data: An Introduction to Cluster Analysis KEDEM and FOKIANOS ·Regression Models for Time Series Analysis KENDALL, BARDEN, CARNE, and LE ·Shape and Shape Theory KHURI ·Advanced Calculus with Applications in Statistics, Second Edition KHURI, MATHEW, and SINHA ·Statistical Tests for Mixed Linear Models KLUGMAN, PANJER, and WILLMOT ·Loss Models: From Data to Decisions KLUGMAN, PANJER, and WILLMOT ·Solutions Manual to Accompany Loss Models: From Data to Decisions *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 5 KOTZ, BALAKRISHNAN, and JOHNSON ·Continuous Multivariate Distributions, Volume 1, Second Edition KOTZ and JOHNSON (editors) ·Encyclopedia of Statistical Sciences: Volumes 1 to 9 with Index KOTZ and JOHNSON (editors) ·Encyclopedia of Statistical Sciences: Supplement Volume KOTZ, READ, and BANKS (editors) ·Encyclopedia of Statistical Sciences: Update Volume 1 KOTZ, READ, and BANKS (editors) ·Encyclopedia of Statistical Sciences: Update Volume 2 KOVALENKO, KUZNETZOV, and PEGG ·Mathematical Theory of Reliability of Time-Dependent Systems with Practical Applications LACHIN ·Biostatistical Methods: The Assessment of Relative Risks LAD ·Operational Subjective Statistical Methods: A Mathematical, Philosophical, and Historical Introduction LAMPERTI ·Probability: A Survey of the Mathematical Theory, Second Edition LANGE, RYAN, BILLARD, BRILLINGER, CONQUEST, and GREENHOUSE · Case Studies in Biometry LARSON ·Introduction to Probability Theory and Statistical Inference, Third Edition LAWLESS ·Statistical Models and Methods for Lifetime Data, Second Edition LAWSON ·Statistical Methods in Spatial Epidemiology LE ·Applied Categorical Data Analysis LE ·Applied Survival Analysis LEE and WANG ·Statistical Methods for Survival Data Analysis, Third Edition LEPAGE and BILLARD ·Exploring the Limits of Bootstrap LEYLAND and GOLDSTEIN (editors) ·Multilevel Modelling of Health Statistics LIAO ·Statistical Group Comparison LINDVALL ·Lectures on the Coupling Method LINHART and ZUCCHINI ·Model Selection LITTLE and RUBIN ·Statistical Analysis with Missing Data, Second Edition LLOYD ·The Statistical Analysis of Categorical Data MAGNUS and NEUDECKER ·Matrix Differential Calculus with Applications in Statistics and Econometrics, Revised Edition MALLER and ZHOU ·Survival Analysis with Long Term Survivors MALLOWS ·Design, Data, and Analysis by Some Friends of Cuthbert Daniel MANN, SCHAFER, and SINGPURWALLA ·Methods for Statistical Analysis of Reliability and Life Data MANTON, WOODBURY, and TOLLEY ·Statistical Applications Using Fuzzy Sets MARDIA and JUPP ·Directional Statistics MASON, GUNST, and HESS ·Statistical Design and Analysis of Experiments with Applications to Engineering and Science, Second Edition McCULLOCH and SEARLE ·Generalized, Linear, and Mixed Models McFADDEN ·Management of Data in Clinical Trials McLACHLAN ·Discriminant Analysis and Statistical Pattern Recognition McLACHLAN and KRISHNAN ·The EM Algorithm and Extensions McLACHLAN and PEEL ·Finite Mixture Models McNEIL ·Epidemiological Research Methods MEEKER and ESCOBAR ·Statistical Methods for Reliability Data MEERSCHAERT and SCHEFFLER ·Limit Distributions for Sums of Independent Random Vectors: Heavy Tails in Theory and Practice *MILLER ·Survival Analysis, Second Edition MONTGOMERY, PECK, and VINING ·Introduction to Linear Regression Analysis, Third Edition MORGENTHALER and TUKEY ·Configural Polysampling: A Route to Practical Robustness *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 6 MUIRHEAD ·Aspects of Multivariate Statistical Theory MURRAY ·X-STAT 2.0 Statistical Experimentation, Design Data Analysis, and Nonlinear Optimization MYERS and MONTGOMERY ·Response Surface Methodology: Process and Product Optimization Using Designed Experiments, Second Edition MYERS, MONTGOMERY, and VINING ·Generalized Linear Models. With Applications in Engineering and the Sciences NELSON ·Accelerated Testing, Statistical Models, Test Plans, and Data Analyses NELSON ·Applied Life Data Analysis NEWMAN ·Biostatistical Methods in Epidemiology OCHI ·Applied Probability and Stochastic Processes in Engineering and Physical Sciences OKABE, BOOTS, SUGIHARA, and CHIU ·Spatial Tesselations: Concepts and Applications of Voronoi Diagrams, Second Edition OLIVER and SMITH ·Influence Diagrams, Belief Nets and Decision Analysis PANKRATZ ·Forecasting with Dynamic Regression Models PANKRATZ ·Forecasting with Univariate Box-Jenkins Models: Concepts and Cases *PARZEN ·Modern Probability Theory and Its Applications PEÑA, TIAO, and TSAY ·A Course in Time Series Analysis PIANTADOSI ·Clinical Trials: A Methodologic Perspective PORT ·Theoretical Probability for Applications POURAHMADI ·Foundations of Time Series Analysis and Prediction Theory PRESS ·Bayesian Statistics: Principles, Models, and Applications PRESS ·Subjective and Objective Bayesian Statistics, Second Edition PRESS and TANUR ·The Subjectivity of Scientists and the Bayesian Approach PUKELSHEIM ·Optimal Experimental Design PURI, VILAPLANA, and WERTZ ·New Perspectives in Theoretical and Applied Statistics PUTERMAN ·Markov Decision Processes: Discrete Stochastic Dynamic Programming *RAO ·Linear Statistical Inference and Its Applications, Second Edition RENCHER ·Linear Models in Statistics RENCHER ·Methods of Multivariate Analysis, Second Edition RENCHER ·Multivariate Statistical Inference with Applications RIPLEY ·Spatial Statistics RIPLEY ·Stochastic Simulation ROBINSON ·Practical Strategies for Experimenting ROHATGI and SALEH ·An Introduction to Probability and Statistics, Second Edition ROLSKI, SCHMIDLI, SCHMIDT, and TEUGELS ·Stochastic Processes for Insurance and Finance ROSENBERGER and LACHIN ·Randomization in Clinical Trials: Theory and Practice ROSS ·Introduction to Probability and Statistics for Engineers and Scientists ROUSSEEUW and LEROY ·Robust Regression and Outlier Detection RUBIN ·Multiple Imputation for Nonresponse in Surveys RUBINSTEIN ·Simulation and the Monte Carlo Method RUBINSTEIN and MELAMED ·Modern Simulation and Modeling RYAN ·Modern Regression Methods RYAN ·Statistical Methods for Quality Improvement, Second Edition SALTELLI, CHAN, and SCOTT (editors) ·Sensitivity Analysis *SCHEFFE ·The Analysis of Variance SCHIMEK ·Smoothing and Regression: Approaches, Computation, and Application SCHOTT ·Matrix Analysis for Statistics SCHUSS ·Theory and Applications of Stochastic Differential Equations SCOTT ·Multivariate Density Estimation: Theory, Practice, and Visualization *SEARLE ·Linear Models SEARLE ·Linear Models for Unbalanced Data *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 7 SEARLE ·Matrix Algebra Useful for Statistics SEARLE, CASELLA, and McCULLOCH ·Variance Components SEARLE and WILLETT ·Matrix Algebra for Applied Economics SEBER and LEE ·Linear Regression Analysis, Second Edition SEBER ·Multivariate Observations SEBER and WILD ·Nonlinear Regression SENNOTT ·Stochastic Dynamic Programming and the Control of Queueing Systems *SERFLING ·Approximation Theorems of Mathematical Statistics SHAFER and VOVK ·Probability and Finance: It’s Only a Game! SMALL and M CLEISH ·Hilbert Space Methods in Probability and Statistical Inference SRIVASTAVA ·Methods of Multivariate Statistics STAPLETON ·Linear Statistical Models STAUDTE and SHEATHER ·Robust Estimation and Testing STOYAN, KENDALL, and MECKE ·Stochastic Geometry and Its Applications, Second Edition STOYAN and STOYAN ·Fractals, Random Shapes and Point Fields: Methods of Geometrical Statistics STYAN ·The Collected Papers of T. W. Anderson: 1943–1985 SUTTON, ABRAMS, JONES, SHELDON, and SONG ·Methods for Meta-Analysis in Medical Research TANAKA ·Time Series Analysis: Nonstationary and Noninvertible Distribution Theory THOMPSON ·Empirical Model Building THOMPSON ·Sampling, Second Edition THOMPSON ·Simulation: A Modeler’s Approach THOMPSON and SEBER ·Adaptive Sampling THOMPSON, WILLIAMS, and FINDLAY ·Models for Investors in Real World Markets TIAO, BISGAARD, HILL, PEÑA, and STIGLER (editors) ·Box on Quality and Discovery: with Design, Control, and Robustness TIERNEY ·LISP-STAT: An Object-Oriented Environment for Statistical Computing and Dynamic Graphics TSAY ·Analysis of Financial Time Series UPTON and FINGLETON ·Spatial Data Analysis by Example, Volume II: Categorical and Directional Data VAN BELLE ·Statistical Rules of Thumb VIDAKOVIC ·Statistical Modeling by Wavelets WEISBERG ·Applied Linear Regression, Second Edition WELSH ·Aspects of Statistical Inference WESTFALL and YOUNG ·Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment WHITTAKER ·Graphical Models in Applied Multivariate Statistics WINKER ·Optimization Heuristics in Economics: Applications of Threshold Accepting WONNACOTT and WONNACOTT ·Econometrics, Second Edition WOODING ·Planning Pharmaceutical Clinical Trials: Basic Statistical Principles WOOLSON and CLARKE ·Statistical Methods for the Analysis of Biomedical Data, Second Edition WU and HAMADA ·Experiments: Planning, Analysis, and Parameter Design Optimization YANG ·The Construction Theory of Denumerable Markov Processes *ZELLNER ·An Introduction to Bayesian Inference in Econometrics ZHOU, OBUCHOWSKI, and M CCLISH ·Statistical Methods in Diagnostic Medicine *Now available in a lower priced paperback edition in the Wiley Classics Library.p&s-cp.qxd 3/25/03 9:47 AM Page 8 The Third Edition ofthe leading reference onsurvival data analysis ‘hestudy ofsurvival dataattempts topredict theprobability ofresponse, survival, ormean lifetime; compare; thesurvival distributions ofexperimental animals orofhuman patients; andidentify riskand/or prognosis factors. Statistical Methods forSurvival Data Analysis, Third Edition examines the statistical methods foranalyzing survival datafromlaboratory studiesofanimals, clinical andepidemi- ological studies ofhumans, andother appropriate applications. Emphasizing applicarions over rigorous mathematics, thisextremely useful reference provides thor- ‘ough discussions ofthemost commonly usedparametric andnonparametric methods insurvival analy- sis,aswell asguidelines fortheplanning anddesign ofclinical trials. The authors give special consideration tothestudy ofsurvival datainbiomedical sciences, though themethods aresuitable for applications inindustrial reliability, thesocialsciences, andbusiness. This Third Edison brings thisstandard intheficlduptodatewith newmaterial andrevised refer- ‘ences including: ‘Anew introduction toleftand interval censored dara Thegeneralized gamma andlog-logistic distribution Estimation procedures forleftandinterval censored dara Parametric models with covariates ‘Cox'sproportional hazardsmodelincluding stratification andtime-dependent covariates, andsome non-proportional hazards models Goodness-of- Fitrestsandmodelselection methods Multiple responses tothelogistic regression model Numerous real-life examples which illustrate keyconcepts ‘Computer programming codes inSAS, BMDP, andSPSS formost examples Related FTPsiteproviding lange datasets These additions andrevisions make Statistical Methods forSurvival Data Analysis, Third Edition, more valuable than everasanessential reference forbiomedical investigators, statisticians, epidemiologists, andresearchers inother disciplines involved orinterested intheanalysis ofsurvival data. isGeorge Lynn Cross Research Professor ofBiostatistics andEpidemiology and Director oftheCenter forAmerican Indian Health Research attheUniversity ofOklahoma Health Sciences Center. Shereceived amaster’s degree from theUniversity ofCalifornia atBerkeley andher doctorate from New York University. Theauthor oftheprevious editions ofStatistical Methods for Survival Data Analysis, Professor LeeisaFellowoftheAmerican Statistical Association andmember of theSociety forEpidemiological Research andtheAmerican Diabetes Association. I 1G,F isanAssociate Professor ofBiostatistics attheUniversity of Oklahoma Health Sciences Center. Hereceived amaster’s degree from theAcademy ofSciences of Chinaandhisdoctorare fromtheUniversity ofMaryland. ‘Subscribe tooutfreeStatisticseNewsletter at ISBN 0-47L-36997-7wwrewdley.com jenewsletters. r 90000 i ONES 1 INTERSCIENCEoWagereneel Ae