Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Probability and Statistics / Survival Analysis

survival analysis_Ibrahim_J

PDF · 491 pages · 2.3 MB
Open PDF file

Lecture notes for a survival analysis course, apparently by or associated with Ibrahim (per the file name), kept in the Probability and Statistics folder. The opening sections cover time-to-event data, right, left and interval censoring, independent versus informative censoring, and Type I/II/III censoring. They also present example clinical datasets (nursing home stays, MAC prevention trial, UMARU drug study, halibut survival) and define the density, survivor, hazard and cumulative hazard functions. Only the first part of a long document was read.

AI-written summary; may contain errors. This description is approximate.

Extracted text (machine-read; may contain errors)
SurvivalAnalysis: Introduction SurvivalAnalysis typicallyfocusesontimetoeventdata. Inthemostgeneralsense,itconsistsoftechniquesforpositive- valuedrandomvariables, suchas ²timetodeath ²timetoonset(orrelapse)ofadisease ²lengthofstayinahospital ²duration ofastrike ²moneypaidbyhealthinsurance ²viralloadmeasuremen ts ²timeto¯nishing adoctoraldissertation! Kindsofsurvivalstudiesinclude: ²clinicaltrials ²prospectivecohortstudies ²retrospectivecohortstudies ²retrospectivecorrelativ estudies Typically,survivaldataarenotfullyobserved,butrather arecensored. 1 Inthiscourse,wewill: ²describesurvivaldata ²comparesurvivalofseveralgroups ²explainsurvivalwithcovariates ²designstudieswithsurvivalendpoints Someknowledgeofdiscrete datamethodswillbeuseful, sinceanalysis ofthe\timetoevent"usesinformation from thediscrete(i.e.,binary)outcome ofwhether theeventoc- curredornot. Someusefulreferences: ²Collett:ModellingSurvivalDatainMedicalResearch ²CoxandOakes:AnalysisofSurvivalData ²Kalb°eisc handPrentice:TheStatisticalAnalysisof FailureTimeData ²Lee:StatisticalMethodsforSurvivalDataAnalysis ²Fleming &Harrington: Counting ProcessesandSur- vivalAnalysis ²Hosmer&Lemesho w:AppliedSurvivalAnalysis ²Kleinbaum:SurvivalAnalysis:Aself-learningtext 2 ²Klein&Moeschberger:SurvivalAnalysis:Techniques forcensoredandtruncateddata ²Cantor:Extending SASSurvivalAnalysisTechniques forMedicalResearch ²Allison:SurvivalAnalysisUsingtheSASSystem ²Jennison &Turnbull:GroupSequentialMethodswith ApplicationstoClinicalTrials ²Ibrahim, Chen,&Sinha:Bayesian SurvivalAnalysis 3 SomeDe¯nitionsandnotation Failuretimerandomvariables arealwaysnon-negative. Thatis,ifwedenotethefailuretimebyT,thenT¸0. Tcaneitherbediscrete (takinga¯nitesetofvalues,e.g. a1;a2;:::;an)orcontinuous(de¯ned on(0;1)). ArandomvariableXiscalledacensoredfailuretime randomvariable ifX=min(T;U),whereUisanon- negativecensoring variable. Inordertode¯neafailuretimerandomvariable, weneed: (1)anunambiguoustimeorigin (e.g.randomization toclinicaltrial,purchaseofcar) (2)atimescale (e.g.realtime(days,years),mileageofacar) (3)de¯nition oftheevent (e.g.death,needanewcartransmission) 4 Illustrationofsurvivaldata X X y y X y X y study opensstudy closes y=censored observation X=event 5 Theillustration ofsurvivaldataontheprevious pageshows severalfeatures whicharetypicallyencounteredinanalysis ofsurvivaldata: ²individuals donotallenterthestudyatthesametime ²whenthestudyends,someindividuals stillhaven'thad theeventyet ²otherindividuals dropoutorgetlostinthemiddleof thestudy,andallweknowaboutthemisthelasttime theywerestill\free"oftheevent The¯rstfeatureisreferredtoas\staggeredentry" Thelasttwofeatures relateto\censoring" ofthefailure timeevents. 6 Typesofcensoring: ²Right-censoring : onlyther.v.Xi=min(Ti;Ui)isobserveddueto {losstofollow-up {drop-out {studytermination Wecallthisright-censoring becausethetrueunobserv ed eventistotherightofourcensoring time;i.e.,allwe knowisthattheeventhasnothappenedattheendof follow-up. Inaddition toobservingXi,wealsogettoseethefail- ureindicator : ±i=8 >< >:1ifTi·Ui 0ifTi>Ui Somesoftwarepackagesinsteadassumewehavea censoringindicator : ci=8 >< >:0ifTi·Ui 1ifTi>Ui Right-censoring isthemostcommon typeofcensoring assumption wewilldealwithinsurvivalanalysis. 7 ²Left-censoring CanonlyobserveYi=max(Ti;Ui)andthefailureindi- cators: ±i=8 >< >:1ifUi·Ti 0ifUi>Ti e.g.(Miller)studyofageatwhichAfricanchildrenlearn atask.Somealreadyknew(left-censored), somelearned duringstudy(exact),somehadnotyetlearnedbyend ofstudy(right-censored). ²Interval-censoring Observe(Li;Ri)whereTi2(Li;Ri) Ex.1:Timetoprostate cancer,observelongitudinal PSAmeasuremen ts Ex.2:Timetoundetectable viralloadinAIDSstudies, basedonmeasuremen tsofviralloadtakenateachclinic visit Ex.3:Detectrecurrence ofcoloncanceraftersurgery. Followpatientsevery3monthsafterresection ofprimary tumor. 8 Independentvsinformativecensoring ²Wesaycensoring isindependent(non-informativ e)if UiisindependentofTi. {Ex.1IfUiistheplanned endofthestudy(say,2 yearsafterthestudyopens),thenitisusuallyinde- pendentoftheeventtimes. {Ex.2IfUiisthetimethatapatientdropsout ofthestudybecausehe/shegotmuchsickerand/or hadtodiscontinuetakingthestudytreatmen t,then UiandTiareprobably notindependent. AnindividualcensoredatUshouldberepre- sentativeofallsubjectswhosurvivetoU. Thismeansthatcensoring atUcoulddependonprog- nosticcharacteristics measured atbaseline, butthatamong allthosewiththesamebaselinecharacteristics, theprob- abilityofcensoring priortoorattimeUshouldbethe same. ²Censoring isconsideredinformativeifthedistribu- tionofUicontainsanyinformation abouttheparameters characterizing thedistribution ofTi. 9 Supposewehaveasampleofobservationsonnpeople: (T1;U1);(T2;U2);:::;(Tn;Un) Therearethreemaintypesof(right)censoring times: ²TypeI:AlltheUi'sarethesame e.g.animalstudies,allanimalssacri¯ced after2years ²TypeII:Ui=T(r),thetimeoftherthfailure. e.g.animalstudies,stopwhen4/6havetumors ²TypeIII:theUi'sarerandomvariables,±i'sarefailure indicators: ±i=8 >< >:1ifTi·Ui 0ifTi>Ui TypeIandTypeIIarecalledsinglycensored data, TypeIIIiscalledrandomly censored (orsometimes pro- gressively censored). 10 Someexampledatasets: ExampleA.Durationofnursinghomestay (Morrisetal.,CaseStudiesinBiometry ,Ch12) TheNational CenterforHealthServices Researchstudied 36for-pro¯t nursinghomestoassessthee®ectsofdi®erent ¯nancial incentivesonlengthofstay.\Treated"nursing homesreceivedhigherperdiemsforMedicaid patients,and bonusesforimprovingapatient'shealthandsendingthem home. Studyincluded 1601patientsadmitted betweenMay1,1981 andApril30,1982. Variablesinclude: LOS-Lengthofstayofaresident(indays) AGE-Ageofaresident RX-Nursinghomeassignmen t(1:bonuses,0:nobonuses) GENDER -Gender(1:male, 0:female) MARRIED -(1:married, 0:notmarried) HEALTH-healthstatus(2:second best,5:worst) CENSOR -Censoring indicator (1:censored, 0:discharged) Firstfewlinesofdata: 378610020 617710040 11 ExampleB.Fecundability Womenwhohadrecentlygivenbirthwereaskedtorecall howlongittookthemtobecomepregnant,andwhether or nottheysmokedduringthattime.Theoutcome ofinter- est(summarized below)istimetopregnancy (measured in menstrual cycles). 19subjectswerenotabletogetpregnantafter12months. CycleSmokersNon-smok ers 129 198 216 107 317 55 4 4 38 5 3 18 6 9 22 7 4 7 8 5 9 9 1 5 10 1 3 11 1 6 12 3 6 12+ 7 12 12 ExampleC:MACPreventionClinicalTrial ACTG196wasarandomized clinicaltrialtostudythee®ects ofcombination regimens onpreventionofMAC(mycobac- teriumaviumcomplex),oneofthemostcommon oppor- tunisticinfections inAIDSpatients. Thetreatmentregimens were: ²clarithrom ycin(new) ²rifabutin (standard) ²clarithrom ycinplusrifabutin Othercharacteristics oftrial: ²PatientsenrolledbetweenApril1993andFebruary1994 ²Follow-upendedAugust1995 ²InFebruary1994,rifabutin dosagewasreduced from3 pills/day(450mg) to2pills/day(300mg) duetoconcern overuveitis1 Themainintent-to-treat analysiscompared the3treatmen t armswithoutadjusting forthischangeindosage. 1Uveitisisanadverseexperienceresultinginin°ammationofthe uvealtractintheeyes(about3-4%ofpatientsreporteduveitis). 13 ExampleD:HMOStudyofHIV-relatedSurvival Thisishypothetical datausedbyHosmer&Lemesho w(de- scribedonpages2-17)containing100observationsonHIV+ subjectsbelonging toanHealthMaintenanceOrganization (HMO). TheHMOwantstoevaluatethesurvivaltimeof thesesubjects.Inthishypothetical dataset, subjectswere enrolled fromJanuary1,1989untilDecember31,1991. StudyfollowupthenendedonDecember31,1995. Variables: ID SubjectID(1-100) TIME Survivaltimeinmonths ENTDATEEntrydate ENDDATEDatefollow-upendedduetodeathorcensoring CENSOR DeathIndicator (1=death, 0=censor) AGE Ageofsubjectinyears DRUG HistoryofIVDrugUse(0=no,1=y es) ThisdatasetisusedbyHosmer &Lemesho wtomotivate someconcepts insurvivalanalysisinChap.1oftheirbook. 14 ExampleE:UMARUImpactStudy(UIS) ThisdatasetcomesfromtheUniversityofMassachusetts AIDSResearchUnit(UMARU)IMPACTStudy,a5-year collaborativeresearchprojectcomprised oftwoconcurren t randomized trialsofresidentialtreatmen tfordrugabuse. (1)ProgramA:Randomized 444subjectstoa3-or6- monthprogram ofhealtheducation andrelapsepreven- tion.Clientsweretaughttorecognize \high-risk" situ- ationsthataretriggerstorelapse,andtaughtskillsto copewiththesesituations withoutusingdrugs. (2)ProgramB:Randomized 184participan tstoa6-or 12-monthprogram withhighlystructured life-styleina communallivingsetting. Variables: ID SubjectID(1-628) AGEAgeinyears BECKTOTABeckDepressionScore HERCOCHeroinorCocaineUsepriortoentry IVHX IVDruguseatAdmission NDRUGTXNumberpreviousdrugtreatments RACESubject'sRace(0=White,1=Other) TREATTreatmentAssignment(0=short,1=long) SITE TreatmentProgram(0=A,1=B) LOT LengthofTreatment(days) TIME TimetoReturntoDrugUse(days) CENSOR IndicatorofDrugUseRelapse(1=yes,0=censored) 15 ExampleF:AtlanticHalibutSurvivalTimes Oneconservationmeasure suggested fortrawl¯shingisa minimumsizelimitforhalibut(32inches).However,thissize limitwouldonlybee®ectiveifcaptured ¯shbelowthelimit surviveduntilthetimeoftheirrelease.Anexperimentwas conducted toevaluatethesurvivalratesofhalibutcaughtby trawlsorlonglines, andtoassessotherfactorswhichmight contributetosurvival(duration oftrawling,maximumdepth ¯shed,sizeof¯sh,andhandling time). AnarticlebySmith,WaiwoodandNeilson,SurvivalAnaly- sisforSizeRegulationofAtlanticHalibutinCaseStudies inBiometry compares parametric survivalmodelstosemi- parametric survivalmodelsinevaluating thisdata. Survival TowDi® Length HandlingTotal Obs Time CensoringDurationinofFishTime log(catch) #(min) Indicator (min.) Depth(cm) (min.) ln(weight) 100353.0 1 301539 5 5.685 109111.0 1 100 544 29 8.690 11364.0 0 100 1053 4 5.323 116500.0 1 100 1044 4 5.323 .... 16 MoreDe¯nitionsandNotation Thereareseveralequivalentwaystocharacterize theprob- abilitydistribution ofasurvivalrandomvariable. Someof thesearefamiliar; othersarespecialtosurvivalanalysis. We willfocusonthefollowingterms: ²Thedensityfunctionf(t) ²ThesurvivorfunctionS(t) ²Thehazardfunction¸(t) ²Thecumulativehazardfunction ¤(t) ²Densityfunction(orProbabilityMassFunc- tion)fordiscreter.v.'s SupposethatTtakesvaluesina1;a2;:::;an. f(t)=Pr(T=t) =8 >>< >>:fjift=aj;j=1;2;:::;n 0ift6=aj;j=1;2;:::;n ²DensityFunctionforcontinuousr.v.'s f(t)=lim ¢t!01 ¢tPr(t·T·t+¢t) 17 ²SurvivorshipFunction :S(t)=P(T¸t). Inothersettings, thecumulativedistribution function, F(t)=P(T·t),isofinterest.Insurvivalanalysis, our interesttendstofocusonthesurvivalfunction,S(t). Foracontinuousrandomvariable: S(t)=Z1 tf(u)du Foradiscreterandomvariable: S(t)=X u¸tf(u) =X aj¸tf(aj) =X aj¸tfj Notes: ²Fromthede¯nition ofS(t)foracontinuousvariable, S(t)=1¡F(t)aslongasF(t)isabsolutely continuous w.r.ttheLebesguemeasure. [Thatis,F(t)hasadensity function.] ²Foradiscretevariable,wehavetodecidewhattodoif aneventoccursexactlyattimet;i.e.,doesthatbecome partofF(t)orS(t)? ²Togetaroundthisproblem, severalbooksde¯ne S(t)=Pr(T>t),orelsede¯neF(t)=Pr(T<t) (eg.Collett) 18 ²HazardFunction¸(t) Sometimes calledaninstantane ousfailurerate,the forceofmortality ,ortheage-speci¯cfailurerate. {Continuousrandomvariables: ¸(t)=lim ¢t!01 ¢tPr(t·T<t+¢tjT¸t) =lim ¢t!01 ¢tPr([t·T<t+¢t]T[T¸t]) Pr(T¸t) =lim ¢t!01 ¢tPr(t·T<t+¢t) Pr(T¸t) =f(t) S(t) {Discreterandomvariables: ¸(aj)´¸j=Pr(T=ajjT¸aj) =P(T=aj) P(T¸aj) =f(aj) S(aj) =f(t) P k:ak¸ajf(ak) 19 ²CumulativeHazardFunction ¤(t) {Continuousrandomvariables: ¤(t)=Zt 0¸(u)du {Discreterandomvariables: ¤(t)=X k:ak<t¸k 20 RelationshipbetweenS(t)and¸(t) We'vealreadyshownthat,foracontinuousr.v. ¸(t)=f(t) S(t) Foraleft-continuoussurvivorfunctionS(t),wecanshow: f(t)=¡S0(t)orS0(t)=¡f(t) Wecanusethisrelationship toshowthat: ¡d dt[logS(t)]=¡0 B@1 S(t)1 CAS0(t) =¡¡f(t) S(t) =f(t) S(t) Soanotherwaytowrite¸(t)isasfollows: ¸(t)=¡d dt[logS(t)] 21 RelationshipbetweenS(t)and¤(t): ²Continuouscase: ¤(t)=Zt 0¸(u)du =Zt 0f(u) S(u)du =Zt 0¡d dulogS(u)du =¡logS(t)+logS(0) )S(t)=e¡¤(t) ²Discretecase: Supposethataj<t·aj+1.Then S(t)=P(T¸a1;T¸a2;:::;T¸aj+1) =P(T¸a1)P(T¸a2jT¸a1)¢¢¢P(T¸aj+1jT¸aj) =(1¡¸1)£¢¢¢£(1¡¸j) =Y k:ak<t(1¡¸k) Coxde¯nes¤(t)=P k:ak<tlog(1¡¸k)sothatS(t)= e¡¤(t)inthediscretecase,aswell. 22 MeasuringCentralTendencyinSurvival ²Meansurvival-callthis¹ ¹=Z1 0uf(u)duforcontinuousT =nX j=1ajfjfordiscreteT ²Mediansurvival-callthis¿,isde¯nedby S(¿)=0:5 Similarly ,anyotherpercentilecouldbede¯ned. Inpractice, wedon'tusuallyhitthemediansurvival atexactlyoneofthefailuretimes.Inthiscase,the estimated mediansurvivalisthesmallesttime¿such that ^S(¿)·0:5 23 Somehazardshapesseeninapplications: ²increasing e.g.agingafter65 ²decreasing e.g.survivalaftersurgery ²bathtub e.g.age-speci¯cmortality ²constant e.g.survivalofpatientswithadvancedchronicdisease 24 Estimatingthesurvivalorhazardfunction Wecanestimate thesurvival(orhazard) function intwo ways: ²byspecifying aparametric modelfor¸(t)basedona particular densityfunctionf(t) ²bydevelopinganempirical estimate ofthesurvivalfunc- tion(i.e.,non-parametric estimation) Ifnocensoring: Theempirical estimate ofthesurvivalfunction, ~S(t),isthe proportionofindividuals witheventtimesgreaterthant. Withcensoring: Iftherearecensored observations,then~S(t)isnotagood estimate ofthetrueS(t),soothernon-parametric methods mustbeusedtoaccountforcensoring (life-table methods, Kaplan-Meier estimator) 25 SomeParametricSurvivalDistributions ²TheExponentialdistribution (1parameter) f(t)=¸e¡¸tfort¸0 S(t)=Z1 tf(u)du =e¡¸t ¸(t)=f(t) S(t) =¸constanthazard! ¤(t)=Zt 0¸(u)du =Zt 0¸du =¸t Check:DoesS(t)=e¡¤(t)? median: solve0:5=S(¿)=e¡¸¿: )¿=¡log(0:5) ¸ mean: Z1 0u¸e¡¸udu=1 ¸ 26 ²TheWeibulldistribution (2parameters) Generalizes exponential: S(t)=e¡¸t· f(t)=¡d dtS(t)=·¸t·¡1e¡¸t· ¸(t)=·¸t·¡1 ¤(t)=Zt 0¸(u)du=¸t· ¸-thescaleparameter ·-theshapeparameter TheWeibulldistribution isconvenientbecauseofitssim- pleform.Itincludes severalhazardshapes: ·=1!constanthazard 0<·<1!decreasing hazard ·>1!increasing hazard 27 ²Rayleighdistribution Another 2-parameter generalization ofexponential: ¸(t)=¸0+¸1t ²compoundexponential T»exp(¸);¸»g f(t)=Z1 0¸e¡¸tg(¸)d¸ ²log-normal ,log-logistic : Possibledistributions forTobtained byspecifying for logTanyconvenientfamilyofdistributions, e.g. logT»normal(non-monotone hazard) logT»logistic 28 Whyuseoneversusanother? ²technicalconvenienceforestimation andinference ²explicitsimpleformsforf(t);S(t),and¸(t). ²qualitativ eshapeofhazardfunction Onecanusuallydistinguish betweenaone-parameter model (liketheexponential)andtwo-parameter (likeWeibullor log-normal) intermsoftheadequacy of¯ttoadataset. Without alotofdata,itmaybehardtodistinguish between the¯tsofvarious2-parameter models(i.e.,Weibullvslog- normal) 29 PlotsofestimatesofS(t) BasedonKM,exponential,Weibull,andlog-normal forstudyofprotease inhibitors inAIDSpatients (ACTG320) KM Curves for Time to PCP - 2 Drug Arm daysProbability 0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00 1KM Exponential Weibull LognormalKM Curves for Time to PCP - 3 Drug Arm daysProbability 0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00 1KM Exponential Weibull Lognormal 30 PlotsofestimatesofS(t) BasedonKM,exponential,Weibull,andlog-normal forstudyofprotease inhibitors inAIDSpatients (ACTG320) KM Curves for Time to MAC - 2 Drug Arm daysProbability 0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00 1KM Exponential Weibull LognormalKM Curves for Time to MAC - 3 Drug Arm daysProbability 0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00 1KM Exponential Weibull Lognormal 31 PlotsofestimatesofS(t) BasedonKM,exponential,Weibull,andlog-normal forstudyofprotease inhibitors inAIDSpatients (ACTG320) KM Curves for Time to CMV - 2 Drug Arm daysProbability 0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00 1KM Exponential Weibull LognormalKM Curves for Time to CMV - 3 Drug Arm daysProbability 0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00 1KM Exponential Weibull Lognormal 32 PreviewofComingAttractions Nextwewilldiscussthemostfamousnon-parametric ap- proachforestimating thesurvivaldistribution, calledthe Kaplan-Meier estimator . Tomotivatethederivationofthisestimator, wewill¯rst consider asetofsurvivaltimeswherethereisnocensoring. Thefollowingaretimestorelapse (weeks)for21leukemia patientsreceiving controltreatmen t(Table1.1ofCox& Oakes): 1,1,2,2,3,4,4,5,5,8,8,8,8,11,11,12,12,15,17,22,23 Howwouldweestimate S(10),theprobabilit ythatanindi- vidualsurvivestotime10orlater? Whatabout~S(8)? Isit12 21or8 21? 33 Let'sconstruct atableof~S(t): Valuesoft^S(t) t·121/21=1.000 1<t·219/21=0.905 2<t·317/21=0.809 3<t·4 4<t·5 5<t·8 8<t·11 11<t·12 12<t·15 15<t·17 17<t·22 22<t·23 EmpiricalSurvivalFunction: Whenthereisnocensoring, thegeneralformulais: ~S(t)=#individualswithT¸t totalsamplesize 34 Inmostsoftwarepackages,thesurvivalfunction isevaluated justaftertimet,i.e.,att+.Inthiscase,weonlycountthe individuals withT>t. Example forleukemiadata(controlarm): 35 StataCommandsforSurvivalEstimation .useleukem .stsetremissstatusiftrt==0 (tokeeponlyuntreated patients) (21observations deleted) .stslist failure _d:status analysis time_t:remiss Beg. NetSurvivor Std. Time Total FailLostFunction Error [95%Conf.Int.] ----------------------------------------------------------- ----------- 1 21 200.9048 0.0641 0.6700 0.9753 2 19 200.8095 0.0857 0.5689 0.9239 3 17 100.7619 0.0929 0.5194 0.8933 4 16 200.6667 0.1029 0.4254 0.8250 5 14 200.5714 0.1080 0.3380 0.7492 8 12 400.3810 0.1060 0.1831 0.5778 11 8200.2857 0.0986 0.1166 0.4818 12 6200.1905 0.0857 0.0595 0.3774 15 4100.1429 0.0764 0.0357 0.3212 17 3100.0952 0.0641 0.0163 0.2612 22 2100.0476 0.0465 0.0033 0.1970 23 1100.0000 . . . ----------------------------------------------------------- ----------- .stsgraph 36 SASCommandsforSurvivalEstimation dataleuk; inputt; cards; 1 1 2 2 3 4 4 5 5 8 8 8 8 11 11 12 12 15 17 22 23 ; proclifetest data=leuk; timet; run; 37 SASOutputforSurvivalEstimation TheLIFETEST Procedure Product-Limit Survival Estimates Survival Standard Number Number tSurvival Failure Error Failed Left 0.0000 1.0000 0 0 0 21 1.0000 . . . 1 20 1.0000 0.9048 0.0952 0.0641 2 19 2.0000 . . . 3 18 2.0000 0.8095 0.1905 0.0857 4 17 3.0000 0.7619 0.2381 0.0929 5 16 4.0000 . . . 6 15 4.0000 0.6667 0.3333 0.1029 7 14 5.0000 . . . 8 13 5.0000 0.5714 0.4286 0.1080 9 12 8.0000 . . . 10 11 8.0000 . . . 11 10 8.0000 . . . 12 9 8.0000 0.3810 0.6190 0.1060 13 8 11.0000 . . . 14 7 11.0000 0.2857 0.7143 0.0986 15 6 12.0000 . . . 16 5 12.0000 0.1905 0.8095 0.0857 17 4 15.0000 0.1429 0.8571 0.0764 18 3 17.0000 0.0952 0.9048 0.0641 19 2 22.0000 0.0476 0.9524 0.0465 20 1 23.0000 01.0000 0 21 0 38 SASOutputforSurvivalEstimation(cont'd) Summary Statistics forTimeVariable t Quartile Estimates Point 95%Confidence Interval Percent Estimate [Lower Upper) 7512.0000 8.0000 17.0000 50 8.0000 4.0000 11.0000 25 4.0000 2.0000 8.0000 Mean Standard Error 8.6667 1.4114 Summary oftheNumberofCensored andUncensored Values Percent TotalFailed Censored Censored 21 21 0 0.00 39 Doesanyonehaveaguessregardinghowtocalcu- latethestandarderroroftheestimatedsurvival? ^S(8+)=P(T>8)=8 21=0:381 (att=8+,wecountthe4eventsattime=8asalready havingfailed) se[^S(8+)]=0:106 40 S-PlusCommandsforSurvivalEstimation >t_c(1,1,2,2,3,4,4,5,5,8,8,8,8,11,11,12,12,15,17,22,23) >surv.fit(t,status=rep(1,21)) 95percent confidence interval isoftype"log" timen.riskn.event survival std.dev lower95%CIupper95%CI 121 20.90476190 0.06405645 0.78753505 1.0000000 219 20.80952381 0.08568909 0.65785306 0.9961629 317 10.76190476 0.09294286 0.59988048 0.9676909 416 20.66666667 0.10286890 0.49268063 0.9020944 514 20.57142857 0.10798985 0.39454812 0.8276066 812 40.38095238 0.10597117 0.22084536 0.6571327 11 8 20.28571429 0.09858079 0.14529127 0.5618552 12 6 20.19047619 0.08568909 0.07887014 0.4600116 15 4 10.14285714 0.07636035 0.05010898 0.4072755 17 3 10.09523810 0.06405645 0.02548583 0.3558956 22 2 10.04761905 0.04647143 0.00703223 0.3224544 23 1 10.00000000 NA NA NA 41 Estimating theSurvivalFunction One-samplenonparametric methods: Wewillconsider threemethodsforestimating asurvivorship function S(t)=Pr(T¸t) withoutresorting toparametric methods: (1)Kaplan-Meier (2)Life-table (Actuarial Estimator) (3)viatheCumulativehazardestimator 42 (1)TheKaplan-Meier Estimator TheKaplan-Meier (orKM)estimator isprobably themostpopularapproach.Itcanbejusti¯ed fromseveralperspectives: ²productlimitestimator ²likelihoodjusti¯cation ²redistribute totherightestimator Wewillstartwithanintuitivemotivationbased onconditional probabilities, thenreviewsomeof theotherjusti¯cations. 43 Motivation: First,consider anexample wherethereisnocensoring. Thefollowingaretimesofremission (weeks)for21leukemia patientsreceiving controltreatmen t(Table1.1ofCox& Oakes): 1,1,2,2,3,4,4,5,5,8,8,8,8,11,11,12,12,15,17,22,23 Howwouldweestimate S(10),theprobabilit ythatanindi- vidualsurvivestotime10orlater? Whatabout~S(8)?Isit12 21or8 21? Let'sconstruct atableof~S(t): Valuesoft^S(t) t·121/21=1.000 1<t·219/21=0.905 2<t·317/21=0.809 3<t·4 4<t·5 5<t·8 8<t·11 11<t·12 12<t·15 15<t·17 17<t·22 22<t·23 44 EmpiricalSurvivalFunction: Whenthereisnocensoring, thegeneralformulais: ~S(t)=#individualswithT¸t totalsamplesize Example forleukemiadata(controlarm): 45 Whatifthereiscensoring? Consider thetreatedgroupfromTable1.1ofCoxandOakes: 6+;6;6;6;7;9+;10+;10;11+;13;16;17+ 19+;20+;22;23;25+;32+;32+;34+;35+ [Note:timeswith+arerightcensored] WeknowS(6)=21/21,becauseeveryonesurvivedatleast untiltime6orgreater. But,wecan'tsayS(7)=17/21, becausewedon'tknowthestatusofthepersonwhowas censored attime6. Ina1958paperintheJournaloftheAmericanStatistical Association,KaplanandMeierproposedawaytononpara- metrically estimate S(t),eveninthepresence ofcensoring. Themethodisbasedontheideasofconditionalproba- bility. 46 Aquickreviewofconditionalprobability: ConditionalProbability:SupposeAandBaretwo events.Then, P(AjB)=P(A\B) P(B) Multiplication lawofprobability:canbeobtained fromtheaboverelationship, bymultiplying bothsidesby P(B): P(A\B)=P(AjB)P(B) Extensiontomorethan2events: SupposeA1;A2:::Akarekdi®erentevents.Then,theprob- abilityofallkeventshappeningtogether canbewrittenas aproductofconditional probabilities: P(A1\A2:::\Ak)=P(AkjAk¡1\:::\A1)£ £P(Ak¡1jAk¡2\:::\A1) ::: £P(A2jA1) £P(A1) 47 Now,let'sapplytheseideastoestimateS(t): Supposeak<t·ak+1.Then S(t)=P(T¸ak+1) =P(T¸a1;T¸a2;:::;T¸ak+1) =P(T¸a1)£kY j=1P(T¸aj+1jT¸aj) =kY j=1[1¡P(T=ajjT¸aj)] =kY j=1[1¡¸j] so^S(t)»=kY j=10 B@1¡dj rj1 CA =Y j:aj<t0 B@1¡dj rj1 CA djisthenumberofdeathsataj rjisthenumberatriskataj 48 IntuitionbehindtheKaplan-Meier Estimator Thinkofdividing theobservedtimespan ofthestudyintoa seriesof¯neintervalssothatthereisaseparate intervalfor eachtimeofdeathorcensoring: D C CDDD Usingthelawofconditional probabilit y, Pr(T¸t)=Y jPr(survivej-thintervalIjjsurvivedtostartofIj) wheretheproductistakenoveralltheintervalsincluding or preceding timet. 49 4possibilities foreachinterval: (1)Noevents(deathorcensoring) -conditional prob- abilityofsurviving theintervalis1 (2)Censoring -assumetheysurvivetotheendofthein- terval,sothattheconditional probabilit yofsurviving theintervalis1 (3)Death,butnocensoring -conditional probabilit y ofnotsurviving theintervalis#deaths(d)dividedby# `atrisk'(r)atthebeginning oftheinterval.Sothecon- ditionalprobabilit yofsurviving theintervalis1¡(d=r). (4)Tieddeathsandcensoring -assumecensorings last totheendoftheinterval,sothatconditional probabilit y ofsurviving theintervalisstill1¡(d=r) GeneralFormulaforjthinterval: Itturnsoutwecanwriteageneralformulafortheconditional probabilit yofsurviving thej-thintervalthatholdsforall4 cases: 1¡dj rj 50 Wecouldusethesameapproachbygrouping theeventtimes intointervals(say,oneintervalforeachmonth),andthen countingupthenumberofdeaths(events)ineachtoesti- matetheprobabilit yofsurviving theinterval(thisiscalled thelifetableestimate ). However,theassumption thatthosecensored lastuntilthe endoftheintervalwouldn'tbequiteaccurate, sowewould endupwithacruderapproximation. Astheintervalsget¯nerand¯ner,theapproximations made inestimating theprobabilities ofgettingthrougheachinter- valbecomesmallerandsmaller, sothattheestimator con- vergestothetrueS(t). Thisintuitionclari¯eswhyanalternativ enamefortheKM istheproductlimitestimator . 51 TheKaplan-Meier estimatorofthesurvivorship function(orsurvivalprobability)S(t)=Pr(T¸t) is: ^S(t)=Q j:¿j<trj¡dj rj =Q j:¿j<t0 @1¡dj rj1 A where ²¿1;:::¿KisthesetofKdistinctdeathtimesobservedin thesample ²djisthenumberofdeathsat¿j ²rjisthenumberofindividuals \atrisk"rightbeforethe j-thdeathtime(everyonedeadorcensored atorafter thattime). ²cjisthenumberofcensored observationsbetweenthe j-thand(j+1)-stdeathtimes.Censorings tiedat¿j areincluded incj Note:twousefulformulasare: (1)rj=rj¡1¡dj¡1¡cj¡1 (2)rj=X l¸j(cl+dl) 52 CalculatingtheKM-CoxandOakesexample Makeatablewitharowforeverydeathorcensoring time: ¿jdjcjrj1¡(dj=rj)^S(¿+ j) 6312118 21=0.857 71017 90116 10 11 13 16 17 19 20 22 23 Notethat: ²^S(t+)onlychangesatdeath(failure) times ²^S(t+)is1uptothe¯rstdeathtime ²^S(t+)onlygoesto0ifthelasteventisadeath 53 KMplotfortreatedleukemiapatients Note:moststatisticalsoftwarepackagessumma- rizetheKMsurvivalfunctionat¿+ j,i.e.,justaf- terthetimeofthej-thfailure. Inotherwords,theyprovide^S(¿+ j). Whenthereisnocensoring, theempirical survivalestimate wouldthenbe: ~S(t+)=#individualswithT>t totalsamplesize 54 OutputfromSTATAKMEstimator: failure time:weeks failure/censor: remiss Beg. NetSurvivor Std. Time Total FailLostFunction Error [95%Conf.Int.] ----------------------------------------------------------- -------- 6 21 310.8571 0.0764 0.6197 0.9516 7 17 100.8067 0.0869 0.5631 0.9228 9 16 010.8067 0.0869 0.5631 0.9228 10 15 110.7529 0.0963 0.5032 0.8894 11 13 010.7529 0.0963 0.5032 0.8894 13 12 100.6902 0.1068 0.4316 0.8491 16 11 100.6275 0.1141 0.3675 0.8049 17 10 010.6275 0.1141 0.3675 0.8049 19 9010.6275 0.1141 0.3675 0.8049 20 8010.6275 0.1141 0.3675 0.8049 22 7100.5378 0.1282 0.2678 0.7468 23 6100.4482 0.1346 0.1881 0.6801 25 5010.4482 0.1346 0.1881 0.6801 32 4020.4482 0.1346 0.1881 0.6801 34 2010.4482 0.1346 0.1881 0.6801 35 1010.4482 0.1346 0.1881 0.6801 55 TwoOtherJusti¯cationsforKMEstimator I.Likelihood-basedderivation(CoxandOakes) Foradiscretefailuretimevariable,de¯ne: djnumberoffailuresataj rjnumberofindividuals atriskataj (including thosecensored ataj). ¸jPr(death) inj-thinterval (conditional onsurvivaltostartofinterval) Thelikelihoodisthatofgindependentbinomials: L(¸)=gY j=1¸dj j(1¡¸j)rj¡dj Therefore, themaximumlikelihoodestimator of¸j is: ^¸j=dj=rj NowweplugintheMLE'sof¸toestimate S(t): ^S(t)=Y j:aj<t(1¡^¸j) =Y j:aj<t0 B@1¡dj rj1 CA 56 II.Redistributetotherightjusti¯cation (Efron,1967) Intheabsence ofcensoring, ^S(t)isjusttheproportionof individuals withT¸t.TheideabehindEfron'sapproach istospreadthecontributions ofcensored observationsout overallthepossibletimestotheirright. Algorithm: ²Step(1):arrangethenobservedtimes(deathsorcensor- ings)inincreasing order.Ifthereareties,putcensored afterdeaths. ²Step(2):Assignweight(1=n)toeachtime. ²Step(3):Movingfromlefttoright,eachtimeyouen- counteracensored observation,distribute itsmasstoall timestoitsright. ²Step(4):Calculate ^Sjbysubtracting the¯nalweight fortimejfrom^Sj¡1 57 Exampleof\redistributetotheright"algorithm Consider thefollowingeventtimes: 2,2.5+,3,3,4,4.5+,5,6,7 Thealgorithm goesasfollows: (Step1) (Step4) Times Step2Step3aStep3b^S(¿j) 21/9=0.11 0.889 2.5+1/9=0.11 0 0.889 32/9=0.22 0.25 0.635 41/9=0.11 0.13 0.508 4.5+1/9=0.11 0.13 00.508 51/9=0.11 0.13 0.17 0.339 61/9=0.11 0.13 0.17 0.169 71/9=0.11 0.13 0.17 0.000 Thiscomesoutthesameastheproductlimitapproach. 58 PropertiesoftheKMestimator Inthecaseofnocensoring: ^S(t)=~S(t)=#deathsattorgreater n wherenisthenumberofindividuals inthestudy. Thisisjustlikeanestimated probabilit yfromabinomial distribution, sowehave: ^S(t)'N(S(t);S(t)[1¡S(t)]=n) Howdoescensoringa®ectthis? ²^S(t)isstillapproximately normal ²Themeanof^S(t)convergestothetrueS(t) ²Thevarianceisabitmorecomplicated (sincethede- nominatornincludes somecensored observations). Oncewegetthevariance,thenwecanconstruct (pointwise) (1¡®)%con¯dence intervals(NOTbands)about^S(t): ^S(t)§z1¡®=2se[^S(t)] 59 Greenwood'sformula(Collett2.1.3) WecanthinkoftheKMestimator as ^S(t)=Y j:¿j<t(1¡^¸j) where^¸j=dj=rj: Sincethe^¸j'sarejustbinomial proportions,wecanapply standard likelihoodtheorytoshowthateach^¸jisapproxi- matelynormal,withmeanthetrue¸j,and var(^¸j)¼^¸j(1¡^¸j) rj Also,the^¸j'sareindependentinlargeenoughsamples. Since^S(t)isafunction ofthe¸j's,wecanestimate itsvari- anceusingthedeltamethod: Deltamethod:IfYisnormalwithmean¹and variance¾2,theng(Y)isapproximately normally distributed withmeang(¹)andvariance[g0(¹)]2¾2. 60 Twospeci¯cexamplesofthedeltamethod: (A)Z=log(Y) thenZ»N2 64log(¹);0 @1 ¹1 A2 ¾23 75 (B)Z=exp(Y) thenZ»N· e¹;[e¹]2¾2¸ Theexamples aboveusethefollowingresultsfromcalculus: d dxlogu=1 u0 @du dx1 A d dxeu=eu0 @du dx1 A 61 Greenwood'sformula(continued) Insteadofdealingwith^S(t)directly,wewilllookatitslog: log[^S(t)]=X j:¿j<tlog(1¡^¸j) Thus,byapproximateindependence ofthe^¸j's, var(log[^S(t)])=X j:¿j<tvar[log(1¡^¸j)] by(A) =X j:¿j<t0 B@1 1¡^¸j1 CA2 var(^¸j) =X j:¿j<t0 B@1 1¡^¸j1 CA2 ^¸j(1¡^¸j)=rj =X j:¿j<t^¸j (1¡^¸j)rj =X j:¿j<tdj (rj¡dj)rj Now,^S(t)=exp[log[^S(t)]].Thusby(B), var(^S(t))=[^S(t)]2var· log[^S(t)]¸ Greenwood'sFormula: var(^S(t))=[^S(t)]2P j:¿j<tdj (rj¡dj)rj 62 Backtocon¯denceintervals Fora95%con¯dence interval,wecoulduse ^S(t)§z1¡®=2se[^S(t)] wherese[^S(t)]iscalculated usingGreenwood'sformula. Problem: Thisapproachcanyieldvalues>1or<0. Betterapproach:Geta95%con¯dence intervalfor L(t)=log(¡log(S(t))) Sincethisquantityisunrestricted, thecon¯dence interval willbeintheproperrangewhenwetransform back. Toseewhythisworks,notethefollowing: ²Since^S(t)isanestimated probabilit y 0·^S(t)·1 ²Takingthelogof^S(t)hasbounds: ¡1·log[^S(t)]·0 ²Takingtheopposite: 0·¡log[^S(t)]·1 ²Takingthelogagain: ¡1·log· ¡log[^S(t)]¸ ·1 Totransform back,reversestepswithS(t)=exp(¡exp(L(t)) 63 Log-logApproachforCon¯denceIntervals: (1)De¯neL(t)=log(¡log(S(t))) (2)Forma95%con¯dence intervalforL(t)basedon^L(t), yielding[^L(t)¡A;^L(t)+A] (3)SinceS(t)=exp(¡exp(L(t)),thecon¯dence bounds forthe95%CIonS(t)are: [exp(¡e(^L(t)+A));exp(¡e(^L(t)¡A))] (notethattheupperandlowerboundsswitch) (4)Substituting ^L(t)=log(¡log(^S(t)))backintotheabove bounds,wegetcon¯dence boundsof ([^S(t)]eA;[^S(t)]e¡A) 64 WhatisA? ²Ais1:96se(^L(t)) ²Tocalculate this,weneedtocalculate var(^L(t))=var· log(¡log(^S(t)))¸ ²Fromourprevious calculations, weknow var(log[^S(t)])=X j:¿j<tdj (rj¡dj)rj ²Applying thedeltamethodasinexample (A),weget: var(^L(t))=var(log(¡log[^S(t)])) =1 [log^S(t)]2X j:¿j<tdj (rj¡dj)rj ²Wetakethesquarerootoftheabovetogetse(^L(t)), andthenformthecon¯dence intervalsas: ^S(t)e§1:96se(^L(t)) ²ThisistheapproachthatStatauses.Splusgivesanop- tiontocalculate thesebounds(useconf.type=''log-log'' insurv.fit ). 65 SummaryofCon¯denceIntervalsonS(t) ²Calculate ^S(t)§1:96se[^S(t)]wherese[^S(t)]iscalcu- latedusingGreenwood'sformula,andreplacenegative lowerboundsby0andupperboundsgreaterthan1by 1. {Recommended byCollett {ThisisthedefaultusingSAS {notverysatisfactory ²Usealogtransformation tostabilize thevarianceand allowfornon-symmetric con¯dence intervals.Thisis whatisnormally doneforthecon¯dence intervalofan estimated oddsratio. {Usevar[log(^S(t))]=P j:¿j<tdj (rj¡dj)rjalreadycalcu- latedaspartofGreenwood'sformula {ThisisthedefaultinSplus ²Usethelog-logtransformation justdescribed {Somewhat complicated, butalwaysyieldsproperbounds {ThisisthedefaultinStata. 66 SoftwareforKaplan-Meier Curves ²Stata-stsetandstscommands ²SAS-proclifetest ²Splus-surv.¯t(time,status) DefaultsforCon¯denceIntervalCalculations ²Stata-\log-log")^L(t)§1:96se[^L(t)] whereL(t)=log[¡log(S(t))] ²SAS-\plain")^S(t)§1:96se[^S(t)] ²Splus-\log")logS(t)§1:96se[log(^S(t))] butSpluswillalsogiveeitheroftheothertwooptionsif yourequestthem. 67 StataCommands Createa¯lecalled\leukemia.dat" withtherawdata,with acolumnfortreatmen t,weekstorelapse(i.e.,duration of remission), andrelapsestatus: .infile trtremissstatususingleukemia.dat .stsetremissstatus (setsupafailure timedataset, withfailtime statusinthatorder, typehelpstsettogetdetails) .stslist (estimated S(t),se[S(t)], and95%CI) .stsgraph,saving(kmtrt) (creates aKaplan-Meier plot,and savestheplotinfilekmtrt.gph, type``helpgphdot'' togetsome printing instructions) .graphusingkmtrt (redisplays thegraphatanylatertime) IfthedatasethasalreadybeencreatedandloadedintoStata, thenyoucansubstitute thefollowingcommands forinitial- izingthedata: .useleukem (findsStatadataset leukem.dta) .describe (provides adescription ofthedataset) .stsetremissstatus (declares datatobefailure type) .stdes (givesadescription ofthesurvival dataset) 68 STATAOutputforTreatedLeukemiaPatients: .useleukem .stsetremissstatusiftrt==1 .stslist failure time:remiss failure/censor: status Beg. NetSurvivor Std. Time Total FailLostFunction Error [95%Conf.Int.] ----------------------------------------------------------- -------- 6 21 310.8571 0.0764 0.6197 0.9516 7 17 100.8067 0.0869 0.5631 0.9228 9 16 010.8067 0.0869 0.5631 0.9228 10 15 110.7529 0.0963 0.5032 0.8894 11 13 010.7529 0.0963 0.5032 0.8894 13 12 100.6902 0.1068 0.4316 0.8491 16 11 100.6275 0.1141 0.3675 0.8049 17 10 010.6275 0.1141 0.3675 0.8049 19 9010.6275 0.1141 0.3675 0.8049 20 8010.6275 0.1141 0.3675 0.8049 22 7100.5378 0.1282 0.2678 0.7468 23 6100.4482 0.1346 0.1881 0.6801 25 5010.4482 0.1346 0.1881 0.6801 32 4020.4482 0.1346 0.1881 0.6801 34 2010.4482 0.1346 0.1881 0.6801 35 1010.4482 0.1346 0.1881 0.6801 69 SASCommandsforKaplanMeierEstimator- PROCLIFETEST TheSAScommand fortheKaplan-Meier estimate is: timefailtime*censor(1); or timefailtime*failind(0); The¯rstvariableisthefailuretime,andthesecondisthe failureorcensoring indicator. Inparenthesesyouneedtoput thespeci¯cnumericvaluethatcorrespondstocensoring. Theupperandlowercon¯dence limitson^S(t)areincluded inthedataset\OUTSUR V"whenspeci¯ed.Theupperand lowerlimitsarecalled:sdfucl,sdflcl. dataleukemia; inputweeksremiss; labelweeks='Time toRemission (inweeks)' remiss='Remission indicator (1=yes,0=no)'; cards; 61 61 ........... (lineseditedouthere) 340 350 ; proclifetest data=leukemia outsurv=confint; timeweeks*remiss(0); title'Leukemia datafromTable1.1ofCoxandOakes'; run; procprintdata=confint; title'95%Confidence Intervals forEstimated Survival'; 70 OutputfromSASProcLifetest Note:thisinformation isnotprintedifyouuseNOPRINT. Leukemia datafromTable1.1ofCoxandOakes TheLIFETEST Procedure Product-Limit Survival Estimates Survival Standard Number Number WEEKS Survival Failure Error Failed Left 0.0000 1.0000 0 0 0 21 6.0000 . . . 1 20 6.0000 . . . 2 19 6.0000 0.8571 0.1429 0.0764 3 18 6.0000* . . . 3 17 7.0000 0.8067 0.1933 0.0869 4 16 9.0000* . . . 4 15 10.0000 0.7529 0.2471 0.0963 5 14 10.0000* . . . 5 13 11.0000* . . . 5 12 13.0000 0.6902 0.3098 0.1068 6 11 16.0000 0.6275 0.3725 0.1141 7 10 17.0000* . . . 7 9 19.0000* . . . 7 8 20.0000* . . . 7 7 22.0000 0.5378 0.4622 0.1282 8 6 23.0000 0.4482 0.5518 0.1346 9 5 25.0000* . . . 9 4 32.0000* . . . 9 3 32.0000* . . . 9 2 34.0000* . . . 9 1 35.0000* . . . 9 0 *Censored Observation 71 OutputfromprintingtheCONFINT¯le 95%Confidence Intervals forEstimated Survival OBS WEEKS _CENSOR_ SURVIVAL SDF_LCL SDF_UCL 1 0 0 1.00000 1.00000 1.00000 2 6 0 0.85714 0.70748 1.00000 3 6 1 0.85714 . . 4 7 0 0.80672 0.63633 0.97711 5 9 1 0.80672 . . 6 10 0 0.75294 0.56410 0.94178 7 10 1 0.75294 . . 8 11 1 0.75294 . . 9 13 0 0.69020 0.48084 0.89955 10 16 0 0.62745 0.40391 0.85099 11 17 1 0.62745 . . 12 19 1 0.62745 . . 13 20 1 0.62745 . . 14 22 0 0.53782 0.28648 0.78915 15 23 0 0.44818 0.18439 0.71197 16 25 1 . . . 17 32 1 . . . 18 32 1 . . . 19 34 1 . . . 20 35 1 . . . Theoutputdatasetwillhaveoneobservationforeachunique combination ofweeksandcensor .Itwillalsoaddan observationforfailuretimeequalto0. 72 SplusCommands Createa¯lecalled\leukemia.dat" withthevariablesnames inthe¯rstrow,asfollows: tc 61 61 etc... InSplus,type y_read.table('leukemia.dat',header=T) surv.fit(y$t,y$c) plot(surv.fit(y$t,y$c)) (theplotcommand willalsoyield95%con¯dence intervals) Tospecifythetypeofcon¯dence intervals,usetheconf.type= optioninthesurv.¯tstatemen ts:e.g.conf.type=\log-log" orconf.type=\plain" 73 >surv.fit(y$t,y$c) 95percent confidence interval isoftype"log" timen.riskn.event survival std.dev lower95%CIupper95%CI 621 30.8571429 0.07636035 0.7198171 1.0000000 717 10.8067227 0.08693529 0.6531242 0.9964437 1015 10.7529412 0.09634965 0.5859190 0.9675748 1312 10.6901961 0.10681471 0.5096131 0.9347692 1611 10.6274510 0.11405387 0.4393939 0.8959949 22 7 10.5378151 0.12823375 0.3370366 0.8582008 23 6 10.4481793 0.13459146 0.2487882 0.8073720 >surv.fit(y$t,y$c,conf.type="log-log") 95percent confidence interval isoftype"log-log" timen.riskn.event survival std.dev lower95%CIupper95%CI 621 30.8571429 0.07636035 0.6197180 0.9515517 717 10.8067227 0.08693529 0.5631466 0.9228090 1015 10.7529412 0.09634965 0.5031995 0.8893618 1312 10.6901961 0.10681471 0.4316102 0.8490660 1611 10.6274510 0.11405387 0.3675109 0.8049122 22 7 10.5378151 0.12823375 0.2677789 0.7467907 23 6 10.4481793 0.13459146 0.1880520 0.6801426 >surv.fit(y$t,y$c,conf.type="plain") 95percent confidence interval isoftype"plain" timen.riskn.event survival std.dev lower95%CIupper95%CI 621 30.8571429 0.07636035 0.7074793 1.0000000 717 10.8067227 0.08693529 0.6363327 0.9771127 1015 10.7529412 0.09634965 0.5640993 0.9417830 1312 10.6901961 0.10681471 0.4808431 0.8995491 1611 10.6274510 0.11405387 0.4039095 0.8509924 22 7 10.5378151 0.12823375 0.2864816 0.7891487 23 6 10.4481793 0.13459146 0.1843849 0.7119737 74 KMSurvivalEstimateandCon¯denceintervals (SPlus) TimeSurvival 0 5 10 15 20 25 30 350.0 0.2 0.4 0.6 0.8 1.0 75 Means,Medians,QuantilesbasedontheKM ²Mean:Pk j=1¿jPr(T=¿j) ²Median -byde¯nition, thisisthetime,¿,suchthat S(¿)=0:5.However,inpractice, itisde¯nedasthe smallest timesuchthat^S(¿)·0:5.Themedianismore appropriate forcensored survivaldatathanthemean. Forthetreatedleukemiapatients,we¯nd: ^S(22)=0:5378 ^S(23)=0:4482 Themedianisthus23.Thiscanalsobeseenvisuallyon thegraphtotheleft. ²Lowerquartile(25thpercentile): thesmallest time(LQ)suchthat^S(LQ)·0:75 ²Upperquartile(75thpercentile): thesmallest time(UQ)suchthat^S(UQ)·0:25 76 The(2)LifetableEstimatorofSurvival: Wesaidthatwewouldconsider thefollowingthreemethods forestimating asurvivorshipfunction S(t)=Pr(T¸t) withoutresorting toparametric methods: (1)pKaplan-Meier (2)=)Life-table (Actuarial Estimator) (3)=)Cumulativehazardestimator 77 (2)TheLifetableorActuarialEstimator ²oneoftheoldesttechniquesaround ²usedbyactuaries, demographers, etc. ²applieswhenthedataaregrouped Ourgoalisstilltoestimate thesurvivalfunction, hazard,and densityfunction, butthisiscomplicated bythefactthatwe don'tknowexactlywhenduringeachtimeintervalanevent occurs. 78 Lee(section 4.2)providesagooddescription oflifetable methods,anddistinguishes severaltypesaccording tothe datasources: Popula tionLifeTables ²cohortlifetable-describesthemortalityexperience frombirthtodeathforaparticular cohortofpeopleborn ataboutthesametime.Peopleatriskatthestartofthe intervalarethosewhosurvivedtheprevious interval. ²currentlifetable-constructed from(1)censusinfor- mationonthenumberofindividuals aliveateachage, foragivenyearand(2)vitalstatistics onthenumber ofdeathsorfailuresinagivenyear,byage.Thistype oflifetable isoftenreportedintermsofahypothetical cohortof100,000people. Generally ,censoring isnotanissueforPopulation LifeTa- bles. Clinical Lifetables-appliestogroupedsurvivaldata fromstudiesinpatientswithspeci¯cdiseases. Because pa- tientscanenterthestudyatdi®erenttimes,orbelostto follow-up,censoring mustbeallowed. 79 Notation ²thej-thtimeintervalis[tj¡1;tj) ²cj-thenumberofcensorings inthej-thinterval ²dj-thenumberoffailuresinthej-thinterval ²rjisthenumberenteringtheinterval Example :2418MaleswithAnginaPectoris(Lee,p.91) Yearafter Diagnosis jdjcjrjr0 j=rj¡cj=2 [0;1)1456024182418.0 [1;2)22263919621942.5 (1962-39 2) [2;3)31522216971686.0 [3;4)41712315231511.5 [4;5)51352413291317.0 [5;6)612510711701116.5 [6;7)783133938871.5 etc.. 80 Estimatingthesurvivorshipfunction WecouldapplytheK-Mformuladirectlytothenumbersin thetableontheprevious page,estimating S(t)as ^S(t)=Y j:¿j<t0 B@1¡dj rj1 CA However,thisapproachisunsatisfactory forgroupeddata.... ittreatstheproblem asthoughitwereindiscretetime,with eventshappeningonlyat1yr,2yr,etc.Infact,whatwe aretryingtocalculate hereistheconditional probabilit yof dyingwithintheinterval,givensurvivaltothebeginning of it. Whatshouldwedowiththecensoredpeople? Wecanassumethatcensoringsoccur: ²atthebeginning ofeachinterval:r0 j=rj¡cj ²attheendofeachinterval:r0 j=rj ²onaveragehalfwaythrough theinterval: r0 j=rj¡cj=2 Thelastassumption yieldstheActuarial Estimator. Itis appropriate ifcensorings occuruniformly throughout thein- terval. 81 Constructingthelifetable First,someadditional notation forthej-thinterval,[tj¡1;tj): ²Midpoint(tmj)-usefulforplotting thedensityand thehazardfunction ²Width(bj=tj¡tj¡1)neededforcalculating thehazard inthej-thinterval Quantitiesestimated: ²Conditional probabilit yofdying ^qj=dj=r0 j ²Conditional probabilit yofsurviving ^pj=1¡^qj ²Cumulativeprobabilit yofsurviving attj: ^S(tj)=Y `·j^p` =Y `·j0 @1¡d` r`01 A 82 Someimportantpointstonote: ²Because theintervalsarede¯nedas[tj¡1;tj),the¯rst intervaltypicallystartswitht0=0. ²Stataestimates thesurvivalfunction attheright-hand endpointofeachinterval,i.e.,S(tj) ²However,SASestimates thesurvivalfunction attheleft- handendpoint,S(tj¡1). ²Theimplication inSASisthat^S(t0)=1and^S(t1)=p1 83 Otherquantitiesestimatedatthe midpointofthej-thinterval: ²Hazard inthej-thinterval: ^¸(tmj)=dj bj(r0j¡dj=2) =^qj bj(1¡^qj=2) thenumberofdeathsintheintervaldividedbytheav- eragenumberofsurvivorsatthemidpoint ²densityatthemidpointofthej-thinterval: ^f(tmj)=^S(tj¡1)¡^S(tj) bj =^S(tj¡1)^qj bj Note:Another waytogetthisis: ^f(tmj)=^¸(tmj)^S(tmj) =^¸(tmj)[^S(tj)+^S(tj¡1)]=2 84 ConstructingtheLifetableusingStata Usestheltablecommand. Iftherawdataarealreadygrouped,thenthefreqstatemen t mustbeusedwhenreadingthedata. .infileyearsstatuscountusingangina.dat (32observations read) .ltableyearsstatus[freq=count] Beg. Std. Interval TotalDeaths LostSurvival Error [95%Conf.Int.] ------------------------------------------------------------------- ------ 012418 456 00.8114 0.0080 0.7952 0.8264 121962 226390.7170 0.0092 0.6986 0.7346 231697 152220.6524 0.0097 0.6329 0.6711 341523 171230.5786 0.0101 0.5584 0.5981 451329 135240.5193 0.0103 0.4989 0.5392 561170 1251070.4611 0.0104 0.4407 0.4813 67938 831330.4172 0.0105 0.3967 0.4376 78722 741020.3712 0.0106 0.3505 0.3919 89546 51680.3342 0.0107 0.3133 0.3553 910427 42640.2987 0.0109 0.2775 0.3201 1011321 43450.2557 0.0111 0.2341 0.2777 1112233 34530.2136 0.0114 0.1917 0.2363 1213146 18330.1839 0.0118 0.1614 0.2075 1314 95 9270.1636 0.0123 0.1404 0.1884 1415 59 6230.1429 0.0133 0.1180 0.1701 1516 30 0300.1429 0.0133 0.1180 0.1701 ------------------------------------------------------------------- -------- ---- 85 Itisalsopossibletogetestimates ofthehazardfunction, ^¸j, anditsstandard errorusingthe\hazard"option: .ltableyearsstatus[freq=count], hazard Beg. Cum. Std. Std. Interval Total Failure Error Hazard Error [95%ConfInt] ------------------------------------------------------------------- ------- 012418 0.1886 0.0080 0.2082 0.0097 0.1892 0.2272 121962 0.2830 0.0092 0.1235 0.0082 0.1075 0.1396 231697 0.3476 0.0097 0.0944 0.0076 0.0794 0.1094 341523 0.4214 0.0101 0.1199 0.0092 0.1020 0.1379 451329 0.4807 0.0103 0.1080 0.0093 0.0898 0.1262 561170 0.5389 0.0104 0.1186 0.0106 0.0978 0.1393 679380.5828 0.0105 0.1000 0.0110 0.0785 0.1215 787220.6288 0.0106 0.1167 0.0135 0.0902 0.1433 895460.6658 0.0107 0.1048 0.0147 0.0761 0.1336 9104270.7013 0.0109 0.1123 0.0173 0.0784 0.1462 10113210.7443 0.0111 0.1552 0.0236 0.1090 0.2015 11122330.7864 0.0114 0.1794 0.0306 0.1194 0.2395 12131460.8161 0.0118 0.1494 0.0351 0.0806 0.2182 1314 950.8364 0.0123 0.1169 0.0389 0.0407 0.1931 1415 590.8571 0.0133 0.1348 0.0549 0.0272 0.2425 1516 300.8571 0.0133 0.0000 . . . ------------------------------------------------------------------- ------ Thereisalsoa\failure "optionwhichgivesthenumberof failures(likethedefault), andalsoprovidesa95%con¯dence intervalonthecumulativefailureprobabilit y. 86 ConstructingthelifetableusingSAS Iftherawdataarealreadygrouped,thentheFREQstate- mentmustbeusedwhenreadingthedata. SASrequires thattheintervalendpointsbespeci¯ed,using oneofthefollowing(seeSASmanualoronlinehelpformore detail): ²intervals-specifythetheintervalendpoints ²width-specifythewidthofeachinterval ²ninterval-specifythenumberofintervals Title'Actuarial Estimator forAnginaPectoris Example'; dataangina; inputyearsstatuscount; cards; 0.51456 1.51226 2.51152 /*anginacases*/ 3.51171 4.51135 5.51125 . . 0.500 1.5039 2.5022 /*censored */ 3.5023 4.5024 5.50107 . . proclifetest data=angina outsurv=survres intervals=0 to15by1method=act; timeyears*status(0); freqcount; 87 SASoutput: Actuarial Estimator forAnginaPectoris Example TheLIFETEST Procedure LifeTableSurvival Estimates Conditional Effective Conditional Probability Interval Number Number Sample Probability Standard [Lower, Upper) Failed Censored Size ofFailure Error 0 1456 02418.0 0.1886 0.00796 1 2226 39 1942.5 0.1163 0.00728 2 3152 22 1686.0 0.0902 0.00698 3 4171 23 1511.5 0.1131 0.00815 4 5135 24 1317.0 0.1025 0.00836 5 6125 107 1116.5 0.1120 0.00944 6 783 133 871.5 0.0952 0.00994 7 874 102 671.0 0.1103 0.0121 8 951 68 512.0 0.0996 0.0132 9 10 42 64 395.0 0.1063 0.0155 10 11 43 45 298.5 0.1441 0.0203 11 12 34 53 206.5 0.1646 0.0258 12 13 18 33 129.5 0.1390 0.0304 13 14 9 27 81.5 0.1104 0.0347 14 15 6 23 47.5 0.1263 0.0482 15 . 0 30 15.0 0 0 Survival Median Median Interval Standard Residual Standard [Lower, Upper) Survival Failure Error Lifetime Error 0 11.0000 0 05.3313 0.1749 1 20.8114 0.1886 0.00796 6.2499 0.2001 2 30.7170 0.2830 0.00918 6.3432 0.2361 3 40.6524 0.3476 0.00973 6.2262 0.2361 4 50.5786 0.4214 0.0101 6.2185 0.1853 5 60.5193 0.4807 0.0103 5.9077 0.1806 6 70.4611 0.5389 0.0104 5.5962 0.1855 7 80.4172 0.5828 0.0105 5.1671 0.2713 8 90.3712 0.6288 0.0106 4.9421 0.2763 9 100.3342 0.6658 0.0107 4.8258 0.4141 10 110.2987 0.7013 0.0109 4.6888 0.4183 11 120.2557 0.7443 0.0111 . . 12 130.2136 0.7864 0.0114 . . 13 140.1839 0.8161 0.0118 . . 14 150.1636 0.8364 0.0123 . . 15 .0.1429 0.8571 0.0133 . . 88 moreSASoutput: (estimated density^fjandhazard^¸j) Evaluated attheMidpoint oftheInterval PDF Hazard Interval Standard Standard [Lower, Upper) PDF Error Hazard Error 0 10.1886 0.00796 0.208219 0.009698 1 20.0944 0.00598 0.123531 0.008201 2 30.0646 0.00507 0.09441 0.007649 3 40.0738 0.00543 0.119916 0.009154 4 50.0593 0.00495 0.108043 0.009285 5 60.0581 0.00503 0.118596 0.010589 6 70.0439 0.00469 0.10.010963 7 80.0460 0.00518 0.116719 0.013545 8 90.0370 0.00502 0.10483 0.014659 9 100.0355 0.00531 0.112299 0.017301 10 110.0430 0.00627 0.155235 0.023602 11 120.0421 0.00685 0.17942 0.030646 12 130.0297 0.00668 0.149378 0.03511 13 140.0203 0.00651 0.116883 0.038894 14 150.0207 0.00804 0.134831 0.054919 15 . . . . . Summary oftheNumberofCensored andUncensored Values Total Failed Censored %Censored 2418 1625 79332.7957 89 Supposewewishtousetheactuarial method,butthedata donotcomegrouped. Consider thetreatednursinghomepatients,withlengthof stay(los)groupedinto100dayintervals: .usenurshome .dropifrx==0 (keeponlythetreated patients) (881observations deleted) .stsetlosfail .ltable losfail,intervals(100) Beg. Std. Interval TotalDeaths LostSurvival Error [95%Conf.Int.] ------------------------------------------------------------------- ----- 0100 710328 00.5380 0.0187 0.5006 0.5739 100200 382 86 00.4169 0.0185 0.3805 0.4529 200300 296 65 00.3254 0.0176 0.2911 0.3600 300400 231 38 00.2718 0.0167 0.2396 0.3050 400500 193 32 10.2266 0.0157 0.1966 0.2581 500600 160 13 00.2082 0.0152 0.1792 0.2388 600700 147 13 00.1898 0.0147 0.1619 0.2195 700800 134 10300.1739 0.0143 0.1468 0.2029 800900 94 4290.1651 0.0143 0.1383 0.1941 9001000 61 4300.1508 0.0147 0.1233 0.1808 10001100 27 0270.1508 0.0147 0.1233 0.1808 ------------------------------------------------------------------- ------ 90 SASCommandsforlifetableanalysis-grouping data Title'Actuarial Estimator fornursing homedata'; datamorris; infile'ch12.dat' ; inputlosagetrtgendermarstat hltstat cens; datamorristr; setmorris; iftrt=1; proclifetest data=morristr outsurv=survres intervals=0 to1100by100method=act; timelos*cens(1); run; procprintdata=survres; run; 91 Actuarial estimator fortreatednursinghomepatients Actuarial Estimator forNursing HomePatients TheLIFETEST Procedure LifeTableSurvival Estimates Effective Conditional Interval Number Number Sample Probability [Lower, Upper) Failed Censored Size ofFailure 0100 330 0 712.0 0.4635 100 200 86 0 382.0 0.2251 200 300 65 0 296.0 0.2196 300 400 38 0 231.0 0.1645 400 500 32 1 192.5 0.1662 500 600 13 0 160.0 0.0813 600 700 13 0 147.0 0.0884 700 800 10 30 119.0 0.0840 800 900 4 29 79.5 0.0503 900 1000 4 30 46.0 0.0870 1000 1100 0 27 13.5 0 Conditional Probability Survival Median Interval Standard Standard Residual [Lower, Upper) Error Survival Failure Error Lifetime 0100 0.0187 1.0000 0 0130.2 100 200 0.0214 0.5365 0.4635 0.0187 306.2 200 300 0.0241 0.4157 0.5843 0.0185 398.8 300 400 0.0244 0.3244 0.6756 0.0175 617.0 400 500 0.0268 0.2711 0.7289 0.0167 . 500 600 0.0216 0.2260 0.7740 0.0157 . 600 700 0.0234 0.2076 0.7924 0.0152 . 700 800 0.0254 0.1893 0.8107 0.0147 . 800 900 0.0245 0.1734 0.8266 0.0143 . 900 1000 0.0415 0.1647 0.8353 0.0142 . 1000 1100 00.1503 0.8497 0.0147 . 92 Actuarial estimator fortreatednursinghomepatients,cont'd Evaluated attheMidpoint oftheInterval Median PDF Hazard Interval Standard Standard Standard [Lower, Upper) Error PDF Error Hazard Error 010015.5136 0.00463 0.000187 0.006033 0.000317 100 20030.4597 0.00121 0.000122 0.002537 0.000271 200 30065.7947 0.000913 0.000108 0.002467 0.000304 300 40074.5466 0.000534 0.000084 0.001792 0.00029 400 500 .0.000451 0.000078 0.001813 0.000319 500 600 .0.000184 0.00005 0.000847 0.000235 600 700 .0.000184 0.00005 0.000925 0.000256 700 800 .0.000159 0.00005 0.000877 0.000277 800 900 .0.000087 0.000043 0.000516 0.000258 900 1000 .0.000143 0.00007 0.000909 0.000454 1000 1100 . 0 . 0 . Summary oftheNumberofCensored andUncensored Values Total Failed Censored %Censored 712 595 11716.4326 93 Actuarial estimator fortreatednursinghomepatients,cont'd OutputfromSURVRESdataset Actuarial Estimator forNursing HomePatients OBS LOSSURVIVAL SDF_LCL SDF_UCL MIDPOINT PDF 1 01.00000 1.00000 1.00000 50 .0046348 2100 0.53652 0.49989 0.57315 150 .0012079 3200 0.41573 0.37953 0.45193 250 .0009129 4300 0.32444 0.29005 0.35883 350 .0005337 5400 0.27107 0.23842 0.30372 450 .0004506 6500 0.22601 0.19528 0.25674 550 .0001836 7600 0.20764 0.17783 0.23745 650 .0001836 8700 0.18928 0.16048 0.21808 750 .0001591 9800 0.17337 0.14536 0.20139 850 .0000872 10900 0.16465 0.13677 0.19253 950 .0001432 111000 0.15033 0.12157 0.17910 1050 .0000000 OBS PDF_LCL PDF_UCL HAZARD HAZ_LCL HAZ_UCL 1.0042685 .0050011 .0060329 .0054123 .0066535 2.0009685 .0014472 .0025369 .0020050 .0030687 3.0007014 .0011245 .0024668 .0018717 .0030619 4.0003686 .0006988 .0017925 .0012248 .0023601 5.0002981 .0006031 .0018130 .0011874 .0024386 6.0000847 .0002825 .0008469 .0003869 .0013069 7.0000847 .0002825 .0009253 .0004228 .0014277 8.0000617 .0002565 .0008772 .0003340 .0014203 9.0000027 .0001717 .0005161 .0000105 .0010218 10.0000069 .0002794 .0009091 .0000191 .0017991 11. . .0000000 . . 94 ExamplesforNursinghomedata: EstimatedSurvival: Estimated Survival 0.00.10.20.30.40.50.60.70.80.91.0 Lower Limit of Time Interval01002003004005006007008009001000 95 Estimatedhazard: Estimated hazard 0.0000.0020.0040.0060.0080.010 Lower Limit of Time Interval01002003004005006007008009001000 96 (3)Estimatingthecumulativehazard (Nelson-Aalen estimator) Supposewewanttoestimate ¤(t)=Rt 0¸(u)du,thecumula- tivehazardattimet. JustaswedidfortheKM,thinkofdividing theobserved timespan ofthestudyintoaseriesof¯neintervalssothat thereisonlyoneeventperinterval: D C CDDD ¤(t)canthenbeapproximated byasum: ^¤(t)=X j¸j¢ wherethesumisoverintervals,¸jisthevalueofthehazard inthej-thintervaland¢isthewidthofeachinterval.Since ^¸¢isapproximately theprobabilit yofdyingintheinterval, wecanfurtherapproximateby ^¤(t)=X jdj=rj Itfollowsthat¤(t)willchangeonlyatdeathtimes,and hencewewritetheNelson-Aalen estimator as: ^¤NA(t)=X j:¿j<tdj=rj 97 D C CDDD rjnnnn- 1n- 1n-2n-2n-3n-4 dj001000011 cj000010100 ^¸(tj)001/n00001 n¡31 n¡4 ^¤(tj)001/n1/n1/n1/n1/n Oncewehave^¤NA(t),wecanalso¯ndanotherestimator of S(t)(Fleming-Harrington): ^SFH(t)=exp(¡^¤NA(t)) Ingeneral, thisestimator ofthesurvivalfunction willbe closetotheKaplan-Meier estimator, ^SKM(t) Wecanalsogotheotherway...wecantaketheKaplan- Meierestimate ofS(t),anduseittocalculate analternativ e estimate ofthecumulativehazardfunction: ^¤KM(t)=¡log^SKM(t) 98 StatacommandsforFHSurvivalEstimate SaywewanttoobtaintheFleming-Harrington estimate of thesurvivalfunction formarried females, inthehealthiest initialsubgroup, whoarerandomized totheuntreatedgroup ofthenursinghomestudy. First,weusethefollowingcommands tocalculate theNelson- Aalencumulativehazardestimator: .usenurshome .keepifrx==0&gender==0 &health==2 &married==1 (1579observations deleted) .stslist,na failure _d:fail analysis time_t:los Beg. NetNelson-Aalen Std. Time Total FailLost Cum.Haz. Error [95%Conf.Int.] ------------------------------------------------------------------- --- 14 12 100.0833 0.0833 0.0117 0.5916 24 11 100.1742 0.1233 0.0435 0.6976 25 10 100.2742 0.1588 0.0882 0.8530 38 9100.3854 0.1938 0.1438 1.0326 64 8100.5104 0.2306 0.2105 1.2374 89 7100.6532 0.2713 0.2894 1.4742 113 6100.8199 0.3184 0.3830 1.7551 123 5101.0199 0.3760 0.4952 2.1006 149 4101.2699 0.4515 0.6326 2.5493 168 3101.6032 0.5612 0.8073 3.1840 185 2102.1032 0.7516 1.0439 4.2373 234 1103.1032 1.2510 1.4082 6.8384 ------------------------------------------------------------------- --- 99 Aftergenerating theNelson-Aalen estimator, wemanually havetocreateavariableforthesurvivalestimate: .stsgennelson=na .gensfh=exp(-nelson) .listsfh sfh 1..9200444 2..8400932 3..7601478 4..6802101 5..6002833 6..5203723 7..4404857 8..3606392 9..2808661 10..2012493 11..1220639 12..0449048 Additional built-infunctions canbeusedtogenerate 95% con¯dence intervalsontheFHsurvivalestimate. 100 Wecancompare theFleming-Harrington survivalestimate totheKMestimate byrerunning thestslistcommand: .stslist .stsgenskm=s .listskmsfh skm sfh 1..91666667 .9200444 2..83333333 .8400932 3. .75.7601478 4..66666667 .6802101 5..58333333 .6002833 6. .5.5203723 7..41666667 .4404857 8..33333333 .3606392 9. .25.2808661 10..16666667 .2012493 11..08333333 .1220639 12. 0.0449048 Inthisexample, itlooksliketheFleming-Harrington estima- torisslightlyhigherthantheKMateverytimepoint,but withlargerdatasets thetwowilltypicallybemuchcloser. 101 SplusCommandsforFleming-Harrington Esti- mator: (Nursing homedata:females, untreated,married, healthy) Fleming-Harrington: >fh<-surv.fit(los,cens,type="f",conf.type="log-log") >fh 95percent confidence interval isoftype"log-log" timen.riskn.event survival std.dev lower95%CIupper95%CI 1412 10.9200444 0.08007959 0.5244209125 0.9892988 2411 10.8400932 0.10845557 0.4750041174 0.9600371 2510 10.7601478 0.12669130 0.4055610500 0.9200425 38 9 10.6802101 0.13884731 0.3367907188 0.8724502 64 8 10.6002833 0.14645413 0.2718422278 0.8187596 89 7 10.5203723 0.15021856 0.2115701242 0.7597900 113 6 10.4404857 0.15045450 0.1564397006 0.6960354 123 5 10.3606392 0.14723033 0.1069925657 0.6278888 149 4 10.2808661 0.14043303 0.0640979523 0.5560134 168 3 10.2012493 0.12990589 0.0293208029 0.4827590 185 2 10.1220639 0.11686728 0.0058990525 0.4224087 234 1 10.0449048 0.06216787 0.0005874321 0.2740658 Kaplan-Meier: >km<-surv.fit(los,cens,conf.type="log-log") >km 95percent confidence interval isoftype"log-log" timen.riskn.event survival std.dev lower95%CIupper95%CI 1412 10.91666667 0.07978559 0.538977181 0.9878256 2411 10.83333333 0.10758287 0.481714942 0.9555094 2510 10.75000000 0.12500000 0.408415913 0.9117204 38 9 10.66666667 0.13608276 0.337018933 0.8597118 64 8 10.58333333 0.14231876 0.270138924 0.8009402 89 7 10.50000000 0.14433757 0.208477143 0.7360731 113 6 10.41666667 0.14231876 0.152471264 0.6653015 123 5 10.33333333 0.13608276 0.102703980 0.5884189 149 4 10.25000000 0.12500000 0.060144556 0.5047588 168 3 10.16666667 0.10758287 0.026510427 0.4129803 185 2 10.08333333 0.07978559 0.005052835 0.3110704 234 1 10.00000000 NA NA NA 102 Comparison ofSurvivalCurves Wespentthelastclasslookingatsomenonparametric ap- proachesforestimating thesurvivalfunction, ^S(t),overtime forasinglesampleofindividuals. Nowwewanttocompare thesurvivalestimates betweentwo groups. Example:Timetoremissionofleukemiapatients 103 Howcanweformabasisforcomparison? Ataspeci¯cpointintime,wecouldseewhether thecon¯- denceintervalsforthesurvivalcurvesoverlap. However,thecon¯dence intervalswehavebeencalculating are\pointwise")theycorrespondtoacon¯dence inter- valfor^S(t¤)atasinglepointintime,t¤. Inotherwords,wecan'tsaythatthetruesurvivalfunction S(t)iscontainedbetweenthepointwisecon¯dence intervals with95%probabilit y. (Aside: ifyou'reinterested, theissueofcon¯dencebands fortheestimated survivalfunction arediscussed inSection 4.4ofKleinandMoeschberger) 104 Lookingatwhether thecon¯dence intervalsfor^S(t¤)overlap betweenthe6MPandplacebogroupswouldonlyfocuson comparing thetwotreatmen tgroupsatasinglepointin time,t¤.Wewantanoverallcomparison. Shouldwebaseouroverallcomparisonof^S(t)on: ²thefurthestdistance betweenthetwocurves? ²themediansurvivalforeachgroup? ²theaveragehazard? (forexponentialdistributions, this wouldbelikecomparing themeaneventtimes) ²addingupthedi®erence betweenthetwosurvivalesti- matesovertime? X j·^S(tjA)¡^S(tjB)¸ ²aweightedsumofdi®erences, wheretheweightsre°ect thenumberatriskateachtime? ²arank-based test?i.e.,wecouldrankalloftheevent times,andthenseewhether thesumofranksforone groupwaslessthantheother. 105 Nonparametric comparisonsofgroups Alloftheseareprettyreasonable options,andwe'llseethat therehavebeenseveralproposalsforhowtocompare the survivaloftwogroups.Forthemoment,wearestickingto nonparametric comparisons. Whynonparametric? ²fairlyrobust ²e±cientrelativetoparametrictests ²oftensimpleandintuitive Beforecontinuingthedescription ofthetwo-sample compar- ison,I'mgoingtotrytoputthisinageneralframeworkto giveaperspectiveofwherewe'reheadinginthisclass. 106 GeneralFrameworkforSurvivalAnalysis Weobserve(Xi;±i;Zi)forindividuali,where ²Xiisacensored failuretimerandomvariable ²±iisthefailure/censoring indicator ²Zirepresentsasetofcovariates NotethatZimightbeascalar(asinglecovariate,saytreat- mentorgender)ormaybea(p£1)vector(represen ting severaldi®erentcovariates). Thesecovariatesmightbe: ²continuous ²discrete ²time-varying(morelater) IfZiisascalarandisbinary,thenwearecomparing the survivaloftwogroups,likeintheleukemiaexample. Moregenerally though, itisusefultobuildamodelthat characterizes therelationship betweensurvivalandallofthe covariatesofinterest. 107 We'llproceedasfollows: ²Twogroupcomparisons ²Multigroup andstrati¯ed comparisons -strati¯ed logrank ²Failuretimeregression models {Coxproportionalhazardsmodel {Accelerated failuretimemodel 108 Twosampletests ²Mantel-Haenszel logranktest ²Peto&Peto'sversionofthelogranktest ²Gehan's Generalized Wilcoxon ²Peto&Peto'sandPrentice'sgeneralized Wilcoxon ²Tarone-WareandFleming-Harrington classes ²Cox'sF-test(non-parametric version) References: Hosmer&Lemesho wSection2.4 Collett Section2.5 Klein&Moeschberger Section7.3 Kleinbaum Chapter 2 Lee Chapter 5 109 Mantel-HaenszelLogranktest Thelogranktestisthemostwellknownandwidelyused. Italsohasanintuitiveappeal,building onstandard meth- odsforbinarydata.(Laterwewillseethatitcanalsobe obtained asthescoretestfromapartiallikelihoodfromthe CoxProportionalHazards model.) Firstconsider thefollowing(2£2)tableclassifying those withandwithouttheeventofinterestinatwogroupsetting: Event Group Yes No Total 0d0n0¡d0n0 1d1n1¡d1n1 Totaldn¡dn 110 Ifthemargins ofthistableareconsidered ¯xed,thend0 followsa ? distribution. Under thenullhypothesisofnoassociationbetweentheeventand group,itfollowsthat E(d0)=n0d n Var(d0)=n0n1d(n¡d) n2(n¡1) Therefore, underH0: Â2 MH=[d0¡n0d=n]2 n0n1d(n¡d) n2(n¡1)»Â2 1 ThisistheMantel-Haenszel statistic andisapproximately equivalenttothePearsonÂ2testforequalityofthetwo groupsgivenby: Â2 p=X(o¡e)2 e Note:recallthatthePearsonÂ2testwasderivedforthe casewhereonlytherowmargins were¯xed,andthusthe varianceabovewasreplaced by: Var(d0¡n0(d0+d1) n)=n0n1d(n¡d) n3 111 Example: Toxicityinaclinicaltrialwithtwotreatmen ts Toxicity Group YesNo Total 0 842 50 1 248 50 Total 10 90 100 Â2 p=4:00(p=0:046) Â2 MH=3:96(p=0:047) 112 NowsupposewehaveK(2£2)tables,allindependent,and wewanttotestforacommon groupe®ect.TheCochran- Mantel-Haenszel testforacommon oddsrationotequalto 1canbewrittenas: Â2 CMH=[PK j=1(d0j¡n0j¤dj=nj)]2 PKj=1n1jn0jdj(nj¡dj)=[n2j(nj¡1)] wherethesubscriptjreferstothej-thtable: Event Group Yes No Total 0d0jn0j¡d0jn0j 1d1jn1j¡d1jn1j Totaldjnj¡djnj Thisstatistic isdistributed approximately asÂ2 1. 113 Howdoesthisapplyinsurvivalanalysis? Supposeweobserve Group1:(X11;±11):::(X1n1;±1n1) Group0:(X01;±01):::(X0n0;±0n0) Wecouldjustcountthenumbersoffailures: eg.,d1= PK j=1±1j Example:Leukemiadata,justcountingupthenumber ofremissions ineachtreatmen tgroup. Fail Group YesNo Total 021 0 21 1 912 21 Total 30 12 42 Â2 p=16:8(p=0:001) Â2 MH=16:4(p=0:001) But,thisdoesn'taccountforthetimeatrisk. Conceptually ,wewouldliketocompare theKMsurvival curves.Let'sputthecomponentsside-by-sideandcompare. 114 Cox&OakesTable1.1Leukemiaexample Ordered Group0Group1 DeathTimesdjcjrjdjcjrj 120210021 220190021 310170021 420160021 520140021 600123121 700121017 840120016 90080116 100081115 112080113 122060012 130041012 151040011 160031011 171030110 19002019 20002018 22102107 23101106 25000015 NotethatIwrotedownthenumberatriskforGroup1fortimes 1-5eventhoughtherewerenoeventsorcensoringsatthosetimes. 115 LogrankTest:FormalDe¯nition Thelogranktestisobtained byconstructing a(2£2)ta- bleateachdistinctdeathtime,andcomparing thedeath ratesbetweenthetwogroups,conditional onthenumberat riskinthegroups.Thetablesarethencombinedusingthe Cochran-Man tel-Haenszel test. Note:ThelogrankissometimescalledtheCox-Manteltest. Lett1;:::;tKrepresenttheKordered, distinctdeathtimes. Atthej-thdeathtime,wehavethefollowingtable: Die/Fail Group Yes No Total 0d0jr0j¡d0jr0j 1d1jr1j¡d1jr1j Totaldjrj¡djrj whered0jandd1jarethenumberofdeathsingroup0and 1,respectivelyatthej-thdeathtime,andr0jandr1jare thenumberatriskatthattime,ingroups0and1. 116 Thelogranktestis: Â2 logrank=[PK j=1(d0j¡r0j¤dj=rj)]2 PKj=1r1jr0jdj(rj¡dj) [r2j(rj¡1)] Assuming thetablesareallindependent,thenthisstatistic willhaveanapproximateÂ2distribution with1df. Basedonthemotivationforthelogranktest, whichofthesurvival-relatedquantitiesarewe comparingateachtimepoint? ²PK j=1wj·^S1(tj)¡^S2(tj)¸ ? ²PK j=1wj·^¸1(tj)¡^¸2(tj)¸ ? ²PK j=1wj·^¤1(tj)¡^¤2(tj)¸ ? 117 Firstseveraltablesofleukemiadata CMHanalysis ofleukemia data TABLE1OFTRTMTBYREMISS TABLE3OFTRTMTBYREMISS CONTROLLING FORFAILTIME=1 CONTROLLING FORFAILTIME=3 TRTMT REMISS TRTMT REMISS Frequency| Frequency| Expected | 0| 1|Total Expected | 0| 1|Total ---------+--------+--------+ ---------+--------+--------+ 0|19|2|21 0|16|1|17 |20|1| |16.553|0.4474| ---------+--------+--------+ ---------+--------+--------+ 1|21|0|21 1|21|0|21 |20|1| |20.447|0.5526| ---------+--------+--------+ ---------+--------+--------+ Total 40 2 42 Total 37 1 38 TABLE2OFTRTMTBYREMISS TABLE4OFTRTMTBYREMISS CONTROLLING FORFAILTIME=2 CONTROLLING FORFAILTIME=4 TRTMT REMISS TRTMT REMISS Frequency| Frequency| Expected | 0| 1|Total Expected | 0| 1|Total ---------+--------+--------+ ---------+--------+--------+ 0|17|2|19 0|14|2|16 |18.05|0.95| |15.135|0.8649| ---------+--------+--------+ ---------+--------+--------+ 1|21|0|21 1|21|0|21 |19.95|1.05| |19.865|1.1351| ---------+--------+--------+ ---------+--------+--------+ Total 38 2 40 Total 35 2 37 118 CMHstatistic=logrankstatistic SUMMARY STATISTICS FORTRTMTBYREMISS CONTROLLING FORFAILTIME Cochran-Mantel-Haenszel Statistics (BasedonTableScores) Statistic Alternative Hypothesis DF Value Prob ----------------------------------------------------------------- 1 Nonzero Correlation 116.793 0.001 2 RowMeanScoresDiffer 116.793 0.001 3 General Association 116.793 0.001<===LOGRANK TEST Note:Although CMHworkstogetthecorrectlogranktest, itwouldrequireinputting thedjandrjateachtimeofdeath foreachtreatmen tgroup.There'saneasierwaytogetthe teststatistic, whichI'llshowyoushortly. 119 Calculatinglogrankstatisticbyhand LeukemiaExample: OrderedGroup0Combined DeathTimesd0jr0jdjrjejoj¡ejvj 12212421.001.000.488 22192400.951.05 31171380.450.55 42162370.861.14 5214235 6012333 7012129 8412428 1008123 1128221 1226218 1304116 1514115 1603114 1713113 221229 231127 Sum 10.2516.257 oj=d0j ej=djr0j=rj vj=r1jr0jdj(rj¡dj)=[r2 j(rj¡1)] Â2 logrank=(10:251)2 6:257=16:793 120 Notesaboutlogranktest: ²Thelogrankstatistic dependsonranksofeventtimes only ²Iftherearenotieddeaths,thenthelogrankhastheform: [PK j=1(d0j¡r0j rj)]2 PKj=1r1jr0j=r2j ²Numerator canbeinterpreted asP(o¡e)where\o"is theobservednumberofdeathsingroup0,and\e"is theexpectednumber,giventheriskset.Theexpected numberequals#deaths£proportioningroup0atrisk. ²The(o¡e)termsinthenumerator canbewrittenas r0jr1j rj(^¸1j¡^¸0j) ²Itdoesnotmatterwhichgroupyouchoosetosumover. Toseethis,notethatifwesummedup(o-e)overthedeath timesforthe6MPgroupwewouldget-10.251,andthesumof thevariancesisthesame.Sowhenwesquarethenumerator, theteststatisticisthesame. 121 Analogous totheCMHtestforaseriesoftablesatdi®erent levelsofaconfounder, thelogranktestismostpowerfulwhen \oddsratios"areconstantovertimeintervals.Thatis,itis mostpowerfulforproportionalhazards . Checkingtheassumptionofproportionalhazards: ²checktoseeiftheestimated survivalcurvescross-if theydo,thenthisisevidence thatthehazardsarenot proportional ²moreformaltest:anyideas? Whatshouldbedoneifthehazardsarenot proportional? ²Ifthedi®erence betweenhazardshasaconsisten tsign, thelogranktestusuallydoeswell. ²Othertestsareavailablethataremorepowerfulagainst di®erentalternativ es. 122 GettingthelogrankstatisticusingStata: Afterdeclaringdataassurvivaltypedatausing the\stset"command,issuethe\ststest"com- mand .stsetremissstatus datasetname:leukem id:-- (meaning eachrecordauniquesubject) entrytime:-- (meaning allentered attime0) exittime:remiss failure/censor: status .stslist,by(trt) Beg. NetSurvivor Std. Time Total FailLostFunction Error [95%Conf.Int.] ------------------------------------------------------------------- --- trt=0 1 21 200.9048 0.0641 0.6700 0.9753 2 19 200.8095 0.0857 0.5689 0.9239 3 17 100.7619 0.0929 0.5194 0.8933 4 16 200.6667 0.1029 0.4254 0.8250 . .(etc) .ststesttrt Log-rank testforequality ofsurvivor functions ------------------------------------------------ |Events trt|observed expected ------+------------------------- 0| 21 10.75 1| 9 19.25 ------+------------------------- Total| 30 30.00 chi2(1) =16.79 Pr>chi2 =0.0000 123 GettingthelogrankstatisticusingSAS ²StillusePROCLIFETEST ²Add\STRATA"command, withtreatmen tvariable ²Givesthechi-square test(2-sided), butalsogivesyou thetermsyouneedtocalculate the1-sidedtest;thisis usefulifwewanttoknowwhichofthetwogroupshas thehigherestimated hazardovertime. ²TheSTRATAcommand alsogivestheGehan-Wilco xon test(whichwewilltalkaboutnext) Title'CoxandOakesexample'; dataleukemia; inputweeksremisstrtmt; cards; 601 611 611 611 /*datafor6MPgroup*/ 711 901 etc 110 110 /*dataforplacebo group*/ 210 210 etc ; proclifetest data=leukemia; timeweeks*remiss(0); stratatrtmt; title'Logrank testforleukemia data'; run; 124 Outputfromleukemiaexample: Logrank testforleukemia data Summary oftheNumberofCensored andUncensored Values TRTMT Total Failed Censored %Censored 6-MP 21 9 1257.1429 Control 21 21 00.0000 Total 42 30 1228.5714 Testing Homogeneity ofSurvival CurvesoverStrata TimeVariable FAILTIME RankStatistics TRTMT Log-Rank Wilcoxon 6-MP -10.251 -271.00 Control 10.251 271.00 Covariance MatrixfortheLog-Rank Statistics TRTMT 6-MP Control 6-MP 6.25696 -6.25696 Control -6.25696 6.25696 TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 16.7929 10.0001 <==Here'stheonewewant!! Wilcoxon 13.4579 10.0002 -2Log(LR) 16.4852 10.0001 125 GettingthelogrankstatisticusingSplus: Insteadofthe\surv.¯t"command,usethe \surv.di®"commandwitha\group"(treatment) variable. Mantel-Haenszel logrank: >logrank<-surv.diff(weeks,remiss,trtmt) >logrank NObserved Expected (O-E)^2/E 021 2110.75 9.775 121 919.25 5.458 Chisq=16.8on1degrees offreedom, p=4.169e-05 126 Generalization oflogranktest =)Linearranktests Thelogrankandothertestscanbederivedbyassigning scorestotheranksofthedeathtimes,andaremembersof ageneralclassoflinearranktests(formoredetail,see Lee,ch5) First,de¯ne ^¤(t)=X j:tj<tdj rj wheredjandrjarethenumberofdeathsandthenumber atrisk,respectivelyatthej-thordereddeathtime. Thenassignthesescores(suggested byPetoandPeto): Event Score Deathattjwj=1¡^¤(tj) Censoring attjwj=¡^¤(tj) Tocalculate thelogranktest,simplysumupthescoresfor group0. 127 Example Group0:15,18,19,19,20 Group1:16+,18+,20+,23,24+ Calculationoflogrankasalinearrankstatistic Ordered DataGroupdjrj^¤(tj)scorewj 15 01100.100 0.900 16+1090.100-0.100 18 0180.225 0.775 18+1070.225-0.225 19 0260.558 0.442 20 0140.808 0.192 20+1030.808-0.808 23 1121.308-0.308 24+1011.308-1.308 ThelogrankstatisticSissumofscoresforgroup0: S=0:900+0:775+0:442+0:442+0:192=2:75 Thevarianceis: Var(S)=n0n1Pn j=1w2 j n(n¡1) Inthiscase,Var(S)=1:210,so Z=2:75p 1:210=2:50=)Â2 logrank=(2:50)2=6:25 128 Whyisthisformofthelogrankequivalent? Thelogrankstatistic SisequivalenttoP(o¡e)overthe distinctdeathtimes,where\o"istheobservednumberof deathsingroup0,and\e"istheexpectednumber,given therisksets. Atdeaths: weightsare1¡^¤ Atcensorings: weightsare¡^¤ Sowearesumming up\1's"fordeaths(togetd0j),andsub- tracting¡^¤atbothdeathsandcensorings. Thisamountsto subtracting dj=rjateachdeathorcensoring timeingroup 0,atorafterthej-thdeath.Sincethereareatotalofr0jof these,wegete=r0j¤dj=rj. Whyisitcalledthelogrank test? SinceS(t)=exp(¡¤(t)),analternativ eestimator ofS(t) is: ^S(t)=exp(¡^¤(t))=exp(¡X j:tj<tdj rj) So,wecanthinkof^¤(t)=¡log(^S(t))asyieldingthe\log- survival"scoresusedtocalculate thestatistic. 129 ComparingtheCMH-typeLogrankand \LinearRank"logrank A.CMH-typeLogrank: Wemotivatedthelogranktestthrough theCMHstatistic fortestingHo:OR=1overKtables,whereKisthe numberofdistinctdeathtimes.Thisturnedouttobewhat wegetwhenweusethelogrank(default) optioninStataor the\strata"statemen tinSAS. B.LinearRanklogrank: Thelinearrankversionofthelogranktestisbasedonadding up\scores" foroneofthetwotreatmen tgroups.Thepar- ticularscoresthatgaveusthesamelogrankstatistic were basedontheNelson-Aalen estimator, i.e.,^¤=P^¸(tj).This iswhatyougetwhenyouusethe\test"statemen tinSAS. Herearesomecomparisons, withanewexample toshow whenthetwotypesoflogrankstatistics willbeequal. 130 First,let'sgobacktoourexample fromChapter 5ofLee: Example Group0:15,18,19,19,20 Group1:16+,18+,20+,23,24+ A.TheCMH-typelogrankstatistic: (usingthestratastatement) RankStatistics TRTMT Log-Rank Wilcoxon Control 2.7500 18.000 Treated -2.7500 -18.000 Covariance MatrixfortheLog-Rank Statistics TRTMT Control Treated Control 1.08750 -1.08750 Treated -1.08750 1.08750 TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 6.9540 10.0084 Wilcoxon 5.5479 10.0185 -2Log(LR) 3.3444 10.0674 131 Thisisexactlythesamechi-square testthatyouwouldget ifyoucalculated thenumerator ofthelogrankasP(oj¡ej) andthevarianceasvj=r1jr0jdj(rj¡dj)=[r2 j(rj¡1)] OrderedGroup0Combined DeathTimesd0jr0jdjrjejoj¡ejvj 15151100.500.500.2500 1814180.500.500.2500 1923261.001.000.4000 2011240.250.750.1870 2300120.000.000.0000 Sum 2.751.0875 Â2 logrank=(2:75)2 1:0875=6:954 132 B.The\linearrank"logrankstatistic: (usingtheteststatement) Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square GROUP 2.7500 1.0897 6.3684 0.0116 Covariance MatrixfortheLOGRANKStatistics Variable TRTMT TRTMT 1.18750 Thisisactually veryclosetowhatwewouldgetifweuse theNelson-Aalen based\scores": Calculationoflogrankasalinearrankstatistic Ordered DataGroupdjrj^¤(tj)scorewj 15 01100.100 0.900 16+1090.100-0.100 18 0180.225 0.775 18+1070.225-0.225 19 0260.558 0.442 20 0140.808 0.192 20+1030.808-0.808 23 1121.308-0.308 24+1111.308-1.308 Sum(grp 0) 2.750 133 Notethatthenumerator istheexactsamenumber(2.75) inbothversionsofthelogranktest.Thedi®erence inthe denominator isduetothewaythattiesarehandled. CMH-typevariance: var=Xr1jr0jdj(rj¡dj) r2j(rj¡1) =Xr1jr0j rj(rj¡1)dj(rj¡dj) rj Linearranktypevariance: var=n0n1Pn j=1w2 j n(n¡1) 134 Nowconsideranexamplewheretherearenotied deathtimes ExampleIGroup0:15,18,19,21,22 Group1:16+,17+,20+,23,24+ A.TheCMH-typelogrankstatistic: (usingthestratastatement) RankStatistics TRTMT Log-Rank Wilcoxon Control 2.5952 15.000 Treated -2.5952 -15.000 Covariance MatrixfortheLog-Rank Statistics TRTMT Control Treated Control 1.21712 -1.21712 Treated -1.21712 1.21712 TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 5.5338 10.0187 Wilcoxon 4.3269 10.0375 -2Log(LR) 3.1202 10.0773 135 B.The\linearrank"logrankstatistic: (usingtheteststatement) Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square TRTMT 2.5952 1.1032 5.5338 0.0187 Covariance MatrixfortheLOGRANKStatistics Variable TRTMT TRTMT 1.21712 Notethatthistime,thevariances ofthetwologrankstatis- ticsareexactlythesame,equalto1.217. Iftherearenotiedeventtimes,thenthe twoversionsofthetestwillyieldidenti- calresults.Themoretieswehave,the moreitmatterswhichversionweuse. 136 Gehan'sGeneralizedWilcoxonTest First,let'sreviewtheWilcoxontestforuncensored data: Denoteobservationsfromtwosamplesby: (X1;X2;:::;Xn)and(Y1;Y2;:::;Ym) Orderthecombinedsampleandde¯ne: Z(1)<Z(2)<¢¢¢<Z(m+n) Ri1=rankofXi R1=m+nX i=1Ri1 RejectH0ifR1istoobigortoosmall,according to R1¡E(R1) r Var(R1)»N(0;1) where E(R1)=m(m+n+1) 2 Var(R1)=mn(m+n+1) 12 137 TheMann-Whitney formoftheWilcoxonisde¯nedas: U(Xi;Yj)=Uij=8 >>>>>< >>>>>:+1ifXi>Yj 0ifXi=Yj ¡1ifXi<Yj and U=nX i=1mX j=1Uij: Thereisasimplecorrespondence betweenUandR1: R1=m(m+n+1)=2+U=2 soU=2R1¡m(m+n+1) Therefore, E(U)=0 Var(U)=mn(m+n+1)=3 138 ExtendingWilcoxontocensoreddata TheMann-Whitney formleadstoageneralization forcen- soreddata.De¯ne U(Xi;Yj)=Uij=8 >>>>>< >>>>>:+1ifxi>yjorx+ i¸yj 0ifxi=yiorlowervaluecensored ¡1ifxi<yjorxi·y+ j Thende¯ne W=nX i=1mX j=1Uij Thus,thereisacontribution toWforeverycomparison wherebothobservationsarefailures(exceptforties),or whereacensored observationisgreaterthanorequaltoa failure. Lookingatallpossiblepairsofindividuals betweenthetwo treatmen tgroupsmakesthisanightmaretocompute by hand! 139 Gehanfoundaneasierwaytocompute theabove.First, poolthesampleof(n+m)observationsintoasinglegroup, thencompare eachindividual withtheremainingn+m¡1: Forcomparing thei-thindividual withthej-th,de¯ne Uij=8 >>>>>< >>>>>:+1ifti>tjort+ i¸tj ¡1ifti<tjorti·t+ j 0 otherwise Then Ui=m+nX j=1Uij Thus,forthei-thindividual, Uiisthenumberofobserva- tionswhicharede¯nitely lessthantiminusthenumberof observationsthatarede¯nitely greaterthanti.Weassume censorings occurafterdeaths,sothatifti=18+andtj=18, thenweadd1toUi. TheGehanstatistic isde¯nedas U=m+nX i=1Ui1fiingroup0g =W Uhasmean0andvariance var(U)=mn (m+n)(m+n¡1)m+nX i=1U2 i 140 Example fromLee: Group0:15,18,19,19,20 Group1:16+,18+,20+,23,24+ TimeGroupUiU2 i 150-9 81 16+1 1 1 180-6 36 18+1 2 4 190-2 4 190-2 4 200 1 1 20+1 525 231 416 24+1 636 SUM -18 208 U=¡18 Var(U)=(5)(5)(208) (10)(9) =57:78 andÂ2=(¡18)2=57:78=5:61 141 SAScode: dataleedata; infile'lee.dat'; inputtimecensgroup; proclifetest data=leedata; timetime*cens(0); stratagroup; run; SASOUTPUT: GehansWilcoxontest RankStatistics TRTMT Log-Rank Wilcoxon Control 2.7500 18.000 Treated -2.7500 -18.000 Covariance MatrixfortheWilcoxon Statistics TRTMT Control Treated Control 58.4000 -58.4000 Treated -58.4000 58.4000 TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 6.9540 10.0084 Wilcoxon 5.5479 10.0185**thisisGehan's test -2Log(LR) 3.3444 10.0674 142 NotesaboutSASWilcoxonTest: SAScalculates theWilcoxonas¡UinsteadofU,probably sothatthesignoftheteststatistic isconsisten twiththe logrank. SASgetssomething slightlydi®erentforthevariance, and thisdoesnotseemtodependonwhether thereareties. Forexample, thehypothetical datasetonp.6without ties yieldsU=¡15andPU2 i=182,so Var(U)=(5)(5)(182) (10)(9)=50:56andÂ2=(¡15)2 50:56=4:45 whileSASgivesthefollowing: RankStatistics TRTMT Log-Rank Wilcoxon Control 2.5952 15.000 Treated -2.5952 -15.000 Covariance MatrixfortheWilcoxon Statistics TRTMT Control Treated Control 52.0000 -52.0000 Treated -52.0000 52.0000 TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 5.5338 10.0187 Wilcoxon 4.3269 10.0375 -2Log(LR) 3.1202 10.0773 143 ObtainingtheWilcoxontestusingStata Usetheststeststatement,withtheappropriate option ststest varlist [ifexp][inrange] [,[logrank|wilcoxon|cox] strata(varlist) detail mat(matname1 matname2) notitle noshow] logrank, wilcoxon, andcoxspecify whichtestofequality isdesired. logrank isthedefault, andcoxyieldsalikelihood ratiotest underacoxmodel. Example:(leukemiadata) .stsetremissstatus .ststesttrt,wilcoxon Wilcoxon (Breslow) testforequality ofsurvivor functions ---------------------------------------------------------- |Events Sumof trt|observed expected ranks ------+-------------------------------------- 0| 21 10.75 271 1| 9 19.25 -271 ------+-------------------------------------- Total| 30 30.00 0 chi2(1) =13.46 Pr>chi2 =0.0002 144 GeneralizedWilcoxon (Peto&Peto,Prentice) Assignthefollowingscores: Foradeathatt: ^S(t+)+^S(t¡)¡1 Foracensoring att:^S(t+)¡1 Theteststatistic isP(scores)forgroup0. TimeGroupdjrj^S(t+)scorewj 1501100.900 0.900 16+1090.900 -0.100 180180.788 0.688 18+1070.788 -0.212 190260.525 0.313 200140.394 -0.081 20+1030.394 -0.606 231120.197 -0.409 24+1010.197 -0.803 Xwj1fjingroup0g=0:900+0:688+2¤(0:313)+(¡0:081) =2:13 Var(S)=n0n1Pn j=1w2 j n(n¡1)=0:765 soZ=2:13=0:765=2:433 145 TheTarone-Wareclassoftests: Thisgeneralclassoftestsislikethelogranktest,butadds weightswj.Thelogrank test,Wilcoxontest,andPeto- PrenticeWilcoxonareincluded asspecialcases. Â2 tw=[PK j=1wj(d1j¡r1j¤dj=rj)]2 PK l=1w2jr1jr0jdj(rj¡dj) r2j(rj¡1) Test Weightwj Logrank wj=1 Gehan's Wilcoxonwj=rj Peto/Pren tice wj=ncS(tj) Fleming-Harrington wj=[^S(tj)]® Tarone-Ware wj=prj Note:theseweightswjarenotthesameasthescoreswjwe'vebeen talkingaboutearlier,andtheyapplytotheCMH-typeformofthe teststatisticratherthanP(scores)overasingletreatmentgroup. 146 Whichtestshouldweused? CMH-typeorLinearRank? Iftherearenotahighproportionofties,thenitdoesn't reallymattersince: ²ThetwoWilcoxonsaresimilartoeachother ²Thetwologranktestsaresimilartoeachother Note:personally,ItendtousetheCMH-typetest,whichyougetwiththestrata statementinSASandtheteststatementinSTATA. LogrankorWilcoxon? ²BothtestshavetherightTypeIpowerfortestingthe nullhypothesisofequalsurvival,Ho:S1(t)=S2(t) ²Thechoiceofwhichtestmaytherefore dependonthe alternativ ehypothesis,whichwilldrivethepowerofthe test. 147 ²TheWilcoxonissensitivetoearlydi®erences between survival,whilethelogrankissensitivetolaterones.This canbeseenbytherelativeweightstheyassigntothetest statistic: LOGRANK numerator=X j(oj¡ej) WILCOXONnumerator=X jrj(oj¡ej) ²Thelogrankismostpowerfulundertheassumption of proportional hazards, whichimpliesanalternativ ein termsofthesurvivalfunctions ofHa:S1(t)=[S2(t)]® ²TheWilcoxonhashighpowerwhenthefailuretimes arelognormally distributed, withequalvarianceinboth groupsbutadi®erentmean.Itwillturnoutthatthisis theassumption ofanaccelerated failuretimemodel. ²Bothtestswilllackpowerifthesurvivalcurves(orhaz- ards)\cross". However,thatdoesnotnecessarily make theminvalid! 148 ComparisonbetweenTESTandSTRATAinSAS for2examples: DatafromLee(n=10) : fromSTRATA: TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 6.9540 10.0084 Wilcoxon 5.5479 10.0185**thisisGehan's test -2Log(LR) 3.3444 10.0674 fromTEST: Univariate Chi-Squares fortheWILCOXON Test Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square GROUP 1.8975 0.7508 6.3882 0.0115 Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square GROUP 2.7500 1.0897 6.3684 0.0116 149 Previousexamplewithleukemiadata: fromSTRATA: TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 16.7929 10.0001 Wilcoxon 13.4579 10.0002 -2Log(LR) 16.4852 10.0001 fromTEST: Univariate Chi-Squares fortheWILCOXON Test Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square GROUP 6.6928 1.7874 14.0216 0.0002 Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square GROUP 10.2505 2.5682 15.9305 0.0001 150 P-sampleandstrati¯edlogranktests Wehavebeendiscussing twosampleproblems. Inpractice, morecomplex settingsoftenarise: ²Therearemorethantwotreatmen tsorgroups,andthe question ofinterestiswhetherthegroupsdi®erfromeach other. ²Weareinterested inacomparison betweentwogroups, butwewishtoadjustforanotherfactorthatmaycon- foundtheanalysis ²Wewanttoadjustforlotsofcovariates. Wewill¯rsttalkaboutcomparing thesurvivaldistributions betweenmorethan2groups,andthenaboutadjusting for othercovariates. 151 P-samplelogrank SupposeweobservedatafromPdi®erentgroups,andthe datafromgroupp(p=1;:::;P)are: (Xp1;±p1):::(Xpnp;±pnp) Wenowconstruct a(P£2)tableateachoftheKdistinct deathtimes,andcompare thedeathratesbetweentheP groups,conditional onthenumberatrisk. Lett1;::::tKrepresenttheKordered, distinctdeathtimes. Atthej-thdeathtime,wehavethefollowingtable: Die/Fail Group Yes No Total 1d1jr1l¡d1jr1j . . . . PdPjrPj¡dPjrPj Totaldjrj¡djrj wheredpjisthenumberofdeathsingrouppatthej-th deathtime,andrpjisthenumberatriskatthattime. ThetablesarethencombinedusingtheCMHapproach. 152 Ifwewerejustfocusingonthisonetable,thenaÂ2 (P¡1)test statistic couldbeconstructed through acomparison of\o"s and\e"s,likebefore. Example: Toxicityinaclinicaltrialwith3treatmen ts TABLEOFGROUPBYTOXICITY GROUP TOXICITY Frequency| RowPct|no |yes |Total ---------+--------+--------+ 1|42|8|50 |84.00|16.00| ---------+--------+--------+ 2|48|2|50 |96.00|4.00| ---------+--------+--------+ 3|38|12|50 |76.00|24.00| ---------+--------+--------+ Total 128 22 150 STATISTICS FORTABLEOFGROUPBYTOXICITY Statistic DFValue Prob ------------------------------------------------------ Chi-Square 28.097 0.017 Likelihood RatioChi-Square 29.196 0.010 Mantel-Haenszel Chi-Square 11.270 0.260 Cochran-Mantel-Haenszel Statistics (BasedonTableScores) Statistic Alternative Hypothesis DFValue Prob ---------------------------------------------------------- 1 Nonzero Correlation 11.270 0.260 2 RowMeanScoresDiffer 28.043 0.018 3 General Association 28.043 0.018 153 FormalCalculations: LetOj=(d1j;:::d(P¡1)j)Tbeavectoroftheobservednum- beroffailuresingroups1to(P¡1),respectively,atthe j-thdeathtime.Giventherisksetsr1j,...rPj,andthefact thattherearedjdeaths,thenOjhasadistribution likea multivariateversionoftheHypergeometric.Ojhasmean: Ej=(djr1j rj;:::;djr(P¡1)j rj)T andvariancecovariancematrix: Vj=0 BBBBBBBB@v11jv12j:::v1(P¡1)j v22j:::v2(P¡1)j ::::::::: v(P¡1)(P¡1)j1 CCCCCCCCA wherethe`-thdiagonal elementis: v``j=r`j(rj¡r`j)dj(rj¡dj)=[r2 j(rj¡1)] andthe`m-tho®-diagonal elementis: v`mj=r`jrmjdj(rj¡dj)=[r2 j(rj¡1)] 154 TheresultingÂ2testforasingle(P£1)tablewouldhave (P-1)degreesandisconstructed asfollows: (Oj¡Ej)TV¡1 j(Oj¡Ej) GeneralizingtoKtables Analogous towhatwedidforthetwosamplelogrank, we replacetheOj,EjandVjwiththesumsovertheKdistinct deathtimes.Thatis,letO=Pk j=1Oj,E=Pk j=1Ej,and V=Pk j=1Vj.Then,theteststatistic is: (O¡E)TV¡1(O¡E) 155 Example: Timetakento¯nishatestwith3di®erentnoisedistractions. Alltestswerestoppedafter12minutes. NoiseLevel Group Group Group 1 2 3 9.0 10.0 12.0 9.5 12.0 12+ 9.0 12+12+ 8.5 11.0 12+ 10.0 12.0 12+ 10.5 10.5 12+ 156 Letsstartthecalculations... Observeddatatable OrderedGroup1Group2Group3Combined Timesd1jr1jd2jr2jd3jr3jdjrj 8.5160606 9.0250606 9.5130606 10.0121606 10.5111506 11.0001406 12.0002316 Expectedtable OrderedGroup1Group2Group3Combined Timeso1je1jo2je2jo3je3jojej 8.5 9.0 9.5 10.0 10.5 11.0 12.0 DoingtheP-sampletestbyhandiscumbersome... Luckily,moststatistical packageswilldoitforyou! 157 P-samplelogrankinStata .stsgraph,by(group) .ststestgroup,logrank Log-rank testforequality ofsurvivor functions ------------------------------------------------ |Events group|observed expected ------+------------------------- 1| 6 1.57 2| 5 4.53 3| 1 5.90 ------+------------------------- Total| 12 12.00 chi2(2) =20.38 Pr>chi2 =0.0000 .ststestgroup,wilcoxon Wilcoxon (Breslow) testforequality ofsurvivor functions ---------------------------------------------------------- |Events Sumof group|observed expected ranks ------+-------------------------------------- 1| 6 1.57 68 2| 5 4.53 -5 3| 1 5.90 -63 ------+-------------------------------------- Total| 12 12.00 0 chi2(2) =18.33 Pr>chi2 =0.0001 158 SASprogramforP-samplelogrank Title'Testing withnoiseexample'; datanoise; inputtesttime finishgroup; cards; 9 11 9.5 11 9.0 11 8.5 11 10 11 10.5 11 10.0 12 12 12 12 02 11 12 12 12 10.5 12 12 13 12 03 12 03 12 03 12 03 12 03 ; proclifetest data=noise; timetesttime*finish(0); stratagroup; run; 159 Testing Homogeneity ofSurvival CurvesoverStrata TimeVariable TESTTIME RankStatistics GROUP Log-Rank Wilcoxon 1 4.4261 68.000 2 0.4703 -5.000 3 -4.8964 -63.000 Covariance MatrixfortheLog-Rank Statistics GROUP 1 2 3 1 1.13644 -0.56191 -0.57454 2 -0.56191 2.52446 -1.96255 3 -0.57454 -1.96255 2.53709 Covariance MatrixfortheWilcoxon Statistics GROUP 1 2 3 1 284.808 -141.495 -143.313 2 -141.495 466.502 -325.007 3 -143.313 -325.007 468.320 TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 20.3844 20.0001 Wilcoxon 18.3265 20.0001 -2Log(LR) 5.5470 20.0624 160 Note:donotuseTestinSASPROCLIFETEST ifyou wantaP-sample logrank. Testwillinterpretthegroup variableasameasured covariate(i.e.,eitherordinalorcon- tinuous). Inotherwords,youwillgetatrendtestwithonly1degree offreedom, ratherthanaP-sample testwith(p-1)df. Forexample, here'swhatwegetifweusetheTESTstate- mentonthenoiseexample: proclifetest data=noise; timetesttime*finish(0); testgroup; run; SASOUTPUT: Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square GROUP 9.3224 2.2846 16.6503 0.0001 Covariance MatrixfortheLOGRANKStatistics Variable GROUP GROUP 5.21957 Forward Stepwise Sequence ofChi-Squares fortheLOGRANKTest Pr>Chi-Square Pr> Variable DFChi-Square Chi-Square Increment Increment GROUP 116.6503 0.0001 16.6503 0.0001 161 TheStrati¯edLogrank Sometimes, eventhoughweareinterestedincomparing two groups(ormaybeP)groups,weknowthereareotherfactors thatalsoa®ecttheoutcome. Itwouldbeusefultoadjustfor theseotherfactorsinsomeway. Example: Forthenursinghomedata,alogranktestcom- paringlengthofstayforthoseunderandover85yearsof agesuggests asigni¯can tdi®erence (p=0.03). However,weknowthatgenderhasastrongassociationwith lengthofstay,andalsoage.Hence,itwouldbeagoodidea toSTRATIFYtheanalysisbygenderwhentryingtoassess theagee®ect. Astrati¯edlogrank allowsonetocompare groups,but allowstheshapesofthehazardsofthedi®erentgroupsto di®eracrossstrata.Itmakestheassumption thatthegroup 1vsgroup2hazardratioisconstantacrossstrata. Inotherwords:¸1s(t) ¸2s(t)=µwhereµisconstantoverthestrata (s=1;:::;S). Thismethodofadjusting forothervariablesisnotas°exible asthatbasedonamodellingapproach. 162 Generalsetupforthestrati¯edlogrank: Supposewewanttoassesstheassociationbetweensurvival andafactor(callthisX)thathastwodi®erentlevels.Sup- posehowever,thatwewanttostratifybyasecondfactor, thathasSdi®erentlevels. First,dividethedataintoSseparate groups.Withingroup s(s=1;:::;S),proceedasthoughyouwereconstructing thelogranktoassesstheassociationbetweensurvivaland thevariableX.Thatis,lett1s;:::;tKssrepresenttheKs ordered, distinctdeathtimesinthes-thgroup. Atthej-thdeathtimeingroups,wehavethefollowing table: Die/Fail XYes No Total 1ds1jrs1j¡ds1jrs1j 2ds2jrs2j¡ds2jrs2j Totaldsjrsj¡dsjrsj 163 LetOsbethesumofthe\o"sobtained byapplying the logrankcalculations intheusualwaytothedatafromgroup s.Similarly ,letEsbethesumofthe\e"s,andVsbethe sumofthe\v"s. Thestrati¯edlogrank is Z=PS s=1(Os¡Es) rPSs=1(Vs) 164 Strati¯edlogrankusingStata: .usenurshome .genage1=0 .replace age1=1ifage>85 .ststestage1,strata(gender) failure _d:cens analysis time_t:los Stratified log-rank testforequality ofsurvivor functions ----------------------------------------------------------- |Events age1|observed expected(*) ------+------------------------- 0| 795 764.36 1| 474 504.64 ------+------------------------- Total|1269 1269.00 (*)sumovercalculations withingender chi2(1) = 3.22 Pr>chi2 =0.0728 165 Strati¯edlogrankusingSAS: datapop1; setpop; age1=0; ifage>85thenage1=1; proclifetest data=pop1 outsurv=survres; timestay*censor(1); testage1; stratagender; RESULTS(justthelogrankpart....youcanalsodoastrati¯ed Wilcoxon) TheLIFETEST Procedure RankTestsfortheAssociation ofLSTAYwithCovariates PooledoverStrata Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square AGE1 29.1508 17.1941 2.8744 0.0900 Covariance MatrixfortheLOGRANKStatistics Variable AGE1 AGE1 295.636 Forward Stepwise Sequence ofChi-Squares fortheLOGRANKTest Pr> Chi-Square Pr> Variable DFChi-Square Chi-Square Increment Increment AGE1 1 2.8744 0.0900 2.8744 0.0900 166 ModelingofSurvivalData Nowwewillexploretherelationship betweensurvivaland explanatory variablesbymodeling.Inthisclass,weconsider twobroadclassesofregression models: ²ProportionalHazards(PH)models ¸(t;Z)=¸0(t)ª(Z) Mostcommonly ,wewritethesecondtermas: ª(Z)=e¯Z SupposeZ=1fortreatedsubjectsandZ=0forun- treatedsubjects.Thenthismodelsaysthatthehazard isincreased byafactorofe¯fortreatedsubjectsversus untreatedsubjects(c¯mightbe<1). Thisisanexample ofasemi-parametric model. ²AcceleratedFailureTime(AFT)models log(T)=¹+¯Z+¾w wherewisan\errordistribution". Typically,weplace aparametric assumption onw: {exponential,Weibull,Gamma {lognormal 167 Covariates: Ingeneral,Zisavectorofcovariatesofinterest. Zmayinclude: ²continuousfactors(eg,age,bloodpressure), ²discretefactors(gender, maritalstatus), ²possibleinteractions (agebysexinteraction) DiscreteCovariates: Justasinstandard linearregression, ifwehaveadiscrete covariateAwithalevels,thenwewillneedtoinclude(a¡1) dummyvariables(U1;U2;:::;Ua)suchthatUj=1ifA= j.Then ¸i(t)=¸0(t)exp(¯2U2+¯3U3+¢¢¢+¯aUa) (Intheabovemodel,thesubgroup withA=1orU1=1is thereference group.) Interactions: Twofactors,AandB,interactifthehazardofdeathde- pendsonthecombination oflevelsofAandB. Weusuallyfollowtheprinciple ofhierarchicalmodels,and onlyincludeinteractions ifallofthecorrespondingmain e®ectsarealsoincluded. 168 Theexample Ijustgavewasbasedonaproportionalhazards model,butthedescription ofthetypesofcovariateswemight wanttoincludeinourmodelappliestoboththeAFTand PHmodel. We'llstartoutbyfocusingontheCoxPHmodel,andad- dresssomeofthefollowingquestions: ²Whatdoestheterm¸0(t)mean? ²What's\proportional" aboutthePHmodel? ²Howdoweestimate theparameters inthemodel? ²Howdoweinterprettheestimated values? ²Howcanweconstruct testsofwhether thecovariates haveasigni¯can te®ectonthedistribution ofsurvival times? ²Howdothesetestscompare tothelogranktestorthe Wilcoxontest? 169 TheCoxProportionalHazardsmodel ¸(t;Z)=¸0(t)exp(¯Z) Thisisthemostcommon modelusedforsurvivaldata. Why? ²°exiblechoiceofcovariates ²fairlyeasyto¯t ²standard softwareexists References: Collett,Chapter 3* Lee,Chapter 10* Hosmer&Lemesho w,Chapters 3-7 Allison,Chapter 5 CoxandOakes,Chapter 7 Kleinbaum,Chapter 3 KleinandMoeschberger,Chapters 8&9 Kalb°eisc handPrentice Note:somebooks(likeCollettandH&L)useh(t;X)as theirstandard notation forthehazardinsteadof¸(t;Z),and H(t)forthecumulativehazardinsteadof¤(t). 170 Whydowecallitproportionalhazards? Thinkofthe¯rstexample, whereZ=1fortreatedandZ= 0forcontrol.Thenifwethinkof¸1(t)asthehazardrate forthetreatedgroup,and¸0(t)asthehazardforcontrol, thenwecanwrite: ¸1(t)=¸(t;Z=1)=¸0(t)exp(¯Z) =¸0(t)exp(¯) Thisimpliesthattheratioofthetwohazardsisaconstant, Á,whichdoesNOTdependontime,t.Inotherwords,the hazardsofthetwogroupsremainproportionalovertime. Á=¸1(t) ¸0(t)=e¯ Áisreferredtoasthehazardratio. Whatistheinterpretationof¯here? 171 TheBaselineHazardFunction Intheexample ofcomparing twotreatmen tgroups,¸0(t)is thehazardrateforthecontrolgroup. Ingeneral,¸0(t)iscalledthebaselinehazardfunction , andre°ectstheunderlying hazardforsubjectswithallco- variatesZ1;:::;Zpequalto0(i.e.,the\reference group"). Thegeneralformis: ¸(t;Z)=¸0(t)exp(¯1Z1+¯2Z2+¢¢¢+¯pZp) Sowhenwesubstitute alloftheZj'sequalto0,weget: ¸(t;Z=0)=¸0(t)exp(¯1¤0+¯2¤0+¢¢¢+¯p¤0) =¸0(t) Inthegeneralcase,wethinkofthei-thindividual havinga setofcovariatesZi=(Z1i;Z2i;:::;Zpi),andwemodeltheir hazardrateassomemultipleofthebaseline hazardrate: ¸i(t;Zi)=¸0(t)exp(¯1Z1i+¢¢¢+¯pZpi) 172 Thismeanswecanwritethelogofthehazardratioforthe i-thindividual tothereference groupas: log0 B@¸i(t) ¸0(t)1 CA=¯1Z1i+¯2Z2i+¢¢¢+¯pZpi TheCoxProportionalHazardsmodelisa linearmodelforthelogofthehazardratio OneofthebiggestadvantagesoftheframeworkoftheCox PHmodelisthatwecanestimate theparameters¯which re°ectthee®ectsoftreatmen tandothercovariateswithout havingtomakeanyassumptions abouttheformof¸0(t). Inotherwords,wedon'thavetoassumethat¸0(t)follows anexponentialmodel,oraWeibullmodel,oranyotherpar- ticularparametric model. That'swhatmakesthemodelsemi-parametric. Questions: 1.Whydon'twejustmodelthehazardratio, Á=¸i(t)=¸0(t),directlyasalinearfunctionofthe covariatesZ? 2.Whydoesn'tthemodelhaveanintercept? 173 Howdoweestimatethemodelparameters? ThebasicideaisthatunderPH,information about¯can beobtained fromtherelativeorderings (i.e.,ranks)ofthe survivaltimes,ratherthantheactualvalues.Why? SupposeTfollowsaPHmodel: ¸(t;Z)=¸0(t)e¯Z NowconsiderT¤=g(T),wheregisamonotonic increasing function. WecanshowthatT¤alsofollowsthePHmodel, withthesamemultiplier,e¯Z. Therefore, whenweconsider likelihoodmethodsforestimat- ingthemodelparameters, weonlyhavetoworryaboutthe ranksofthesurvivaltimes. 174 LikelihoodEstimationforthePHModel Kalb°eisc handPrenticederivealikelihoodinvolvingonly ¯andZ(not¸0(t))basedonthemarginal distribution of theranksoftheobservedfailuretimes(intheabsence of censoring). Cox(1972)derivedthesamelikelihood,andgeneralized it forcensoring, usingtheideaofapartiallikelihood Supposeweobserve(Xi;±i;Zi)forindividuali,where ²Xiisacensored failuretimerandomvariable ²±iisthefailure/censoring indicator (1=fail,0=censor) ²Zirepresentsasetofcovariates Thecovariatesmaybecontinuous,discrete, ortime-varying. 175 SupposethereareKdistinctfailure(ordeath)times,and let¿1;::::¿KrepresenttheKordered, distinctdeathtimes. Fornow,assumetherearenotieddeathtimes. LetR(t)=fi:xi¸tgdenotethesetofindividuals who are\atrisk"forfailureattimet. Moreaboutrisksets: ²IwillrefertoR(¿j)astherisksetatthejthfailuretime ²IwillrefertoR(Xi)astherisksetatthefailuretimeof individuali ²Therewillstillberjindividuals inR(¿j). ²rjisanumber,whileR(¿j)identi¯estheactualsubjects atrisk 176 Whatisthepartiallikelihood? Intuitively,itisaproductoverthesetofobserveddeath timesoftheconditional probabilities ofseeingtheobserved deaths,giventhesetofindividuals atriskatthosetimes. Ateachdeathtime¿j,thecontribution tothelikelihoodis: Lj(¯)=Pr(individual jfailsj1failurefromR(¿j)) =Pr(individual jfailsjatriskat¿j) P `2R(¿j)Pr(individual `failsjatriskat¿j) =¸(¿j;Zj) P `2R(¿j)¸(¿j;Z`) UnderthePHassumption, ¸(t;Z)=¸0(t)e¯Z,soweget: Lpartial(¯)=KY j=1¸0(¿j)e¯Zj P `2R(¿j)¸0(¿j)e¯Z` =KY j=1e¯Zj P `2R(¿j)e¯Z` 177 Anotherderivation: Ingeneral,thelikelihoodcontributions forcensored datafall intotwocategories: ²IndividualiscensoredatXi: Li(¯)=S(Xi)=exp[¡ZXi 0¸i(u)du] ²IndividualfailsatXi: Li(¯)=S(Xi)¸i(Xi)=¸i(Xi)exp[¡ZXi 0¸i(u)du] Thus,everyonecontributesS(Xi)tothelikelihood,andonly thosewhofailcontribute¸i(Xi). Thismeanswegetatotallikelihoodof: L(¯)=nY i=1¸i(Xi)±iexp[¡ZXi 0¸i(u)du] Theabovelikelihoodholdsforallcensored survivaldata, withgeneralhazardfunction¸(t).Inotherwords,wehaven't usedtheCoxPHassumption atallyet. 178 Now,let'smultiplyanddividebytheterm·P j2R(Xi)¸i(Xi)¸±i: L(¯)=nY i=12 4¸i(Xi) P j2R(Xi)¸i(Xi)3 5±i2 64X j2R(Xi)¸i(Xi)3 75±i exp[¡ZXi 0¸i(u)du] Cox(1972)arguedthatthe¯rstterminthisproductcon- tainedalmostalloftheinformation about¯,whilethesec- ondtwotermscontainedtheinformation about¸0(t),i.e., thebaseline hazard. Ifwejustfocusonthe¯rstterm,thenundertheCoxPH assumption: L(¯)=nY i=12 64¸i(Xi) P j2R(Xi)¸i(Xi)3 75±i =nY i=12 64¸0(Xi)exp(¯Zi) P j2R(Xi)¸0(Xi)exp(¯Zj)3 75±i =nY i=12 64exp(¯Zi) P j2R(Xi)exp(¯Zj)3 75±i Thisisthepartiallikelihoodde¯nedbyCox.Notethatit doesnotdependontheunderlying hazardfunction¸0(¢). Coxrecommends treating thisasanordinary likelihoodfor makinginferences about¯inthepresence ofthenuisance parameter¸0(¢). 179 Asimpleexample: individual Xi±iZi 1 914 2 805 3 617 4 1013 Nowlet'scompilethepiecesthatgointothepartiallikeli- hoodcontributions ateachfailuretime: ordered failure Likelihoodcontribution jtimeXiR(Xi)ij· e¯Zi=P j2R(Xi)e¯Zj¸±i 16f1,2,3,4g3e7¯=[e4¯+e5¯+e7¯+e3¯] 28f1,2,4g2 1 39f1,4g1e4¯=[e4¯+e3¯] 410f4g4 e3¯=e3¯=1 Thepartiallikelihoodwouldbetheproductofthesefour terms. 180 Notesonthepartiallikelihood: L(¯)=nY j=12 664e¯Zj P `2R(Xj)e¯Z`3 775±j =KY j=1e¯Zj P `2R(¿j)e¯Z` wheretheproductisovertheKdeath(orfailure)times. ²contributions onlyatthedeathtimes ²thepartiallikelihoodisNOTaproductofindependent terms,butofconditional probabilities ²Thereareotherchoicesbesidesª(Z)=e¯Z,butthis isthemostcommon andtheoneforwhichsoftwareis generally available. 181 PartialLikelihoodinference Inference canbeconducted bytreatingthepartiallikelihood asthoughitsatis¯ed alltheregularlikelihoodproperties. Thelog-partiallikelihoodis: `(¯)=log2 664nY j=1e¯Zj P `2R(¿j)e¯Z`3 775±j =log2 664KY j=1e¯Zj P `2R(¿j)e¯Z`3 775 =KX j=12 664¯Zj¡log[X `2R(¿j)e¯Z`]3 775 =KX j=1lj(¯) whereljisthelog-partial likelihoodcontribution atthej-th ordereddeathtime. 182 Supposethereisonlyonecovariate(¯isone-dimensional): Thepartiallikelihoodscoreequations are: U(¯)=@ @¯`(¯)=nX j=1±j2 664Zj¡P `2R(¿j)Z`e¯Z` P `2R(¿j)e¯Z`3 775 WecanexpressU(¯)intuitivelyasasumof\observed"mi- nus\expected"values: U(¯)=@ @¯`(¯)=nX j=1±j(Zj¡¹Zj) where¹Zjisthe\weightedaverage"ofthecovariateZover alltheindividuals intherisksetattime¿j.Notethat¯is involvedthrough theterm¹Zj. Themaximumpartiallikelihoodestimators canbefoundby solvingU(¯)=0. 183 Analogous tostandard likelihoodtheory,itcanbeshown (though noteasily)that (c¯¡¯) se(^¯)»N(0;1) Thevarianceof^¯canbeobtained byinvertingthesecond derivativeofthepartiallikelihood, var(^¯)»2 64¡@2 @¯2`(¯)3 75¡1 Fromtheaboveexpression forU(¯),wehave: @2 @¯2`(¯)=nX j=1±j2 664¡P `2R(¿j)(Zj¡¹Zj)2e¯Z` P `2R(¿j)e¯Z`3 775 Note: Thetruevarianceof^¯endsupbeingafunctionof¯,which isunknown.Wecalculate the\observed"information by substituting inourpartiallikelihoodestimateof¯intothe aboveformulaforthevariance 184 SimpleExamplefor2-groupcomparison:(noties) Group0:4+;7;8+;9;10+=)Zi=0 Group1:3;5;5+;6;8+=)Zi=1 orderedfailureNumberatriskLikelihoodcontribution jtimeXiGroup0Group1h e¯Zi=P j2R(Xi)e¯Zji±i 13 55 e¯=[5+5e¯] 25 44 e¯=[4+4e¯] 36 42 e¯=[4+2e¯] 47 41 e¯=[4+1e¯] 59 20 e0=[2+0]=1=2 Again,wetaketheproductoverthelikelihoodcontributions, thenmaximize togetthepartialMLEfor¯. Whatdoes¯representinthiscase? 185 Notes ²The\observed"information matrixisgenerally usedbe- causeinpractice, people¯ndithasbetterproperties. Also,the\expected"isveryhardtocalculate. ²Thereisaniceanalogy withthescoreandinforma- tionmatrices frommorestandard regression problems, exceptthatherewearesumming overobserveddeath times,ratherthanindividuals. ²NewtonRaphson isusedbymanyofthecomputer pack- agestosolvethepartiallikelihoodequations. 186 FittingCoxPHmodelwithStata Usesthe\stcox"command. First,trytyping\helpstcox" ------------------------------------------------------------------- --- helpforstcox ------------------------------------------------------------------- --- Estimate Coxproportional hazards model --------------------------------------- stcox[varlist] [ifexp][inrange] [,nohrstrata(varnames) robustcluster(varname) noadjust mgale(newvar) esr(newvars) schoenfeld(newvar) scaledsch(newvar) basehazard(newvar) basechazard(newvar) basesurv(newvar) {breslow |efron|exactm|exactp} cmdestimate noshow offsetlevel(#) maximize-options ] stphtest [,kmlogranktime(varname) plot(varname) detail graph-options ksm-options] stcoxisforusewithsurvival-time data;seehelpst.Youmust havestsetyourdatabeforeusingthiscommand; seehelpstset. Description ----------- stcoxestimates maximum-likelihood proportional hazards modelsonstdata. Options (manymore!) ------- nohrreports theestimated coefficients ratherthanhazardratios; i.e., bratherthanexp(b). Standard errorsandconfidence intervals are similarly transformed. Thisoptionaffects howresults aredisplayed, nothowtheyareestimated. 187 Ex.LeukemiaData .stcoxtrt Iteration 0:loglikelihood =-93.98505 Iteration 1:loglikelihood =-86.385606 Iteration 2:loglikelihood =-86.379623 Iteration 3:loglikelihood =-86.379622 Refining estimates: Iteration 0:loglikelihood =-86.379622 Coxregression --Breslow methodforties No.ofsubjects = 42 Numberofobs= 42 No.offailures = 30 Timeatrisk = 541 LRchi2(1) =15.21 Loglikelihood =-86.379622 Prob>chi2 =0.0001 ------------------------------------------------------------------- ----------- _t| _d|Haz.Ratio Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- trt|.2210887 .0905501 -3.685 0.000 .0990706 .4933877 ------------------------------------------------------------------- ----------- .stcoxtrt,nohr (sameiterations forlog-likelihood) Coxregression --Breslow methodforties No.ofsubjects = 42 Numberofobs= 42 No.offailures = 30 Timeatrisk = 541 LRchi2(1) =15.21 Loglikelihood =-86.379622 Prob>chi2 =0.0001 ------------------------------------------------------------------- ----------- _t| _d|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- trt|-1.509191 .4095644 -3.685 0.000 -2.311923 -.7064599 ------------------------------------------------------------------- ----------- 188 FittingPHmodelsinSAS-PROCPHREG Ex.Leukemiadata Title'CoxandOakesexample'; dataleukemia; inputweeksremisstrtmt; cards; 601 611 611 611 /*datafor6MPgroup*/ 711 901 etc 110 110 /*dataforplacebo group*/ 210 210 etc ; procphregdata=leukemia; modelweeks*remiss(0)=trtmt; title'CoxPHModelforleukemia data'; run; 189 PROCPHREGOutput: ThePHREGProcedure DataSet:WORK.LEUKEM Dependent Variable: FAILTIME TimetoRelapse Censoring Variable: FAIL Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 42 30 12 28.57 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 187.970 172.759 15.211with1DF(p=0.0001) Score . . 15.931with1DF(p=0.0001) Wald . . 13.578with1DF(p=0.0002) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DFEstimate ErrorChi-Square Chi-Square Ratio TRTMT 1-1.509191 0.40956 13.57826 0.0002 0.221 190 FittingPHmodelsinS-plus:coxphfunction Herearesomeofthedatainleuk.dat: tfx 110 110 210 210 310 ... 1901 2001 2211 2311 2501 3201 3201 3401 3501 leuk_read.table("leuk.dat",header=T) #specify Breslow handling ofties print(coxph(Surv(t,f) ~x,leuk,method="breslow")) #specify Efronhandling ofties(default) print(coxph(Surv(t,f) ~x,leuk)) 191 coxphOutput: Call: coxph(formula =Surv(t, f)~x,data=leuk,method="breslow") coefexp(coef) se(coef) z p x-1.51 0.221 0.41-3.680.00023 Likelihood ratiotest=15.2 on1df,p=0.0000961 n=42 Call: coxph(formula =Surv(t, f)~x,data=leuk) coefexp(coef) se(coef) z p x-1.57 0.208 0.412-3.810.00014 Likelihood ratiotest=16.4 on1df,p=0.0000526 n=42 192 Comparethiswiththelogranktest fromProcLifetest (Usingthe\Test"statement) TheLIFETEST Procedure RankTestsfortheAssociation ofFAILTIME withCovariates PooledoverStrata Univariate Chi-Squares fortheLOGRANKTest Test Standard Pr> Variable Statistic Deviation Chi-Square Chi-Square TRTMT 10.2505 2.5682 15.9305 0.0001 Notes: ²Thelogranktest=scoretestfromProcphreg! Ingeneral,thescoretestwouldbeforallofthevariables inthemodel,butinthiscase,wehaveonly\trtmt". ²Statadoesnotprovideascoretestinitsoutputfrom theCoxmodel.However,thestcoxcommand with thebreslow optionfortiesyieldsthesameLRtestas theCMH-versionlogranktestfromtheststest,cox command. 193 MoreNotes: ²TheCoxProportionalhazardsmodelhastheadvantage overasimplelogranktestofgivingusanestimate of the\riskratio"(i.e.,Á=¸1(t)=¸0(t)).Thisismore informativ ethanjustateststatistic, andwecanalso formcon¯dence intervalsfortheriskratio. ²Inthiscase,^Á=0:221,whichcanbeinterpreted tomean thatthehazardforrelapseamongpatientstreatedwith 6-MPislessthan25%ofthatforplacebopatients. ²Fromthestslistcommand inStataorProclifetest inSAS,wewereabletogetestimates oftheentiresur- vivaldistribution ^S(t)foreachtreatmen tgroup;wecan't immediately getthisfromourCoxmodelwithout fur- therassumptions.Whynot? 194 Adjustmentsforties Theproportionalhazardsmodelassumes acontinuoushaz- ard{tiesarenotpossible.Therearefourproposedmodi¯- cationstothelikelihoodtoadjustforties. (1)Cox's(1972)modi¯cation: \discrete" method (2)Peto-Breslowmethod (3)Efron's(1977)method (4)Exactmethod(Kalb°eischandPrentice) (5)Exactmarginalmethod(stata) Somenotation: ¿1;::::¿K theKordered, distinctdeathtimes dj thenumberoffailuresat¿j Hj the\history" oftheentiredataset,uptothe j-thdeathorfailuretime,including thetime ofthefailure,butnottheidentitiesofthedj whofailthere. ij1;:::ijdjtheidentitiesofthedjindividuals whofailat¿j 195 (1)Cox's(1972)modi¯cation: \discrete" method Cox'smethodassumes thatiftherearetiedfailuretimes, theytrulyhappenedatthesametime.Itisbasedona discretelikelihood. Thepartiallikelihoodis: L(¯)=KY j=1Pr(ij1;:::ijdjfailjdjfailat¿j;fromR) =KY j=1Pr(ij1;:::ijdjfailjinR(¿j)) P `2s(j;dj)Pr(`1;::::`djfailjinR(¿j)) =KY j=1exp(¯Zij1)¢¢¢exp(¯Zijdj) P `2s(j;dj)exp(¯Z`1)¢¢¢exp(¯Z`dj) =KY j=1exp(¯Sj) P `2s(j;dj)exp(¯Sj`) where ²s(j;dj)isthesetofallpossiblesetsofdjindividuals that canpossiblybedrawnfromtherisksetattime¿j ²SjisthesumoftheZ'sforallthedjindividuals who failat¿j ²Sj`isthesumoftheZ'sforallthedjindividuals inthe `-thsetdrawnoutofs(j;dj) 196 Whatdoesthisallmean??!! Let'smodifyourprevious simpleexample toincludeties. SimpleExample(withties) Group0:4+;6;8+;9;10+=)Zi=0 Group1:3;5;5+;6;8+=)Zi=1 Ordered failureNumberatriskLikelihoodContribution jtimeXiGroup0Group1e¯Sj=P `2s(j;dj)e¯Sj` 1355 e¯=[5+5e¯] 2544 e¯=[4+4e¯] 3642 e¯=[6+8e¯+e2¯] 4920 e0=2=1=2 Thetieoccursatt=6,whenR(¿j)=fZ=0:(6;8+;9;10+); Z=1:(6;8+)g.Ofthe³6 2´=15possiblepairsofsubjects atriskatt=6,thereare6pairsformedwherebotharefrom group0(Sj=0),8pairsformedwithoneineachgroup (Sj=1),and1pairsformedwithbothingroup1(Sj=2). Problem: Withlargenumbersofties,thedenominator can havemanymanytermsandbedi±culttocalculate. 197 (2)Breslowmethod:(default) BreslowandPetosuggested replacing thetermP `2s(j;dj)e¯Sj` inthedenominator bythetermµP `2R(¿j)e¯Z`¶dj,sothatthe followingmodi¯edpartiallikelihoodwouldbeused: L(¯)=KY j=1e¯Sj P `2s(j;dj)e¯Sj`¼KY j=1e¯Sj µP `2R(¿j)e¯Z`¶dj Justi¯cation: Supposeindividuals 1and2failfromf1;2;3;4gattime¿j. LetÁ(i)bethehazardratioforindividual i(compared to baseline). e¯Sj P `2s(j;dj)e¯Sj`=Á(1) Á(1)+Á(2)+Á(3)+Á(4)£Á(2) Á(2)+Á(3)+Á(4) +Á(2) Á(1)+Á(2)+Á(3)+Á(4)£Á(1) Á(1)+Á(3)+Á(4) ¼2Á(1)Á(2) [Á(1)+Á(2)+Á(3)+Á(4)]2 ThePeto(Breslow)approximation willbreakdownwhen thenumberoftiesarelargerelativetothesizeoftherisk sets,andthentendstoyieldestimates of¯whicharebiased toward0. 198 (3)Efron's(1977)method: Efronsuggested anevencloserapproximation tothediscrete likelihood: L(¯)=KY j=1e¯Sj à P `2R(¿j)e¯Z`+j¡1 djP `2D(¿j)e¯Z`!dj LiketheBreslowapproximation, Efron'smethodwillyield estimates of¯whicharebiasedtoward0whenthereare manyties. However,Allison(1995)recommends theEfronapproxima- tionsinceitismuchfasterthantheexactmethodsandtends toyieldmuchcloserestimates thanthedefaultBreslowap- proach. 199 (4)Exactmethod(Kalb°eischandPrentice): The\discrete" optionthatwediscussed in(1)isanexact methodbasedonadiscretelikelihood(assuming thattied eventstrulyAREtied). Thissecondexactmethodisbasedonthecontinuouslike- lihood,undertheassumption thatiftherearetiedevents, thatisduetotheimprecise natureofourmeasuremen t,and thattheremustbesometrueordering. Allpossibleorderings ofthetiedeventsarecalculated, and theprobabilities ofeacharesummed. Example with2tiedevents(1,2)fromriskset(1,2,3,4): e¯Sj P `2s(j;dj)e¯Sj`=e¯S1 e¯S1+e¯S2+e¯S3+e¯S4£e¯S2 e¯S2+e¯S3+e¯S4 +e¯S2 e¯S1+e¯S2+e¯S3+e¯S4£e¯S1 e¯S1+e¯S3+e¯S4 200 BottomLine:ImplicationsofTies (SeeAllison(1995),p.127-137) (1)Whentherearenoties,alloptionsgiveexactlythe sameresults. (2)Whenthereareonlyafewties,itwon'tmake muchdi®erence whichmethodisused.However,since theexactmethodswon'ttakemuchextracomputing time,youmightaswelluseoneofthem. (3)Whentherearemanyties(relativetothenumber atrisk),theBreslowoption(default) performspoorly (Farewell&Prentice,1980;Hsieh,1995).Bothofthe approximatemethods,BreslowandEfron,yieldcoe±- cientsthatareattenuated(biasedtoward0). (4)Thechoiceofwhichexactmethodtouseshould bebasedonsubstantivegrounds -arethetiedevent timestrulytied?...oraretheytheresultofimprecise measuremen t? (5)Computingtimeofexactmethodsismuchlonger thanthatoftheapproximatemethods.However,inmost casesitwillstillbelessthan30secondsevenfortheexact methods. (6)Bestapproximatemethod-theEfronapproxi- mationnearlyalwaysworksbetterthantheBreslow method,withnoincreaseincomputing time,sousethis optionifexactmethodsaretoocomputer-in tensive. 201 Example:Thefecundabilitystudy Womenwhohadrecentlygivenbirth(orhadtriedtoget pregnantforatleastayear)wereaskedtorecallhowlong ittookthemtobecomepregnant,andwhether ornotthey smokedduringthattime.Theoutcome ofinterestistimeto pregnancy (measured inmenstrual cycles). datafecund; input smoke cycle status count; cards; 0 1 1 198 0 2 1 107 0 3 1 55 0 4 1 38 0 5 1 18 0 6 1 22 .......................................... 1 10 1 1 1 11 1 1 1 12 1 3 1 12 0 7 ; procphreg; modelcycle*status(0) =smoke/ties=breslow; /*default */ freqcount; procphreg; modelcycle*status(0) =smoke/ties=discrete; freqcount; procphreg; modelcycle*status(0) =smoke/ties=exact; freqcount; procphreg; modelcycle*status(0) =smoke/ties=efron; freqcount; 202 SASOutputforFecundabilitystudy: AccountingforTies ******************************************************************* ******** TiesHandling: BRESLOW Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio SMOKE 1-0.329054 0.11412 8.31390 0.0039 0.720 ******************************************************************* ******** TiesHandling: DISCRETE Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio SMOKE 1-0.461246 0.13248 12.12116 0.0005 0.630 ******************************************************************* ******** TiesHandling: EXACT Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio SMOKE 1-0.391548 0.11450 11.69359 0.0006 0.676 ******************************************************************* ******** TiesHandling: EFRON Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio SMOKE 1-0.387793 0.11402 11.56743 0.0007 0.679 ******************************************************************* ******** Forthisparticulardataset,doesitseemlikeit wouldbeimportanttoconsiderthee®ectoftied failuretimes?Whichmethodwouldbebest? 203 StataCommandsforPHModelwithTies: Stataalsoo®ersfouroptionsforadjustmen tswithtieddata: ²breslow (default) ²efron ²exactp(sameasthe\discrete" optioninSAS) ²exactm-anexactmarginal likelihoodcalculation (di®erentthanthe\exact"optioninSAS) FecundabilityDataExample: .stcoxsmoker, efronnohr failure _d:status analysis time_t:cycle Iteration 0:loglikelihood =-3113.5313 Iteration 1:loglikelihood =-3107.3102 Iteration 2:loglikelihood =-3107.2464 Iteration 3:loglikelihood =-3107.2464 Refining estimates: Iteration 0:loglikelihood =-3107.2464 Coxregression --Efronmethodforties No.ofsubjects = 586 Numberofobs= 586 No.offailures = 567 Timeatrisk = 1844 LRchi2(1) =12.57 Loglikelihood =-3107.2464 Prob>chi2 =0.0004 ------------------------------------------------------------------- ----------- _t| _d|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- smoker|-.3877931 .1140202 -3.401 0.001 -.6112685 -.1643177 ------------------------------------------------------------------- ----------- 204 Aspecialcase:thetwo-sampleproblem Previously ,wederivedthelogranktestfromanintuitiveper- spective,assuming thatwehave(X01;±01):::(X0n0;±0n0)from group0and(X11;±11);:::;(X1n1;±1n1)fromgroup1. JustasaÂ2testforbinarydatacanbederivedfromalogistic model,wewillseeherethatthelogranktestcanbederived asaspecialcaseoftheCoxProportionalHazards model. First,let'sre-de¯ne ournotation intermsof(Xi;±i;Zi): (X01;±01);:::;(X0n0;±0n0)=)(X1;±1;0);:::;(Xn0;±n0;0) (X11;±11);:::;(X1n1;±1n1)=)(Xn0+1;±n0+1;1);:::;(Xn0+n1;±n0+n1;1) Inotherwords,wehaven0rowsofdata(Xi;±i;0)forthe group0subjects,thenn1rowsofdata(Xi;±i;1)forthe group1subjects. Usingtheproportionalhazardsformulation,wehave ¸(t;Z)=¸0(t)e¯Z Group0hazard: ¸0(t) Group1hazard: ¸0(t)e¯ 205 Thelog-partial likelihoodis: logL(¯)=log2 664KY j=1e¯Zj P `2R(¿j)e¯Z`3 775 =KX j=12 664¯Zj¡log[X `2R(¿j)e¯Z`]3 775 Takingthederivativewithrespectto¯,weget: U(¯)=@ @¯`(¯) =nX j=1±j2 664Zj¡P `2R(¿j)Z`e¯Z` P `2R(¿j)e¯Z`3 775 =nX j=1±j(Zj¡¹Zj) where¹Zj=P `2R(¿j)Z`e¯Z` P `2R(¿j)e¯Z` U(¯)iscalledthe\score" . 206 Aswediscussed earlierintheclass,oneusefulformofa likelihood-basedtestisthescoretest.Thisisobtained by usingthescoreU(¯)evaluatedatHoasateststatistic. Let'slookmorecloselyattheformofthescore: ±jZjobservednumberofdeathsingroup1at¿j ±j¹Zjexpectednumberofdeathsingroup1at¿j Why?UnderH0:¯=0,¹Zjissimplythenumberof individuals fromgroup1intherisksetattime¿j(callthis r1j),dividedbythetotalnumberintherisksetatthattime (callthisrj).Thus,¹Zjapproximates theprobabilit ythat giventhereisadeathat¿j,itisfromgroup1. Thus,thescorestatisticisoftheform: nX j=1(Oj¡Ej) Whenthereareties,thelikelihoodhastobereplaced byone thatallowsforties. InSASorStata: discrete/exactp !Mantel-Haenszel logranktest breslow!linearrankversionofthelogranktest 207 Ialreadyshowedyoutheequivalenceofthelinearranklo- granktestandtheBreslow(default) CoxPHmodelinSAS (p.24-25) HereistheoutputfromSASfortheleukemiadatausingthe method=discrete option: Logrank testwithproclifetest -stratastatement TestofEquality overStrata Pr> Test Chi-Square DFChi-Square Log-Rank 16.7929 10.0001 Wilcoxon 13.4579 10.0002 -2Log(LR) 16.4852 10.0001 ThePHREGProcedure DataSet:WORK.LEUKEM Dependent Variable: FAILTIME TimetoRelapse Censoring Variable: FAIL Censoring Value(s): 0 TiesHandling: DISCRETE Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 165.339 149.086 16.252with1DF(p=0.0001) Score . . 16.793with1DF(p=0.0001) Wald . . 14.132with1DF(p=0.0002) 208 MoreontheCoxPHmodel I.Con¯denceintervalsandhypothesistests {Twomethodsforcon¯denceintervals {Waldtestsandlikelihoodratiotests {Interpretationofparameterestimates {AnexamplewithrealdatafromanAIDS clinicaltrial II.Predictedsurvivalunderproportionalhazards III.PredictedmediansandP-yearsurvival 209 I.ConstructingCon¯denceintervalsandtestsfor theHazardRatio(seeH&L4.2,Collett3.4): Manysoftwarepackagesprovideestimates of¯,butthehaz- ardratioHR=exp(¯)isusuallytheparameter ofinterest. Wecanusethedeltamethodtogetstandard errorsfor exp(^¯): Var(dHR)=Var(exp(^¯))=exp(2^¯)Var(^¯) Constructingcon¯denceintervalsforexp(¯) Twooptions:(assumingthat¯isascalar) I.Usingse(exp^¯)obtained aboveviathedeltamethodas se(exp^¯)=r [Var(exp(^¯))],calculate theendpointsas: [L;U]=[dOR¡1:96se(dOR);dOR+1:96se(dOR)] II.Formacon¯dence intervalfor^¯,andthenexponentiate theendpoints. [L;U]=[e^¯¡1:96se(^¯);e^¯+1:96se(^¯)] Whichapproachdoyouthinkwouldbethemost preferable? 210 HypothesisTests: Foreachcovariateofinterest,thenullhypothesisis Ho:HRj=1,¯j=0 AWaldtest2oftheabovehypothesisisconstructed as: Z=^¯j se(^¯j)orÂ2=0 BB@^¯j se(^¯j)1 CCA2 Thistestfor¯j=0assumes thatallothertermsinthe modelareheld¯xed. Note:ifwehaveafactorAwithalevels,thenwewouldneed toconstruct aÂ2testwith(a¡1)df,usingateststatistic basedonaquadratic form: Â2 (a¡1)=c¯0 AVar(c¯A)¡1c¯A where¯A=(¯2;:::;¯a)0arethe(a¡1)coe±cientscor- respondingtoZ2;:::;Za(orZ1;:::;Za¡1,dependingonthe reference group). 2The¯rstfollowsanormaldistribution, andthesecondfollowsaÂ2with1df. STATAgivestheZstatistic,whileSASgivestheÂ2 1teststatistic(thep-values arealsogiven,anddon'tdependonwhichform,ZorÂ2,isprovided) 211 LikelihoodRatioTests: Supposethereare(p+q)explanatory variablesmeasured: Z1;:::;Zp;Zp+1;:::;Zp+q andproportionalhazardsareassumed. Considerthefollowingmodels: ²Model1:(containsonlythe¯rstpcovariates) ¸i(t;Z) ¸0(t)=exp(¯1Z1+¢¢¢+¯pZp) ²Model2:(containsall(p+q)covariates) ¸i(t;Z) ¸0(t)=exp(¯1Z1+¢¢¢+¯p+qZp+q) Thesearenestedmodels.Forsuchnestedmodels,wecan construct alikelihoodratiotestof H0:¯p+1=¢¢¢=¯p+q=0 as: Â2 LR=¡2· log(^L(1))¡log(^L(2))¸ UnderHo,thisteststatistic isapproximately distributed as Â2withqdf. 212 SomeexamplesusingtheStatastcoxcommand: Model1: .usemac .stsetmactime macstat .stcoxkarnofrifclari,nohr failure _d:macstat analysis time_t:mactime Coxregression --Breslow methodforties No.ofsubjects = 1151 Numberofobs=1151 No.offailures = 121 Timeatrisk = 489509 LRchi2(3) =32.01 Loglikelihood =-754.52813 Prob>chi2 =0.0000 ------------------------------------------------------------------- ---- _t| _d|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ---- karnof|-.0448295 .0106355 -4.215 0.000-.0656747 -.0239843 rif|.8723819 .2369497 3.6820.000 .4079691 1.336795 clari|.2760775 .2580215 1.0700.285-.2296354 .7817903 ------------------------------------------------------------------- ---- 213 Model2: .stcoxkarnofrifclaricd4,nohr failure _d:macstat analysis time_t:mactime Coxregression --Breslow methodforties No.ofsubjects = 1151 Numberofobs=1151 No.offailures = 121 Timeatrisk = 489509 LRchi2(4) =63.74 Loglikelihood =-738.66225 Prob>chi2 =0.0000 ------------------------------------------------------------------- ------ _t| _d|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ------ karnof|-.0368538 .0106652 -3.456 0.001 -.0577572 -.0159503 rif|.880338 .2371111 3.713 0.000 .4156089 1.345067 clari|.2530205 .2583478 0.979 0.327 -.253332 .7593729 cd4|-.0183553 .0036839 -4.983 0.000 -.0255757 -.0111349 ------------------------------------------------------------------- ------ 214 Notes: ²Ifweomitthenohroption,wewillgettheestimated hazardratioalongwith95%con¯dence intervalsusing MethodII(i.e.,formingaCIforthelogHR(beta),and thenexponentiatingthebounds) ------------------------------------------------------------------- ----- _t| _d|Haz.Ratio Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----- karnof|.9638171 .0102793 -3.456 0.001 .9438791 .9841762 rif|2.411715 .5718442 3.713 0.000 1.515293 3.838444 clari|1.28791 .3327287 0.979 0.327 .7762102 2.136936 cd4|.9818121 .0036169 -4.983 0.000 .9747486 .9889269 ------------------------------------------------------------------- ----- ²Wecanalsocompute thehazardratioourselves,byex- ponentiatingthecoe±cients: HRcd4=exp(¡0:01835)=0:98 WhyisthisHRsocloseto1,andyetstill highlysigni¯cant? WhatistheinterpretationofthisHR? ²Thelikelihoodratiotestforthee®ectofCD4istwice thedi®erence inminuslog-likelihoodsbetweenthetwo models: Â2 LR=2¤(754:533¡(738:66))=31:74 Howdoesthisteststatisticcompare totheWaldÂ2test? 215 ²Inthemacstudy,therewerethreetreatmen tarms(rif, clari,andtherif+clari combination). Because wehave onlyincluded therifandclarie®ectsinthemodel, thecombination therapyisthe\reference" group. ²Wecanconduct anoveralltestoftreatmen tusingthe testcommand inStata: .testrifclari (1)rif=0.0 (2)clari=0.0 chi2(2)=17.01 Prob>chi2=0.0002 fora2dfWaldchi-square testofwhetherbothtreatmen t coe±cientsareequalto0.Thistestcommand canbe usedtoconductanoveralltestforanynumberofe®ects. ²Thetestcommand canalsobeusedtotestwhether thereisadi®erence betweentherifandclaritreat- mentarms: .testrif=clari (1)rif-clari=0.0 chi2(1)=8.76 Prob>chi2=0.0031 216 SomeexamplesusingSASPROCPHREG procphregdata=alloi; modeldthtime*dthstat(0)=mlogrna cd4grp1 cd4grp2 combther /risklimits; cd4level: testcd4grp1, cd4grp2; title1'Proportional hazards regression modelfortimetoDeath'; title2'Baseline viralloadandCD4predictors'; procphregdata=alloi; modeldthtime*dthstat(0)=mlogrna cd4grp1 cd4grp2 combther decrs8incrs8 /risklimits; cd4level: testcd4grp1, cd4grp2; wk8resp: testdecrs8, incrs8; Notes: ²The\risklimits"optiononthemodelstatementprovides95% con¯denceintervalsusingMethodIIfrompage2.(i.e.,forming aCIforthelogHR(beta),andthenexponentiatingthebounds) ²The\test"statementhasthefollowingform: Label:testvarname1, varname2, ...,varnamek; forakdfWaldchi-squaretestofwhetherthekcoe±cientsare allequalto0. ²WecanusethesameapproachdescribedbyFreedmantoassess thee®ectsofintermediateendpoints(incrs8,decrs8)onthe treatmente®ect(i.e.,assesstheiruseassurrogatemarkers). Thepercentageoftreatmente®ectexplained, °,isestimated by: ^°=1¡^¯trt;M2 ^¯trt;M1 whereM1isthemodelwithouttheintermediateendpointand M2isthemodelwiththemarker. 217 OUTPUTFROMPROCPHREG (Model1) Proportional hazards regression modelfortimetoDeath Baseline viralloadandCD4predictors DataSet:WORK.ALLOI Dependent Variable: DTHTIME Timetodeath(days) Censoring Variable: DTHSTAT Deathstatus(1=died,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 690 89 601 87.10 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1072.543 924.167 148.376 with4DF(p=0.0001) Score . . 189.702 with4DF(p=0.0001) Wald . . 127.844 with4DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Variable DF Estimate Error Chi-Square Chi-Square MLOGRNA 1 0.833237 0.17808 21.89295 0.0001 CD4GRP1 1 2.364612 0.32436 53.14442 0.0001 CD4GRP2 1 1.171137 0.34434 11.56739 0.0007 COMBTHER 1 -0.497161 0.24389 4.15520 0.0415 218 OUTPUTFROMPROCPHREG ,continued Outputfrom\risklimits" and\test"statements Analysis ofMaximum Likelihood Estimates Conditional RiskRatioand 95%Confidence Limits Risk Variable Ratio Lower UpperLabel MLOGRNA 2.301 1.623 3.262logbaseline rna(rocheassay) CD4GRP1 10.640 5.634 20.093CD4<=100 CD4GRP2 3.226 1.643 6.335100<CD4<=200 COMBTHER 0.608 0.377 0.981Combination therapy withAZT/ddI/ddC/Nvp LinearHypotheses Testing Wald Pr> Label Chi-Square DFChi-Square CD4LEVEL 55.0794 2 0.0001 219 OUTPUTFROMPROCPHREG ,(Model2) Proportional hazards regression modelfortimetoDeath Baseline viralloadandCD4predictors DataSet:WORK.ALLOI Dependent Variable: DTHTIME Timetodeath(days) Censoring Variable: DTHSTAT Deathstatus(1=died,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 690 89 601 87.10 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1072.543 912.009 160.535 with6DF(p=0.0001) Score . . 198.537 with6DF(p=0.0001) Wald . . 132.091 with6DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Variable DF Estimate Error Chi-Square Chi-Square MLOGRNA 1 0.893838 0.18062 24.48880 0.0001 CD4GRP1 1 2.023005 0.33594 36.26461 0.0001 CD4GRP2 1 1.001046 0.34907 8.22394 0.0041 COMBTHER 1 -0.456506 0.24687 3.41950 0.0644 DECRS8 1 -0.410919 0.26383 2.42579 0.1194 INCRS8 1 -0.834101 0.32884 6.43367 0.0112 220 OUTPUTFROMPROCPHREG ,continued Outputfrom\risklimits" and\test"statements Analysis ofMaximum Likelihood Estimates Conditional RiskRatioand 95%Confidence Limits Risk Variable Ratio Lower UpperLabel MLOGRNA 2.444 1.716 3.483logbaseline rna(rocheassay) CD4GRP1 7.561 3.914 14.606CD4<=100 CD4GRP2 2.721 1.373 5.394100<CD4<=200 COMBTHER 0.633 0.390 1.028Combination therapy withAZT/ddI/ddC/Nvp DECRS8 0.663 0.395 1.112Decrease>=0.5 logrnaatweek8? INCRS8 0.434 0.228 0.827Increase>=50 CD4cells,week8? LinearHypotheses Testing Wald Pr> Label Chi-Square DFChi-Square CD4LEVEL 37.6833 20.0001 WK8RESP 10.4312 20.0054 Thepercentageoftreatmen te®ectexplained byincluding theRNAandCD4responsetotreatmen tbyWeek8is: ^°=1¡¡0:456 ¡0:497¼0:08 or8%.Thepercentageoftreatmen te®ectontimeto¯rst opportunistic infection ordeathismuchhigher(about24%). 221 II.PredictedSurvivalusingPH TheCoxPHmodelsaysthat¸i(t;Z)=¸0(t)exp(¯Z). Whatdoesthisimplyaboutthesurvivalfunction,Sz(t),for thei-thindividual withcovariatesZi? Forthebaseline (reference) group,wehave: S0(t)=e¡Rt 0¸0(u)du=e¡¤0(t) Thisisbyde¯nition ofasurvivalfunction (seeintronotes). Forthei-thpatientwithcovariatesZi,wehave: Si(t)=e¡Rt0¸i(u)du=e¡¤i(t) =e¡Rt0¸0(u)exp(¯Zi)du =e¡exp(¯Zi)Rt0¸0(u)du =" e¡Rt0¸0(u)du#exp(¯Zi) =[S0(t)]exp(¯Zi) (Thisusesthemathematical relationship [eb]a=eab) 222 Sayweareinterestedinthesurvivalpatternforsinglemales inthenursinghomestudy.Basedontheprevious formula, ifwehadanestimate forthesurvivalfunction intherefer- encegroup,i.e.,^S0(t),wecouldgetestimates ofthesurvival function foranysetofcovariatesZi. Howcanweestimatethesurvivalfunction, S0(t)? WecouldusetheKMestimator, butthereareafewdisad- vantagesofthatapproach: ²Itwouldonlyusethesurvivaltimesforobservationscon- tainedinthereference group,andnotalltherestofthe survivaltimes. ²Itwouldtendtobesomewhat choppy,sinceitwould re°ectthesmallersamplesizeofthereference group. ²It'spossiblethattherearenosubjectsinthedataset whoareinthe\reference" group(ex.saycovariatesare ageandsex;thereisnooneofage=0inourdataset). 223 Instead, wewilluseabaselinehazardestimator whichtakes advantageoftheproportional hazardsassumption togeta smootherestimate. ^Si(t)=[^S0(t)]exp(c¯Zi) Usingtheaboveformula,wesubstitutec¯basedon¯ttingthe CoxPHmodel,andcalculate ^S0(t)byoneofthefollowing approaches: ²Breslowestimator (Stata) ²Kalb°eisc h/Prenticeestimator (SAS) 224 (1)BreslowEstimator: ^S0(t)=exp¡^¤0(t) where^¤0(t)istheestimated cumulativebaselinehazard: ^¤(t)=X j:¿j<t0 BB@dj P k2R(¿j)exp(¯1Z1k+:::¯pZpk)1 CCA (2)Kalb°eisch/PrenticeEstimator ^S0(t)=Y j:¿j<t^®j where^®j;j=1;:::daretheMLE'sobtained byassum- ingthatS(t;Z)satis¯es S(t;Z)=[S0(t)]e¯Z=2 64Y j:¿j<t®j3 75e¯Z =Y j:¿j<t®e¯Z j 225 BreslowEstimator:furthermotivation TheBreslowestimator isbasedonextending theconcept oftheNelson-Aalen estimator totheproportional hazards model. Recallthatforasinglesamplewithnocovariates,theNelson- AalenEstimator ofthecumulativehazardis: ^¤(t)=X j:¿j<tdj rj wheredjandrjarethenumberofdeathsandthenumber atrisk,respectively,atthej-thdeathtime. Whentherearecovariatesandassuming thePHmodelabove, onecangeneralize thistoestimate thecumulativebaseline hazardbyadjusting thedenominator: ^¤(t)=X j:¿j<t0 BB@dj P k2R(¿j)exp(¯1Z1k+:::¯pZpk)1 CCA Heuristic: Theexpectednumberoffailuresin(t;t+±t)is dj¼±t£X k2R(t)¸0(t)exp(zk^¯) Hence, ±t£¸0(tj)¼dj P k2R(t)exp(zk^¯) 226 Kalb°eisch/PrenticeEstimator:furthermotivation Thismethodisanalogous totheKaplan-Meier Estimator. Consider adiscretetimemodelwithhazard(1¡®j)atthe j-thobserveddeathtime. (Note:weuse®j=(1¡¸j)tosimplifythealgebra!) Thus,forsomeone withz=0,thesurvivorshipfunction is S0(t)=Y j:¿j<t®j andforsomeone withZ6=0,itis: S(t;Z)=S0(t)e¯Z=2 64Y j:¿j<t®j3 75e¯Z =Y j:¿j<t®e¯Z j Thelikelihoodcontributions underthismodelare: ²forsomeone censored att:S(t;Z) ²forsomeone whofailsattj: S(t(j¡1);Z)¡S(tj;Z)=2 64Y k<j®j3 75e¯z [1¡®e¯Z j] Thesolution for®jsatis¯es: X k2Djexp(Zk¯) 1¡®exp(Zk¯) j=X k2Rjexp(Zk¯) (NotewhathappenswhenZ=0) 227 Obtaining ^S0(t)fromsoftwarepackages ²StataprovidestheBreslowestimator ofS0(t;Z),butnot predicted survivalsatspeci¯edcovariatevalues..... you havetoconstruct theseyourself ²SASusestheKalb°eisc h/Prenticeestimator ofthebase- linehazard,andcanprovideestimates ofsurvivalatar- bitraryvaluesofthecovariateswithalittlebitofpro- gramming. Inpractice, theyareincredibly close!(seeFleming and Harrington 1984,Communic ationsinStatistics ) 228 UsingStatatoPredictSurvival TheStatacommand basesurv calculates thepredicted sur- vivalvaluesforthereference group,i.e.,thosesubjectswith allcovariates=0. (1)BaselineSurvival: Toobtaintheestimated baseline survival^S0(t),follow theexample below(forthenursinghomedata): .usenurshome .stsetlosfail .stcoxmarried health, basesurv(prsurv) .sortlos .listlosprsurv 229 EstimatingtheBaselineSurvivalwithStata los prsurv 1. 1.99252899 2. 1.99252899 3. 1.99252899 4. 1.99252899 5. 1.99252899 . . . 22. 1.99252899 23. 2.98671824 24. 2.98671824 25. 2.98671824 26. 2.98671824 27. 2.98671824 28. 2.98671824 29. 2.98671824 30. 2.98671824 31. 2.98671824 32. 2.98671824 33. 2.98671824 34. 2.98671824 35. 2.98671824 36. 2.98671824 37. 2.98671824 38. 2.98671824 39. 2.98671824 40. 3.98362595 41. 3.98362595 . . . Statacreatesapredicted baseline survivalestimate for everyobservedeventtimeinthedataset, evenifthere areduplicates. 230 (2)PredictedSurvivalforSubgroups Toobtaintheestimated survival^Si(t)foranyothersub- group(i.e.,notthereference orbaseline group),follow theStatacommands below: .predict betaz,xb .gennewterm=exp(betaz) .genpredsurv=prsurv^newterm .sortmarried healthlos .listmarried healthlospredsurv 231 PredictingSurvivalforSubgroupswithStata married health lospredsurv 1. 0 2 1.9896138 8. 0 2 2.981557 11. 0 2 3.9772769 13. 0 2 4.9691724 16. 0 2 5.9586483 ........................................................... ..... 300. 0 3 1.9877566 302. 0 3 2.9782748 304. 0 3 3.9732435 305. 0 3 4.9637272 312. 0 3 5.9513916 ........................................................... ..... 768. 0 4 1.9855696 777. 0 4 2.9744162 779. 0 4 3.9685058 781. 0 4 4.9573418 785. 0 4 5.9428996 . . . 1468. 1 4 1.9806339 1469. 1 4 2.9657326 1472. 1 4 3.9578599 1473. 1 4 5.9239448 ........................................................... ..... 1559. 1 5 1.9771894 1560. 1 5 2.9596928 1562. 1 5 3.9504684 1564. 1 5 4.9331349 232 UsingSAStoPredictSurvival TheSAScommand BASELINE calculates thepredicted sur- vivalvaluesattheeventtimesforagivensetofcovariate values. (1)Togettheestimated baseline survival^S0(t),createa datasetwith0'sforvaluesofallcovariatesinthemodel (2)Togettheestimated survival^Si(t)foranyothersub- group(i.e.,notthereference orbaselinegroup),createa datasetwhichinputsthebaseline valuesofthecovari- atesforthesubgroup ofinterest. Foreithercase,wethensupplythecorrespondingdataset nametotheBASELINE command underPROCPHREG. Bygivingtheinputdatasetseverallines,eachcorresponding toadi®erentcombination ofcovariatevalues,wecancom- putepredicted survivalvaluesformorethanonegroupat once. 233 (1)BaselineSurvivalEstimate (notethatthebaselinesurvivalfunction doesnotcorrespond toanyobservationsinoursample,sincehealthstatusvalues rangefrom2-5) ***Estimating Baseline Survival Function underPH; datainrisks; inputmarried health; cards; 00 ; procphregdata=pop out=survres; modellos*fail(0)=married health; baseline covariates=inrisks out=outph survival=ps/nomean; procprintdata=outph; title1'Nursinghome data:Baseline Survival Estimate'; 234 EstimatingtheBaselineSurvivalwithSAS Nursinghome data:Baseline Survival Estimate OBSMARRIED HEALTH LOS PS 1 0 0 01.00000 2 0 0 10.99253 3 0 0 20.98672 4 0 0 30.98363 5 0 0 40.97776 6 0 0 50.97012 7 0 0 60.96488 8 0 0 70.95856 9 0 0 80.95361 10 0 0 90.94793 11 0 0 100.94365 12 0 0 110.93792 13 0 0 120.93323 14 0 0 130.92706 15 0 0 140.92049 16 0 0 150.91461 17 0 0 160.91017 18 0 0 170.90534 19 0 0 180.90048 20 0 0 190.89635 21 0 0 200.89220 22 0 0 210.88727 23 0 0 220.88270 . . . 235 (2)PredictedSurvivalEstimateforSubgroup ThefollowingSAScommands willgenerate thepredicted survivalprobabilit yforeachcombination ofcovariates,at everyobservedeventtimeinthedataset. ***Estimating Baseline Survival Function underPH; datainrisks; inputmarried health; cards; 02 05 12 15 ; procphregdata=pop out=survres; modellos*fail(0)=married health; baseline covariates=inrisks out=outph survival=ps/nomean; procprintdata=outph; title1'Nursinghome data:predicted survival bysubgroup'; 236 SurvivalEstimatesbyMaritalandHealthStatus Nursinghome data:Predicted Survival bySubgroup OBSMARRIED HEALTH LOS PS 1 0 2 01.00000 2 0 2 10.98961 3 0 2 20.98156 4 0 2 30.97728 ........................................................... ..... 171 0 21840.50104 172 0 21850.49984 ........................................................... ..... 396 0 5 01.00000 397 0 5 10.98300 398 0 5 20.96988 399 0 5 30.96295 ........................................................... ..... 474 0 5 780.50268 475 0 5 800.49991 ........................................................... ..... 791 1 2 01.00000 792 1 2 10.98605 793 1 2 20.97527 794 1 2 30.96955 ........................................................... ..... 897 1 21080.50114 898 1 21090.49986 ........................................................... ..... 1186 1 5 01.00000 1187 1 5 10.97719 1188 1 5 20.95969 1189 1 5 30.95047 ........................................................... ..... 1233 1 5 470.50519 1234 1 5 480.49875 237 Wecangetavisualpictureofwhatthepropor- tionalhazardsassumptionimpliesbylookingat thesefoursubgroups Subgroup Single, healthy Single, unhealth Married, healthy Married, unhealt0.00.10.20.30.40.50.60.70.80.91.0 LOS01002003004005006007008009001000 238 III.PredictedmediansandP-yearsurvival PredictedMedians Supposewewantto¯ndthepredicted mediansurvivalforan individual withaspeci¯edcombination ofcovariates(e.g.,a singlepersonwithhealthstatus5). Threepossibleapproaches: (1)Calculate themedianfromthesubsetofindividuals with thespeci¯edcovariatecombination (usingKMapproach) (2)Generate predicted survivalcurvesforeachcombination ofcovariates,andobtainthemedians directly OBSMARRIED HEALTH LOSPREDSURV 171 0 21840.50104 172 0 21850.49984 474 0 5 780.50268 475 0 5 800.49991 897 1 21080.50114 898 1 21090.49986 1233 1 5 470.50519 1234 1 5 480.49875 Recallthatpreviously wede¯nedthemedianasthe smallestvalueoftforwhich^S(t)·0:5,sothemedians fromabovewouldbe185,80,109,and48daysforsingle healthy,singleunhealth y,marriedhealthy,andmarried unhealth y,respectively. 239 (3)Generate thepredicted survivalcurvefromtheestimated baseline hazard,asfollows: Wewanttheestimated median(M)foranindividual withcovariatesZi.Weknow S(M;Z)=[S0(M)]e¯Zi=0:5 Hence,Msatis¯es(multiplying bothsidesbye¡¯Zi): S0(M)=[0:5]e¡¯Z Ex.Supposewewanttoestimate themediansurvival forasingleunhealth ysubjectfromthenursinghome data.Thereciprocalofthehazardratioforunhealth y (health=5) is:e¡0:165¤5=0:4373,(where^¯=0:165for healthstatus) So,wewantMsuchthatS0(M)=(0:5)0:4373=0:7385 Sothemedianforsingleunhealth ysubjectisthe73.8th percentileofthebaseline group. OBSMARRIED HEALTH LOSPREDSURV 79 0 0 780.74028 80 0 0 800.73849 81 0 0 810.73670 Sotheestimated medianwouldstillbe80days.Note:simi- larlogiccanbefollowedtoestimate otherquantilesbesides themedian. 240 EstimatingP-yearsurvival Supposewewantto¯ndtheP-yearsurvivalrateforanindi- vidualwithaspeci¯edcombination ofcovariates,^S(P;Zi) Foranindividual withZi=0,theP-yearsurvivalcanbe obtained fromthebaseline survivorshipfunction, ^S0(P) Forindividuals withZi6=0,itcanbeobtained as: ^S(P;Zi)=[^S0(P)]ec¯Zi Notes: ²Although Isay\P-year"survival,theunitsoftimeina particular datasetmaybedays,weeks,ormonths.The answerherewillbeinthesameunitsoftimeasthe originaldata. ²Ifc¯Ziispositive,thentheP-yearsurvivalrateforthei- thindividual willbelowerthanforabaselineindividual. Whyisthistrue? 241 ModelSelection inSurvivalAnalysis Supposewehaveacensored survivaltimethatwewantto modelasafunction ofa(possiblylarge)setofcovariates. Twoimportantquestions are: ²Howtodecidewhichcovariatestouse ²Howtodecideifthe¯nalmodel¯tswell Toaddressthesetopics,we'llconsider anewexample: SurvivalofAtlanticHalibut-Smithetal Survival TowDi® Length HandlingTotal Obs Time CensoringDurationinofFishTime log(catch) #(min) Indicator (min.) Depth(cm) (min.) ln(weight) 100353.0 1 301539 5 5.685 109111.0 1 100 544 29 8.690 11364.0 0 100 1053 4 5.323 116500.0 1 100 1044 4 5.323 ... Hosmer&Lemesho w Chapter 5: ModelDevelopment Chapter 6: Assessmen tofModelAdequacy (sections 6.1-6.2) 242 ProcessofModelSelection Collett(Section 3.6)hasanexcellentdiscussion ofvarious approachesformodelselection. Inpractice, modelselection proceedsthrough acombination of ²knowledgeofthescience ²trialanderror,common sense ²automatic variableselection procedures {forwardselection {backwardselection {stepwiseselection Manyadvocatetheapproachof¯rstdoingaunivariateanal- ysisto\screen" outpotentiallysigni¯can tvariablesforcon- sideration inthemultivariatemodel(seeCollett). Let'sstartwiththisapproach. 243 UnivariateKMplotsofAtlanticHalibutsurvival (continuousvariableshavebeendichotomized)Survival Distribution Function 0.00.10.20.30.40.50.60.70.80.91.0 SURVTIME0100200300400500600700800900100011001200 STRATA:TOWDUR=0TOWDUR=1 Survival Distribution Function 0.00.10.20.30.40.50.60.70.80.91.0 SURVTIME0100200300400500600700800900100011001200 STRATA:LENGTHGP=0LENGTHGP=1Survival Distribution Function 0.00.10.20.30.40.50.60.70.80.91.0 SURVTIME0100200300400500600700800900100011001200 STRATA:DEPTHGP=0DEPTHGP=1 Survival Distribution Function 0.00.10.20.30.40.50.60.70.80.91.0 SURVTIME0100200300400500600700800900100011001200 STRATA:HANDLGP=0HANDLGP=1Survival Distribution Function 0.00.10.20.30.40.50.60.70.80.91.0 SURVTIME0100200300400500600700800900100011001200 STRATA:LOGCATGP=0LOGCATGP=1 Whichcovariateslookliketheymightbeimportant? 244 AutomaticVariableselectionprocedures inStataandSAS StatisticalSoftware: ²Stata:swcommand beforecoxcommand ²SAS:selection= optiononmodelstatemen tof procphreg Options: (1)forward (2)backward (3)stepwise (4)bestsubset(SASonly,usingscoreoption) Onedrawbackoftheseoptionsisthattheycanonlyhandle variablesoneatatime.Whenmightthatbeadisadvantage? 245 Collett'sModelSelectionApproach Section3.6.1 Thisapproachassumes thatallvariablesareconsidered to beonanequalfooting,andthereisnoapriorireasonto includeanyspeci¯cvariables(liketreatmen t). Approach: (1)Fitaunivariatemodelforeachcovariate,andidentify thepredictors signi¯can tatsomelevelp1,say0:20. (2)Fitamultivariatemodelwithallsigni¯can tunivariate predictors, andusebackwardselection toeliminate non- signi¯can tvariablesatsomelevelp2,say0.10. (3)Starting with¯nalstep(2)model,consider eachofthe non-signi¯can tvariablesfromstep(1)usingforwardse- lection,withsigni¯cance levelp3,say0.10. (4)Do¯nalpruning ofmain-e®ects model(omitvariables thatarenon-signi¯can t,addanythataresigni¯can t), usingstepwise regression withsigni¯cance levelp4.At thisstage,youmayalsoconsider addinginteractions be- tweenanyofthemaine®ectscurrentlyinthemodel, underthehierarchicalprinciple. Collettrecommends usingalikelihoodratiotestforallvari- ableinclusion/exclusion decisions. 246 StataCommandforForwardSelection: ForwardSelection =)usepe(®)option,where®isthe signi¯cance levelforenteringavariableintothemodel. .usehalibut .stsetsurvtime censor .swcoxsurvtime towdurdepthlengthhandling logcatch, >dead(censor) pe(.05) beginwithemptymodel p=0.0000<0.0500 adding handling p=0.0000<0.0500 adding logcatch p=0.0010<0.0500 adding towdur p=0.0003<0.0500 adding length CoxRegression --entrytime0 Numberofobs=294 chi2(4) =84.14 Prob>chi2=0.0000 LogLikelihood =-1257.6548 PseudoR2=0.0324 ------------------------------------------------------------------- -------- survtime | censor|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- -------- handling |.0548994 .0098804 5.556 0.000 .0355341 .0742647 logcatch |-.1846548 .051015 -3.620 0.000 .2846423 -.0846674 towdur|.5417745 .1414018 3.831 0.000 .2646321 .818917 length|-.0366503 .0100321 -3.653 0.000 -.0563129 -.0169877 ------------------------------------------------------------------- -------- 247 StataCommandforBackwardSelection: BackwardSelection =)usepr(®)option,where®is thesigni¯cance levelforavariabletoremaininthemodel. .swcoxsurvtime towdurdepthlengthhandling logcatch, >dead(censor) pr(.05) beginwithfullmodel p=0.1991>=0.0500 removing depth CoxRegression --entrytime0 Numberofobs=294 chi2(4) =84.14 Prob>chi2=0.0000 LogLikelihood =-1257.6548 PseudoR2=0.0324 ------------------------------------------------------------------- ------- survtime | censor|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ------- towdur|.5417745 .1414018 3.831 0.000 .2646321 .818917 logcatch |-.1846548 .051015 -3.620 0.000-.2846423 -.0846674 length|-.0366503 .0100321 -3.653 0.000-.0563129 -.0169877 handling |.0548994 .0098804 5.556 0.000 .0355341 .0742647 ------------------------------------------------------------------- ------- 248 StataCommandforStepwiseSelection: StepwiseSelection =)usebothpe(:)andpr(:)options, withpr(:)>pe(:) .swcoxsurvtime towdurdepthlengthhandling logcatch, >dead(censor) pr(0.10) pe(0.05) beginwithfullmodel p=0.1991>=0.1000 removing depth CoxRegression --entrytime0 Numberofobs=294 chi2(4) =84.14 Prob>chi2=0.0000 LogLikelihood =-1257.6548 PseudoR2=0.0324 ------------------------------------------------------------------- ------ survtime | censor|Coef. Std.Err. zP>|z|[95%Conf.Interval] ---------+--------------------------------------------------------- ------ towdur|.5417745 .1414018 3.831 0.000.2646321 .818917 handling |.0548994 .0098804 5.556 0.000.0355341 .0742647 length|-.0366503 .0100321 -3.653 0.000-.0563129 -.0169877 logcatch |-.1846548 .051015 -3.620 0.000-.2846423 -.0846674 ------------------------------------------------------------------- ------ Itisalsopossibletodoforwardstepwiseregression byin- cludingbothpr(:)andpe(:)optionswithforward option 249 SASprogrammingstatementsformodelselection datafish; infile'fish.dat'; inputIDSURVTIME CENSORTOWDURDEPTHLENGTHHANDLING LOGCATCH; run; title'Survival ofAtlantic Halibut'; ***automatic variable selection procedures; procphregdata=fish; modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch /selection=stepwise slentry=0.1 slstay=0.1 details; title2'Stepwise selection'; run; procphregdata=fish; modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch /selection=forward slentry=0.1 details; title2'Forward selection'; run; procphregdata=fish; modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch /selection=backward slstay=0.1 details; title2'Backward selection'; run; procphregdata=fish; modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch /selection=score; title2'Bestsubsets selection'; run; 250 Finalmodelforstepwiseselectionapproach Survival ofAtlantic Halibut Stepwise selection ThePHREGProcedure Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio TOWDUR 10.007740 0.00202 14.68004 0.0001 1.008 LENGTH 1-0.036650 0.01003 13.34660 0.0003 0.964 HANDLING 10.054899 0.00988 30.87336 0.0001 1.056 LOGCATCH 1-0.184655 0.05101 13.10166 0.0003 0.831 Analysis ofVariables NotintheModel Score Pr> Variable Chi-Square Chi-Square DEPTH 1.6661 0.1968 Residual Chi-square =1.6661 with1DF(p=0.1968) NOTE:No(additional) variables metthe0.1levelforentryintothe model. Summary ofStepwise Procedure Variable Number Score Wald Pr> StepEntered Removed InChi-Square Chi-Square Chi-Square 1HANDLING 147.1417 .0.0001 2LOGCATCH 218.4259 .0.0001 3TOWDUR 311.0191 .0.0009 4LENGTH 413.4222 .0.0002 251 OutputfromPROCSAS\score"option NUMBEROF SCORE VARIABLES INCLUDED VARIABLES VALUE INMODEL 147.1417 HANDLING 129.9604 TOWDUR 112.0058 LENGTH 1 4.2185 DEPTH 1 1.4795 LOGCATCH --------------------------------- 265.6797 HANDLING LOGCATCH 259.9515 TOWDURHANDLING 256.1825 LENGTHHANDLING 251.6736 TOWDURLENGTH 247.2229 DEPTHHANDLING 232.2509 TOWDURLOGCATCH 230.6815 TOWDURDEPTH 216.9342 DEPTHLENGTH 214.4412 LENGTHLOGCATCH 2 9.1575 DEPTHLOGCATCH ------------------------------------- 376.8829 LENGTHHANDLING LOGCATCH 376.3454 TOWDURHANDLING LOGCATCH 375.5291 TOWDURLENGTHHANDLING 369.0334 DEPTHHANDLING LOGCATCH 360.0340 TOWDURDEPTHHANDLING 356.4207 DEPTHLENGTHHANDLING 355.8374 TOWDURLENGTHLOGCATCH 352.4130 TOWDURDEPTHLENGTH 334.7563 TOWDURDEPTHLOGCATCH 324.2039 DEPTHLENGTHLOGCATCH -------------------------------------------- 494.0062 TOWDURLENGTHHANDLING LOGCATCH 481.6045 DEPTHLENGTHHANDLING LOGCATCH 477.8234 TOWDURDEPTHHANDLING LOGCATCH 475.5556 TOWDURDEPTHLENGTHHANDLING 459.1932 TOWDURDEPTHLENGTHLOGCATCH ------------------------------------------------- 596.1287 TOWDURDEPTHLENGTHHANDLING LOGCATCH ------------------------------------------------------ 252 Bestmultivariatemodelforall3options Survival ofAtlantic Halibut BestMultivariate Model ThePHREGProcedure DataSet:WORK.FISH Dependent Variable: TIME Censoring Variable: CENSOR Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 294 273 21 7.14 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 2599.449 2515.310 84.140with4DF(p=0.0001) Score . . 94.006with4DF(p=0.0001) Wald . . 90.247with4DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio TOWDUR 10.007740 0.00202 14.68004 0.0001 1.008 LENGTH 1-0.036650 0.01003 13.34660 0.0003 0.964 HANDLING 10.054899 0.00988 30.87336 0.0001 1.056 LOGCATCH 1-0.184655 0.05101 13.10166 0.0003 0.831 253 Notes: ²Whenthehalibutdatawasanalyzed withtheforward, backwardandstepwiseoptions,thesame¯nalmodelwas reached.However,thiswillnotalwaysbethecase. ²Variablescanbeforcedintothemodelusingthelockterm optioninStataandtheinclude optioninSAS.Any variables thatyouwanttoforceinclusion ofmustbe listed¯rstinyourmodelstatemen t. ²StatausestheWaldtestforbothforwardandbackward selection, although ithasanoptiontousethelikelihood ratiotestinstead(lrtest).SASusesthescoretestto decidewhatvariablestoaddandtheWaldtestforwhat variablestoremove. ²Ifyou¯tarangeofmodelsmanually,youcanapplythe AICcriteriadescribedbyCollett: minimize AIC=¡2log(^L)+(®¤q) whereqisthenumberofunknownparameters inthe modeland®istypicallybetween2and6(theysuggest ®=3). Themodelisthenchosenwhichminimizes theAIC(sim- ilartomaximizing log-likelihood,butwithapenaltyfor numberofvariablesinthemodel) 254 Questions: ²Whenmightwewanttoforcecertainvariablesintothe model? (1)toexamine interactions (2)tokeepmaine®ectsinthemodel (3)tocalculate ascoretestforaparicular e®ect ²Woulditbepossibletogetdi®erent¯nalmodelsfrom SASandStata? ²Basedonwhatwe'veseeninthebehaviorofWaldtests, wouldSASorStatabemorelikelytoaddacovariateto amodelinaforwardselection model? ²IfweusetheAICcriteriawith®=3,howdoesthat compare tothelikelihoodratiotest? 255 Assessingoverallmodel¯t Howdoweknowifthemodel¯tswell? ²Alwayslookatunivariateplots(Kaplan-Meiers) ConstructaKaplan-Meier survivalplotforeachoftheimpor- tantpredictors,liketheonesshownatthebeginningofthese notes. ²Checkproportionalit yassumption (thiswillbethetopic ofthenextlecture) ²Checkresiduals! (a)generalized (Cox-Snell) (b)martingale (c)deviance (d)Schoenfeld (e)weightedSchoenfeld 256 Residuals forsurvivaldataareslightlydi®erentthanfor othertypesofmodels,duetothecensoring. Beforewestart talkingaboutresiduals, weneedanimportantbasicresult: InverseCDF: IfTi(thesurvivaltimeforthei-thindividual)has survivorshipfunctionSi(t),thenthetransformed randomvariableSi(Ti)(i.e.,thesurvivalfunction evaluatedattheactualsurvivaltimeTi)should befromauniformdistributionon[0;1],andhence ¡log[Si(Ti)]shouldbefromaunitexponentialdis- tribution Moremathematically: IfTi»Si(t) thenSi(Ti)»Uniform[0;1] and¡logSi(Ti)»Exponential (1) 257 (a)Generalized(Cox-Snell)Residuals : Theimplication ofthelastresultisthatifthemodeliscor- rect,theestimated cumulativehazardforeachindividual at thetimeoftheirdeathorcensoring shouldbelikeacensored samplefromaunitexponential.Thisquantityiscalledthe generalizedorCox-Snellresidual. Hereishowthegeneralized residualmightbeused.Suppose we¯taPHmodel: S(t;Z)=[S0(t)]exp(¯Z) or,intermsofhazards: ¸(t;Z)=¸0(t)exp(¯Z) =¸0(t)exp(¯1Z1+¯2Z2+¢¢¢+¯kZk) After¯tting,wehave: ²^¯1;:::;^¯k ²^S0(t) 258 So,foreachpersonwithcovariatesZi,wecanget ^S(t;Zi)=[^S0(t)]exp(¯Zi) Thisgivesapredicted survivalprobabilit yateachtimetin thedataset(seenotesfromtheprevious lecture). Thenwecancalculate ^¤i=¡log[^S(Ti;Zi)] Inotherwords,¯rstwe¯ndthepredictedsur- vivalprobabilityattheactualsurvivaltimefor anindividual,thenlog-transformit. 259 Example:Nursinghomedata Saywehave ²asinglemale ²withactualduration ofstayof941days(Xi=941) Wecompute theentiredistribution ofsurvivalprobabilities forsinglemales,andobtain^S(941)=0:260. ¡log[^S(941;singlemale)]=¡log(0:260)=1:347 Werepeatthisforeveryoneinourdataset. Theseshouldbe likeacensored samplefromanexponential(1)distribution ifthemodel¯tsthedatawell. Basedonthepropertiesofaunitexponentialmodel ²plotting¡log(^S(t))vstshouldyieldastraightline ²plotting log[¡logS(t)]vslog(t)shouldyieldastraight linethrough theoriginwithslope=1. Toconvinceyourselfofthis,startwithS(t)=e¡¸tand calculate log[¡logS(t)].Whatdoyougetfortheslopeand intercept? (Note:thisdoesnotnecessarily meanthattheunderlying distribution oftheoriginalsurvivaltimesisexponential!) 260 ObtainingthegeneralizedresidualsfromStata ²FitaCoxPHmodelwiththestcoxcommand, along withthemgale(newvar)option ²Usethepredict command withthecsnelloption ²De¯neasurvivaldatasetusingtheCox-Snellresiduals asthe\pseudo" failuretimes ²Calculate theestimated KMsurvival ²Takethelog[¡log(S(t))]basedontheabove ²Generate thelogoftheCox-Snellresiduals ²Graphlog[¡logS(t)]vslog(t) .stcoxtowdurhandling lengthlogcatch, mgale(mg) .predict csres,csnell .stsetcsrescensor .stslist .stsgensurvcs=s .genlls=log(-log(survcs)) .genloggenr=log(csres) .graphllsloggenr 261 LLS -6-5-4-3-2-1012 Log of SURVIVAL-5 -4 -3 -2 -1 0 1 Allisonstates\Cox-Snellresiduals...arenotveryinformativefor Coxmodelsestimatedbypartiallikelihood."Heinsteadprefers devianceresiduals(later). 262 ObtainingthegeneralizedresidualsfromSAS Thegeneralizedresiduals canbeobtained fromSAS after¯ttingaPHmodelusingtheoutputstatemen twith thelogsurv option. procphregdata=fish; modelsurvtime*censor(0) =towdurhandling logcatch length; outputout=phres logsurv=genres; ***takenegative logPr(survival) ateachpersons survtime; dataphres; setphres; genres=-genres; ***Nowwetreatthegeneralized residuals astheinputdataset; ***toevaluate whether theassumption ofanexponential; ***distribution isappropriate; proclifetest data=phres outsurv=survres; timegenres*censor(0); datasurvres; setsurvres; lls=log(-log(survival)); loggenr=log(genres); procgplotdata=survres; plotlls*loggenr; run; 263 (b)MartingaleResiduals (seeFleming andHarrington, p.164) Martingale residuals arede¯nedforthei-thindividual as: ri=±i¡^¤(Ti) Properties: ²ri'shavemean0 ²rangeofri'sisbetween¡1and1 ²approximately uncorrelated (inlargesamples) ²Interpretation: -theresidualricanbeviewedasthe di®erence betweentheobservednumberofdeaths(0or 1)forsubjectibetweentime0andTi,andtheexpected numbersbasedonthe¯ttedmodel. 264 Themartingaleresiduals canbeobtained fromStata usingthemgaleoptionshownpreviously . Oncethemartingale residualiscreated,youcanplotitversus thepredicted logHR(i.e.,¯Zi),oranyoftheindividual covariates. .stcoxtowdurhandling lengthlogcatch, mgale(mg) .predict betaz=xb .graphmgbetaz .graphmglogcatch .graphmgtowdur .graphmghandling .graphmglength 265 Themartingaleresiduals canbeobtained fromSAS after¯ttingaPHmodelusingtheoutputstatemen twith theresmart option. Onceyouhavethem,youcan ²plotagainstpredicted values ²plotagainstcovariates procphregdata=fish; modelsurvtime*censor(0) =towdurhandling logcatch length; outputout=phres resmart=mres xbeta=xb; procgplotdata=phres; plotmres*xb; /*predicted values*/ plotmres*towdur; plotmres*handling; plotmres*logcatch; plotmres*length; run; Allisonstillprefersthedeviance residuals (next) 266 MartingaleResiduals M a r t i n g a l e R e s i d u a l-6-5-4-3-2-101 Towing Duration0102030405060708090100110120M a r t i n g a l e R e s i d u a l-6-5-4-3-2-101 Length of fish (cm)20 30 40 50 60 M a r t i n g a l e R e s i d u a l-6-5-4-3-2-101 Log(weight) of catch2 3 4 5 6 7 8M a r t i n g a l e R e s i d u a l-6-5-4-3-2-101 Handling time0 10 20 30 40 M a r t i n g a l e R e s i d u a l-6-5-4-3-2-101 Linear Predictor-3 -2 -1 0 267 (c)DevianceResiduals Oneproblem withthemartingale residuals isthattheytend tobeasymmetric. Asolution istousedevianceresiduals .Forpersoni, thesearede¯nedasafunction ofthemartingale residuals (ri): ^Di=sign(^ri)r ¡2[^ri+±ilog(±i¡^ri)] InStata,thedeviance residuals aregenerated usingthesame approachastheCox-Snellresiduals. .stcoxtowdurhandling lengthlogcatch, mgale(mg) .predict devres, deviance andthentheycanbeplottedversusthepredicted log(HR) ortheindividual covariates,asshownfortheMartingale residuals. InSAS,justuseresdev optioninsteadofresmart. Deviance residuals behavemuchlikeresiduals fromOLSre- gression(i.e.,mean=0, s.d.=1). Theyarenegativeforobser- vationswithsurvivaltimesthataresmallerthanexpected. 268 DevianceResiduals D e v i a n c e R e s i d u a l-3-2-10123 Towing Duration0102030405060708090100110120D e v i a n c e R e s i d u a l-3-2-10123 Length of fish (cm)20 30 40 50 60 D e v i a n c e R e s i d u a l-3-2-10123 Log(weight) of total catch2 3 4 5 6 7 8D e v i a n c e R e s i d u a l-3-2-10123 Handling time0 10 20 30 40 D e v i a n c e R e s i d u a l-101234 Linear Predictor0.00.10.20.30.40.50.60.70.80.9 269 (d)SchoenfeldResiduals Thesearede¯nedateachobservedfailuretimeas: rs ij=Zij(ti)¡¹Zj(ti) Notes: ²representthedi®erence betweentheobservedcovariate andtheaverageovertherisksetatthattime ²calculated foreachcovariate ²notde¯nedforcensored failuretimes. ²usefulforassessing timetrendorlackorproportionalit y, basedonplotting versuseventtime ²sumtozero,haveexpectedvaluezero,andareuncorre- lated(inlargesamples) InStata,theSchoenfeldresiduals aregenerated inthestcox command itself,usingtheschoenf( newvar(s) )option: .stcoxtowdurhandling lengthlogcatch, schoenf(towres handres lenres logres) .graphtowressurvtime InSAS,addtotheoutputline RESSCH=name1 name2...namek foruptokregressors inthemodel. 270 SchoenfeldResiduals -60-50-40-30-20-1001020304050 Survival Time0100200300400500600700800900100011001200-20-1001020 Survival Time0100200300400500600700800900100011001200 -3-2-1012345 Survival Time0100200300400500600700800900100011001200-20-100102030 Survival Time0100200300400500600700800900100011001200 271 (e)WeightedSchoenfeldResiduals Theseareactually usedmoreoftenthantheprevious un- weightedversion,becausetheyaremorelikethetypicalOLS residuals (i.e.,symmetric around0). Theyarede¯nedas: rw ij=ncVrs ij wherecVistheestimated varianceof^¯.Theweightedresid- ualscanbeusedinthesamewayastheunweightedonesto assesstimetrendsandlackofproportionalit y. InStata,usethecommand: .stcoxtowdurlengthlogcatch handling depth,scaledsch(towres2 >lenres2 logres2 handres2 depres2) .graphlogres2 survtime InSAS,addtotheoutputline WTRESSCH=name1 name2...namek foruptokregressors inthemodel. 272 WeightedSchoenfeldResiduals -0.08-0.07-0.06-0.05-0.04-0.03-0.02-0.010.000.010.020.030.040.050.060.07 Survival Time0100200300400500600700800900100011001200-0.4-0.3-0.2-0.10.00.10.20.30.4 Survival Time0100200300400500600700800900100011001200 -2-10123 Survival Time0100200300400500600700800900100011001200-0.4-0.3-0.2-0.10.00.10.20.30.40.5 Survival Time0100200300400500600700800900100011001200 273 UsingResidualplotstoexplorerelationships Ifyoucalculate martingale ordeviance residuals withoutany covariatesinthemodelandthenplotagainstcovariates,you obtainagraphical impression oftherelationship betweenthe covariateandthehazard. InSplus,itiseasytodothis(alsopossibleinstatausingthe \estimate" option) **readinthedataset andfitacoxPHmodel fish_read.table('fish.data',header=T) x_fish$towdur fishres_coxreg(fish$time, fish$censor, x,resid="martingale",iter.max=0) **the2commands belowsetupthepostscript file,with4graphs postscript("fishres.plt",horizontal=F,height=10,width=7) par(mfrow=c(2,2),oma=c(0,0,2,0)) **plotthemartingale residuals vseachoftheothercovariates **andaddalowesssmoothed fittotheplot plot(fish$depth, fishres$resid, xlab="depth") lines(lowess(fish$depth,fishres$resid,iter=0)) plot(fish$length, fishres$resid, xlab="length") lines(lowess(fish$length,fishres$resid,iter=0)) plot(fish$handling, fishres$resid, xlab="handling") lines(lowess(fish$handling,fishres$resid,iter=0)) plot(fish$logcatch, fishres$resid, xlab="logcatch") lines(lowess(fish$logcatch,fishres$resid,iter=0)) 274 SplusPlotsofMartingaleResidualsforCoxModel containingonlytowingdurationasapredictor, vsothercovariates ·· ··· ····· · ·· ··· ··· ··· ·········· ··· ··· ·· ··· ··· · ······ ·· ···· ····················· · ··· ·· ·· · ···· ·· ····· ·· ·· · ·· ······ ·· · · ··· ·· ···· ···· ·· ··· ·· ·· ······· ···· ···· ····· · ·· ·· · ·· ·· · ·· ······················· ····· ··· ······ ········ ············· ·· ·· ······· ······ ······ ······ ···· ·· ··· · ··· depthfishres$resid 0103050-4-3-2-101 ··· ·· ··· · ····· · ·· ····· ··· ··· ········· · ··· ··· ·· ··· ··· · ······ ·· ···· ········ ·· ··· ·········· · ··· ·· ··· · ····· ·· ····· ·· ·· · ·· ······· ·· ·· · ··· ·· ···· ···· ·· ··· ·· ·· ······· ···· ···· ····· · ·· ·· · ·· ·· · ··· ············· ······ ······· ····· ··· ········ ········ · ···· ········· ·· ·· ····· ···· ······ ······ ···· ··· ···· ·· ··· · ··· Lengthfishres$resid 304050-4-3-2-101 ·· ·· · ····· · ·· ··· ··· ··· ·········· ··· ··· ·· ··· ··· · ······ ·· ···· ····················· · ··· ·· ·· · ···· ·· ····· ·· ·· · ·· ······ ·· · · ··· ·· ···· ···· ·· ··· ·· ·· ······· ···· ···· ····· · ·· ·· · ·· ·· · ·· ······················· ····· ··· ······ ········ ············· ·· ·· ······· ······ ······ ······ ···· ·· ··· · ··· logcatchfishres$resid 345678-4-3-2-101 ··· ··· ····· · ·· ··· ··· ··· ·········· ··· ··· ·· ··· ··· · ······ ·· ···· ······················ · ··· ·· ·· · ···· ·· ····· ·· ·· · ·· ······ ·· ·· · ··· ·· ···· ···· ·· ··· ·· ·· ······· ···· ···· ····· · ·· ·· · ·· ·· · ·· ························ ····· ··· ······ ········ ············· ·· ·· ······· ······ ······ ······· ···· ·· ·· · · ··· handlingfishres$resid 0102030-4-3-2-101 275 (f)Deletiondiagnostics Deletion diagnostics arede¯nedgenerally as: ±i=^¯¡^¯(i) Inotherwords,theyarethedi®erence betweentheestimated regression coe±cientusingallobservationsandthatwithout thei-thindividual. Thiscanbeusefulforassessing thein- °uence ofanindividual. InSASPROCPHREG, weusethedfbetaoption: (Notethatthereisaseparatedfbetacalculated foreachof thepredictors.) procphregdata=fish; modelsurvtime*censor(0)=towdur handling logcatch length; idid; outputout=phinfl dfbeta=dtow dhanddlogcdlength ld=lrchange; procunivariate data=phinfl; vardtowdhanddlogcdlength lrchange; idid; run; Theprocunivariateprocedurewillsupplythe5smallest val- uesandthe5largestvalues.The\id"statemen tmeansthat thesewillbelabeledwiththevalueofidfromthedataset. 276 (g)OtherIn°uencediagnostics Otherin°uencediagnostics: TheLDoptionisanothermethodforcheckingin°uence. It calculates howmuchthelog-likelihood(x2)wouldchangeif thei-thpersonwasremovedfromthesample. LDi=2· logL(c¯)¡logL(c¯¡i)¸ c¯=MLEforallparameters witheveryoneincluded c¯¡i=MLEwithi-thsubjectomitted Again,theprocunivariateprocedureinSASwillidentify theobservationswiththelargestandsmallest valuesofthe lrchange diagnostic measure. 277 Canweimprovethemodel? Theplotsappeartohavesomestructure, whichindicatethat wecouldbeleavingsomething out.Itisalwaysagoodidea tocheckforinteractions: Inthiscase,thereareseveralimportantinteractions. Iused abackwardselection modelforcingallmaine®ectstobe included, andconsidering allpairwise interactions. Hereare theresults: Parameter Standard Wald Pr>Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio TOWDUR 1-0.075452 0.01740 18.79679 0.0001 0.927 DEPTH 10.123293 0.06400 3.71107 0.0541 1.131 LENGTH 1-0.077300 0.02551 9.18225 0.0024 0.926 HANDLING 10.004798 0.03221 0.02219 0.8816 1.005 LOGCATCH 1-0.225158 0.07156 9.89924 0.0017 0.798 TOWDEPTH 10.002931 0.0004996 34.40781 0.0001 1.003 TOWLNGTH 10.001180 0.0003541 11.10036 0.0009 1.001 TOWHAND 10.001107 0.0003558 9.67706 0.0019 1.001 DEPLNGTH 1-0.006034 0.00136 19.77360 0.0001 0.994 DEPHAND 1-0.004104 0.00118 12.00517 0.0005 0.996 Interpretation: Handling alonedoesn'tseemtoa®ectsurvival,unlessitis combinedwithalongertowingduration orshallowertrawl- ingdepths. 278 Analternativemodelingstrategywhenwehave fewercovariates Withadatasetwithonly5maine®ects,itwouldmakesense toconsider interactions fromthestart.Howmanywould therebe? ²Fitmodelwithallmaine®ectsandpairwise interactions ²Thenusebackwardselection toeliminate non-signi¯can t pairwise interactions (remembertoforcethemaine®ects intothemodelatthisstage) ²Oncenon-signi¯can tpairwise interactions havebeenelim- inated,youcouldconsider backwardsselection toelim- inateanynon-signi¯can tmaine®ectsthatarenotin- volvedinremaining interaction terms ²Afterobtaining ¯nalmodel,useresiduals tocheck¯tof model. 279 Assessing thePHAssumption Sofar,we'vebeenconsidering thefollowingCoxPHmodel: ¸(t;Z)=¸0(t)exp(¯Z) =¸0(t)exp(X¯jZj) where¯jistheparameter forthethej-thcovariate(Zj). Importantfeaturesofthismodel: (1)thebaselinehazarddependsont,butnotonthecovari- atesZ1;:::;Zp (2)thehazardratio,i.e.,exp(¯Z),dependsonthecovariates Z=(Z1;:::;Zp),butnotontimet. Assumption (2)iswhatledustocallthisaproportional hazards model.That'sbecausewecouldtaketheratioof thehazardsfortwoindividuals withcovariatesZiandZi0, andwriteitasaconstantintermsofthecovariates. 280 ProportionalHazardsAssumption HazardRatio: ¸(t;Zi) ¸(t;Zi0)=¸0(t)exp(¯Zi) ¸0(t)exp(¯Zi0) =exp(¯Zi) exp(¯Zi0) =exp[¯(Zi¡Zi0)] =exp[X¯j(Zij¡Zi0j)]=µ Inthelastformula,Zijisthevalueofthej-thcovariatefor thei-thindividual. Forexample,Z42mightbethevalueof gender (0or1)forthethe4-thperson. Wecanalsowritethehazardforthei-thpersonasaconstant timesthehazardforthei0-thperson: ¸(t;Zi)=µ¸(t;Zi0) Thus,theHRbetweentwotypesofindividuals isconstant (i.e.,=µ)overtime.Thesearemathematical waysofstating theproportionalhazardsassumption. 281 Thereareseveraloptionsforcheckingtheassumption ofpro- portionalhazards: I.Graphical (a)Plotsofsurvivalestimates fortwosubgroups (b)Plotsoflog[¡log(^S)]vslog(t)fortwosubgroups (c)PlotsofweightedSchoenfeldresiduals vstime (d)Plotsofobservedsurvivalprobabilities versusex- pectedunderPHmodel(seeKleinbaum,ch.4) II.Useofgoodnessof¯ttests-wecanconstruct agoodness-of-¯t testbasedoncomparing theobserved survivalprobabilit y(fromstslist)withtheexpected (fromstcox)undertheassumption ofproportionalhaz- ards-seeKleinbaumch.4 III.Includinginteractiontermsbetweenacovari- ateandt(time-dep endentcovariates) 282 Howdoweinterprettheabove? Kleinbaum(andothertexts)suggestastrategy ofassuming thatPHholdsunlessthereisverystrongevidence tocounter thisassumption: ²estimated survivalcurvesarefairlyseparated, thencross ²estimated logcumulativehazardcurvescross,orlook veryunparallel overtime ²weightedSchoenfeldresiduals clearlyincreaseordecrease overtime(youcould¯taOLSregression lineandseeif theslopeissigni¯can t) ²testfortime£covariateinteraction termissigni¯can t (thisrelatestotime-dep endentcovariates) IfPHdoesn'texactlyholdforaparticular covariatebutwe ¯tthePHmodelanyway,thenwhatwearegettingissort ofanaverageHR,averagedovertheeventtimes. Inmostcases,thisisnotsuchabadestimate. Allisonclaims thattoomuchemphasis isputontestingthePHassumption, andnotenoughtootherimportantaspectsofthemodel. 283 Implicationsofproportionalhazards Consider aPHmodelwithasinglecovariate,Z: ¸(t;Z)=¸0(t)e¯Z Whatdoesthisimplyfortherelationbetweenthesurvivor- shipfunctions atvariousvaluesofZ? UnderPH, log[¡log[S(t;Z)]]=log[¡log[S0(t)]]+¯Z Ingeneral,wehavethefollowingrelationship: ¤i(t)=Zt 0¸i(u)du =Zt 0¸0(u)exp(¯Zi)du =exp(¯Zi)Zt 0¸0(u)du =exp(¯Zi)¤0(t) Thismeansthattheratioofthecumulativehazardsisthe sameastheratioofhazardrates: ¤i(t) ¤0(t)=exp(¯Zi)=exp(¯1Z1i+¢¢¢+¯pZpi) 284 Usingtheaboverelationship, wecanshowthat: ¯Zi=log0 B@¤i(t) ¤0(t)1 CA =log¤i(t)¡log¤0(t) =log[¡logSi(t)]¡log[¡logS0(t)] solog[¡logSi(t)]=log[¡logS0(t)]+¯Zi Thus,toassessifthehazards areactually proportional to eachotherovertime(usinggraphical optionI(b)) ²calculate KaplanMeierCurvesforvariouslevelsofZ ²compute log[¡log(^S(t;Z))](i.e.,logcumulativehazard) ²plotvslog-time toseeiftheyareparallel(linesorcurves) Note:IfZiscontinuous,breakintocategories. 285 Question:Whynotjustcomparetheunderlying hazardratestoseeiftheyareproportional? Here'stwosimulatedexamples withhazardswhicharetruly proportionalbetweenthetwogroups: Weibull-typehazard:U-shapedhazard: Plots of hazard function vs time Simulated data with HR=2 for men vs women GenderWomen MenHAZARD 0.0000.0020.0040.0060.0080.010 Length of Stay (days)010020030040050060070080090010001100Plots of hazard function vs time Simulated data with HR=2 for men vs women GenderWomen MenHAZARD 0.0000.0020.0040.0060.0080.010 Length of Stay (days)010020030040050060070080090010001100 Reason1:It'shardtoeyeballthese¯guresand seethatthehazardratesareproportional-it wouldbeeasiertolookforaconstantshiftbe- tweenlines. 286 Reason2:Estimatedhazardratestendtobe moreunstablethanthecumulativehazardrate Consider thenursinghomeexample (wherewethinkPHis reasonable). Ifwegroupthedataintointervalsandcalculate thehazardrateusingactuarial method,wegettheseplots: 200dayintervals:100dayintervals: Plots of hazard function vs time GenderWomen Men0.0000.0010.0020.0030.0040.0050.006 Length of Stay (days)01002003004005006007008009001000Plots of hazard function vs time GenderWomen Men0.0000.0010.0020.0030.0040.0050.0060.0070.0080.009 Length of Stay (days)01002003004005006007008009001000 50dayintervals:25dayintervals: Plots of hazard function vs time GenderWomen Men0.0000.0020.0040.0060.0080.0100.012 Length of Stay (days)010020030040050060070080090010001100Plots of hazard function vs time GenderWomen Men0.0000.0020.0040.0060.0080.0100.0120.014 Length of Stay (days)010020030040050060070080090010001100 287 Incontrast,thelogcumulativehazardplotsare easiertointerpretandtendtogivemorestable estimates Ex:NursingHome-genderandmaritalstatus proclifetest data=pop outsurv=survres; timelos*fail(0); stratagender; formatgendersexfmt.; title'Duration ofLengthofStayinnursing homes'; datasurvres; setsurvres; labellog_los='Log(Length ofstayindays)'; iflos>0thenlog_los=log(los); ifsurvival<1 thenlls=log(-log(survival)); procgplotdata=survres; plotlls*log_los=gender; formatgendersexfmt.; title2'Plotsoflog-log KMversuslog-time'; run; Thestatementsformaritalstatusaresimilar,substituting married forgender. Note:Thisisequivalenttocomparing plotsofthelogcumu- lativehazard,log(^¤(t)),betweenthecovariatelevels,since ¤(t)=Zt 0¸(u;Z)du=¡log[S(t)] 288 Assessmentofproportionalhazardsforgender andmaritalstatusinnursinghomedata(Mor- ris) Plots of log-log KM versus log-time GenderWomen MenLLS -6-5-4-3-2-101 Log(Length of stay in days)0 1 2 3 4 5 6 Plots of log-log KM versus log-time Marital Status SingleMarriedLLS -6-5-4-3-2-101 Log(Length of stay in days)0 1 2 3 4 5 6 289 Assessingproportionalitywithseveralcovariates Ifthereisenoughdataandyouonlyhaveacoupleofcovari- ates,createanewcovariatethattakesadi®erentvaluefor everycombination ofcovariatevalues. Example: Healthstatusandgenderfornursinghome datapop; infile'ch12.dat'; inputlosagerxgendermarried healthfail; ifgender=0 andhealth=2 thenhlthsex=1; ifgender=1 andhealth=2 thenhlthsex=2; ifgender=0 andhealth=5 thenhlthsex=3; ifgender=1 andhealth=5 thenhlthsex=4; procformat; valuehsfmt 1='Healthier Women' 2='Healthier Men' 3='Sicker Women' 4='Sicker Men'; proclifetest data=pop outsurv=survres; timelos*fail(0); stratahlthsex; formathlthsex hsfmt.; title'Length ofStayinnursing homes'; datasurvres; setsurvres; labellog_los='Log(Length ofstayindays)'; labelhlthsex='Health/Gender Status'; iflos>0thenlog_los=log(los); ifsurvival<1 lls=log(-log(survival)); procgplotdata=survres; plotlls*log_los=hlthsex; formathlthsex hsfmt.; title2'Plotsoflog-log KMversuslog-time'; run; 290 Log[-log(surviv al)]PlotsforHealthstatus*gender Plots of log-log KM versus log-time Health/Gender Status Healthier Women Healthier Men Sicker Women Sicker MenLLS -5-4-3-2-101 Log(Length of stay in days)0 1 2 3 4 5 6 Iftherearetoomanycovariates(ornotenoughdata)forthis, thenthereisawaytotestproportionalit yforeachvariable, oneatatime,usingthestrati¯cation option. 291 Whatifproportionalhazardsfails? ²doastrati¯ed analysis ²includeatime-varyingcovariatetoallowchanging haz- ardratiosovertime ²includeinteractions withtime Thesecondtwooptionsrelatetotime-dep endentcovariates, whichwillbecoveredinfuturelectures. Wewillfocusonthe¯rstalternativ e,andthenthesecond twooptionswillbebrie°ydescribed. 292 Strati¯edAnalyses Suppose: ²wearehappywiththeproportionalit yassumption onZ1 ²proportionalit ysimplydoesnotholdbetweenvarious levelsofasecondvariableZ2. IfZ2isdiscrete(withalevels)andthereisenoughdata,¯t thefollowingstrati¯edmodel: ¸(t;Z1;Z2)=¸Z2(t)e¯Z1 Forexample, anewtreatmen tmightleadtoa50%decrease inhazardofdeathversusthestandard treatmen t,butthe hazardforstandard treatmen tmightbedi®erentforeach hospital. Astrati¯edmodelcanbeusefulbothforprimary analysisandforcheckingthePHassumption. 293 AssessingPHAssumptionforSeveralCovariates Supposewehaveseveralcovariates(Z=Z1,Z2,...Zp),and wewanttoknowifthefollowingPHmodelholds: ¸(t;Z)=¸0(t)e¯1Z1+:::+¯pZp Tostart,we¯tamodelwhichstrati¯es byZk: ¸(t;Z)=¸0Zk(t)e¯1Z1+:::+¯k¡1Zk¡1+¯k+1Zk+1+:::+¯pZp Sincewecanestimate thesurvivalfunction foranysubgroup, wecanusethistoestimate thebaseline survivalfunction, S0Zk(t),foreachlevelofZk. Thenwecompute¡logS(t)foreachlevelofZk,controlling fortheothercovariatesinthemodel,andgraphically check whether thelogcumulativehazardsareparallelacrossstrata levels. 294 Ex:PHassumptionforgender(nursinghomedata): ²includemarried andhealthascovariatesinaCoxPH model,butstratifybygender. ²calculate thebaseline survivalfunction foreachlevelof thevariablegender(i.e.,malesandfemales) ²plotthelog-cumulativehazards formalesandfemales andevaluatewhether thelines(curves)areparallel Intheaboveexample, wemakethePHassumption formarried andhealth,butnotforgender. ThisislikegettingaKMsurvivalestimate foreachgen- derwithoutassuming PH,butismore°exiblesincewecan controlforothercovariates. Wewouldrepeatthestrati¯cation foreachvariableforwhich wewantedtocheckthePHassumption. 295 SASCodeforAssessing PHwithinStrati¯ed Model: datapop; infile'ch12.dat'; inputlosagerxgendermarried healthfail; iflos<=0thendelete; datainrisks; inputmarried health; cards; 02 ; procformat; valuesexfmt 1='Male' 0='Female'; procphregdata=pop; modellos*fail(0)=married health; stratagender; baseline covariates=inrisks out=outph loglogs=lls /nomean; procprintdata=outph; title'LogCumulative HazardEstimates byGender'; title2'Controlling forMarital andHealthStatus'; dataoutph; setoutph; iflos>0thenlog_los=log(los); labellog_los='Log(LOS)' lls='Log Cumulative Hazard'; procgplotdata=outph; plotlls*log_los=gender; formatgendersexfmt.; title1'Log-log Survival versuslog-time byGender'; run; 296 Log[-log(surviv al)]PlotsforGender ControllingforMaritalandHealthStatus GENDERFemale Male-5-4-3-2-1012 Log(LOS)5.85.96.06.16.26.36.46.56.66.76.86.97.0 297 ModelswithTime-dependentInteractions Consider aPHmodelwithtwocovariatesZ1andZ2.The standard PHmodelassumes ¸(t;Z)=¸0(t)e¯1Z1+¯2Z2 However,ifthelog-hazards arenotreallyparallelbetween thegroupsde¯nedbyZ2,thenyoucantryaddinganinter- actionwithtime: ¸(t;Z)=¸0(t)e¯1Z1+¯2Z2+¯3Z2¤t Atestofthecoe±cient¯3wouldbeatestoftheproportional hazardsassumption forZ2. If¯3ispositive,thenthehazardratiowouldbeincreasing overtime;ifnegative,thendecreasing overtime. Changes incovariatestatussometimes occurnaturally dur- ingastudy(ex.patientgetsakidneytransplan t),andare handled byintroducingtime-dependentcovariates . 298 Using StatatoAssessProportionalHazards Statahastwocommands whichcanbeusedtographically assesstheproportionalhazardsassumption, usinggraphical options(b)and(d)describedpreviously: ²stphplot: plots¡log[¡log(¡(S(t))]curvesforeach category ofanominal orordinalindependentvariable versuslog(time). Optionally ,theseestimates canbead- justedforothercovariates. ²stcoxkm: plotsKaplan-Meier observedsurvivalcurves andcompares themtotheCoxpredicted curvesforthe samevariable.(Noneedtorunstcoxpriortothiscom- mand,itwillbedoneautomatically) Foreithercommand, youmusthavestsetyourdata¯rst. Youmustspecifyby()withstcoxkm andyoumustspecify eitherby()orstrata() withstphplot . 299 AssessingPHAssumptionforaSingleCovariate byComparing¡log[¡log(S(t))]Curves .usenurshome .stsetlosfail .stphplot, by(gender) Notethatthelineswillbegoingfromtoplefttobottomright, ratherthanbottomlefttotopright,sinceweareplotting ¡log[¡log(S(t))]ratherthanlog[¡log(S(t))]. Thiswillgiveaplotsimilartothatonp.10(top). Ofcourse,you'llwanttomakeyourplotprettierbyadding titlesandlabels,asfollows: .stphplot, by(gender) xlabylabb2(log(Length ofStay)) >title(Evaluation ofPHAssumption) saving(phplot) 300 AssessingPHAssumptionforSeveralCovariates byComparing¡log[¡log(S(t))]Curves .usenurshome .stsetlosfail .genhlthsex=1 .replace hlthsex=2 ifhealth==2 &gender==1 .replace hlthsex=3 ifhealth==5 &gender==0 .replace hlthsex=4 ifhealth==5 &gender==1 .tabhlthsex .stphplot, by(hlthsex) Thiswillgiveaplotsimilartothatonp.12. 301 AssessingPHAssumptionforaSingleCovariate ControllingfortheLevelsofOtherCovariates .usenurshome .stsetlosfail .stphplot, strata(gender) adjust(married health) Thiswillproduceaplotsimilartothatonp.18. 302 AssessingPHAssumptionforaCovariate ByComparingCoxPHSurvivaltoKMSurvival Toconstruct plotsbasedonoptionI(d),usethestcoxkm command, eitherforasinglecovariateorforanewlygen- eratedcovariate(likehlthsex )whichrepresentscombined levelsofmorethanonecovariate. .usenurshome .stsetlosfail .stcoxkm, by(gender) .stcoxkm, by(hlthsex) Asusual,you'llwanttoaddtitles,labels,andsaveyour graphforlateruse. 303 Timevarying(ortime-dependent)covariates References: Allison(*) p.138-153 Hosmer&Lemesho wChapter 7,Section3 Kalb°eisc h&Prentice Section5.3 Collett Chapter 7 Kleinbaum Chapter 6 Cox&Oakes Chapter 8 Andersen &Gill Page168(Advanced!) Sofar,we'vebeenconsidering thefollowingCoxPHmodel: ¸(t;Z)=¸0(t)exp(¯Z) =¸0(t)exp(X¯jZj) where¯jistheparameter forthethej-thcovariate(Zj). Importantfeaturesofthismodel: (1)thebaselinehazarddependsont,butnotonthecovari- atesZ1;:::;Zp (2)thehazardratioexp(¯Z)dependsonthecovariatesZ1;:::;Zp, butnotontimet. Nowwewanttorelaxthesecondassumption, andallowthe hazardratiotodependontimet. 304 Exampletomotivatetime-dependentcovariates Stanford Hearttransplan texample: Variables: ²survival-timefromprogramenrollmentuntildeathorcen- soring ²dead-indicatorofdeath(1)orcensoring(0) ²transpl-whetherpatienteverhadtransplant (1ifyes,2ifno) ²surgery-previousheartsurgerypriortoprogram ²age-ageattimeofacceptanceintoprogram ²wait-timefromacceptanceintoprogramuntiltransplant surgery(=.forthosewithouttransplant) Initially,aCoxPHmodelwas¯tforpredicting survivaltime: ¸(t;Z)=¸0(t)exp(¯1¤transpl+¯2¤surgery+¯3¤age) However,thismodelcouldgivemisleading results,sincepa- tientswhodiedmorequicklyhadlesstimeavailabletoget transplan ts.Amodelwithatimedependentindicator of whether apatienthadatransplan tateachpointintime mightbemoreappropriate: ¸(t;Z)=¸0(t)exp(¯1¤trnstime+¯2¤surgery+¯3¤age) wheretrnstime=1iftranspl=1andwait>t 305 SAScodeforthesetwomodels Time-independentcovariatefortranspl : procphregdata=stanford; modelsurvival*dead(0)=transpl surgery age; run; Time-dependentcovariatefortranspl : procphregdata=stanford; modelsurvival*dead(0)=trnstime surgery age; ifwait>survival orwait=.thentrnstime=0; elsetrnstime=1; run; 306 Ifweaddtime-dep endentcovariatesorinteractions withtime totheCoxproportionalhazardsmodel,thenitisnot\pro- portionalhazards" modelanylonger. Werefertoitasan\extended Coxmodel". Comparison withasinglebinarypredictor (likehearttrans- plant): ²Astandard CoxPHmodelwouldcompare thesurvival distributions betweenthosewithoutatransplan t(ever) tothosewithatransplan t.Asubject'stransplan tstatus attheendofthestudywoulddetermine whichcategory theywereputintofortheentirestudyfollow-up. ²Anextended Coxmodelwouldcompare theriskofan eventbetweentransplan tandnon-transplan tateach eventtime,butwouldre-evaluatewhichriskgroupeach personbelongedinbasedonwhether they'dhadatrans- plantbythattime. 307 RecidivismExample: (seeAllison,p.42) Recidivismstudy: 432maleinmates werefollowedforoneyearafterrelease fromprison,toevaluateriskofre-arrest asfunction of¯nan- cialaid(fin),ageatrelease(age),race(race),full-time workexperiencepriorto¯rstarrest(wexp),maritalsta- tus(mar),parolestatus(paro=1ifreleased withparole, 0otherwise), andnumberofpriorconvictions(prio).Data werealsocollected onemploymentstatusovertimeduring theyear. Time-independentmodel: Atimeindependentmodelmightincludetheemployment statusoftheindividual atthebeginning ofthestudy(1if employed,0ifunemployed),orperhapsatanypointduring theyear. Time-dependentmodel: However,employmentstatuschangesovertime,anditmay bethemorerecentemploymentstatusthatwoulda®ectthe hazardforre-arrest. Forexample, wemightwanttode¯ne atime-dep endentcovariateforeachmonthofthestudythat indicates whether theindividual wasemployedduringthe pastmonth. 308 ExtendedCoxModel Framework: Forindividuali,supposewehavetheirfailuretime,failure indicator, andasummary oftheircovariatevaluesovertime: (Xi;±i;fZi(t);t2[0;Xi]g); fZi(t);t2[0;Xi]grepresentsthecovariatepathforthe i-thindividual whiletheyareinthestudy,andthecovariates cantakedi®erentvaluesatdi®erenttimes. Assumptions: ²conditional onanindividual's covariatehistory,thehaz- ardforfailureattimetdependsonlyonthevalueofthe covariatesatthattime: ¸(t;fZi(u);u2[0;t]g)=¸(t;Zi(t)) ²theCoxmodelforthehazardholds: ¸(t;Zi(t))=¸0(t)e¯Zi(t) Survivorfunction: S(t;Z)=expf¡Zt 0exp(¯Z(u))¸0(u)dug anddependsonthevaluesofthetimedependentvariables overtheintervalfrom0tot. ThisistheclassicformulationofthetimevaryingCoxre- gression survivalmodel. 309 Kindsoftime-varyingcovariates: ²internalcovariates: variablesthatrelatetotheindividuals, andcanonlybe measured whenanindividual isalive,e.g.whiteblood cellcount,CD4count ²externalcovariates: {variablewhichchangesinaknownway,e.g.age,dose ofdrug {variablethatexiststotallyindependentlyofallindi- viduals,e.g.airtemperature 310 ApplicationsandExamples Theextended Coxmodelisused: I.Whenimportantcovariateschangeduringastudy ²FraminghamHeartstudy 5209subjectsfollowedsince1948toexamine relation- shipbetweenriskfactorsandcardiovasculardisease. A particular example: Outcome: timetocongestiv eheartfailure Predictors: age,systolicbloodpressure, #cigarettes perday ²LiverCirrhosis (Andersen andGill,p.528) Clinicaltrialcomparing treatmen ttoplaceboforcirrho- sis.Theoutcome ofinterestistimetodeath.Patients wereseenattheclinicafter3,6and12months,then yearly. Fixedcovariates: treatmen t,gender,age(atdiagno- sis) Time-varyingcovariates: alcoholconsumption, nu- tritional status,bleeding, albumin, bilirubin, alkaline phosphatase andprothrom bin. ²RecidivismStudy(Allison, p.42) 311 II.Forcross-overstudies,toindicate changeintreatmen t ²Stanfordheartstudy(CoxandOakesp.129) Between1967and1980,249patientsenteredaprogram atStanford Universitywheretheywereregistered tore- ceiveahearttransplan t.Ofthese,184receivedtrans- plants,57diedwhilewaiting,and8droppedoutofthe program forotherreasons. Doesgettingahearttrans- plantimprovesurvival?Hereisasampleofthedata: Waiting transplant? survival post total final time transplant survival status ------------------------------------------------------------ 49 2 . . 1 5 2 . . 1 0 1 15 15 1 35 1 3 38 1 17 2 . . 1 11 1 46 57 1 etc (survivalisnotindicated aboveforthosewithouttransplants,butwasavail- ableinthedataset) Naiveapproach:Compare thetotalsurvivaloftrans- plantedandnon-transplan ted. Problem: LengthBias! 312 III.ForCompetingRisksAnalysis Forexample, incancerclinicaltrials,\tumorresponse"(or shrinking ofthetumor)isusedasanoutcome. However, clinicians wanttoknowwhether tumorresponsecorrelates withsurvival. Forthispurpose,wecan¯tanextended Coxmodelfortime todeath,withtumorresponseasatimedependentcovariate. IV.FortestingthePHassumption Forexample, wecan¯tthesetwomodels: (1)TimeindependentcovariateZ1 ¸(t;Z)=¸0(t)exp(¯1¤Z1) ThehazardratioforZ1isexp(¯1). (2)TimedependentcovariateZ1 ¸(t;Z)=¸0(t)exp(¯1¤Z1+¯2¤Z1¤t) ThehazardratioforZ1isexp(¯1+¯2t). (note:wemaywanttoreplacetby(t¡t0),sothatexp(¯1) representsHRatsomeconvenienttime,likethemediansurvival time.) Atestoftheparameter¯2isatestofthePHassumption. (howdowegetthetest?...usingtheWaldtestfromthe outputofsecondmodel,orLRtestformedbycomparing thelog-likelihoodsofthetwomodels) 313 Partiallikelihoodwithtime-varyingcovariates Startingoutjustasbefore... SupposethereareKdistinctfailure(ordeath)times,and let(¿1;::::¿K)representtheKordered, distinctdeathtimes. Fornow,assumetherearenotieddeathtimes. RiskSet:LetR(t)=fi:xi¸tgdenotethesetof individuals whoare\atrisk"forfailureattimet. Failure: Letijdenotethelabeloridentityoftheindividual whofailsattime¿j,including thevalueoftheirtime-varying covariateduringtheirtimeinthestudy fZij(t);t2[0;¿j]g History: LetHjdenotethe\history" oftheentiredata set,uptothej-thdeathorfailuretime,including thetime ofthefailure,butnottheidentityoftheonewhofails,also including thevaluesofallcovariatesforeveryoneuptoand including time¿j. PartialLikelihood:Wehaveseenpreviously thatthe partiallikelihoodcanbewrittenas L(¯)=dY j=1P(ijjHj) =dY j=1¸(¿j;Zj(¿j)) P `2R(¿j)¸(¿j;Z`(¿j)) 314 UnderthePHassumption, thisis: L(¯)=dY j=1exp(¯Zjj) P `2R(¿j)exp(¯Z`j) whereZ`jisashort-cut waytodenotethevalueoftheco- variatevectorforthe`-thpersonatthej-thdeathtime, ie: Z`j=Z`(¿j) WhatifZisnotmeasured forperson`attime¿j? ²usethemostrecentvalue(assumes stepfunction) ²interpolate ²imputebasedonsomemodel Inference (i.e.estimating theregression coe±cients,con- structing scoretests,etc.)proceedssimilarly tostandard case.Themaindi®erence isthatthevaluesofZwillchange ateachriskset. AllisonnotesthatitisveryeasytowritedownaCoxmodel withtime-dep endentcovariates,butmuchharderto¯t(com- putationally) andinterpret. 315 OldExamplerevisited: Group0:4+;7;8+;9;10+ Group1:3;5;5+;6;8+ LetZ1begroup,andaddanother¯xedcovariateZ2 IDfailcensorZ1Z2e(¯1Z1+¯2Z2) 13111e¯1+¯2 24001e¯2 35111e¯1+¯2 45010e¯1 56111e¯1+¯2 671001 78001e¯2 88010e¯1 99101e¯2 10100001 ordered Partial failure Individuals Likelihood time(¿j)atriskfailureIDcontribution 3 5 6 7 9 316 Examplecontinued: NowsupposeZ2(acompletely di®erentcovariate)isatime varyingcovariate: Z2(t) IDfailcensorZ13456789 13110 240011 3511111 4501000 56110000 671000011 7800000000 8801000011 99100001111 1010000111111 ordered Partial failure Individuals Likelihood time(¿j)atriskfailureIDcontribution 3 5 6 7 9 317 SASsolutiontopreviousexamples Title'Phregression: smallclassexample'; dataph; inputtimestatusgroupz3z4z5z6z7z8z9; cards; 3110...... 40011..... 511111.... 501000.... 6110000... 71000011.. 800000000. 801000011. 9100001111 10000111111 run; procphreg; modeltime*status(0)=group z3; run; procphreg; modeltime*status(0)=group z; z=z3; if(time>=4)thenz=z4; if(time>=5)thenz=z5; if(time>=6)thenz=z6; if(time>=7)thenz=z7; if(time>=8)thenz=z8; if(time>=9)thenz=z9; run; 318 SASoutputfrom¯ttingbothmodels Modelwithz3: Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 16.953 13.699 3.254with2DF(p=0.1965) Score . . 3.669with2DF(p=0.1597) Wald . . 2.927with2DF(p=0.2315) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio GROUP 11.610529 1.21521 1.75644 0.1851 5.005 Z3 11.360533 1.42009 0.91788 0.3380 3.898 Modelwithtime-dependentZ: Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 16.953 14.226 2.727with2DF(p=0.2558) Score . . 2.725with2DF(p=0.2560) Wald . . 2.271with2DF(p=0.3212) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio GROUP 11.826757 1.22863 2.21066 0.1371 6.214 Z 10.705963 1.20630 0.34249 0.5584 2.026 319 TheStanfordHeartTransplantdata Title'Stanford hearttransplant data:C&OTable8.1'; dataheart; infile'heart.dat'; inputwaittranspostsurvstatus; run; dataheart; setheart; iftrans=2 thensurv=wait; run; ***naiveanalysis; procphreg; modelsurv*status(2)=tstat; tstat=2-trans; ***analysis withtime-dependent covariate; procphreg; modelsurv*status(2)=tstat; tstat=0; if(trans=1 andsurv>=wait)thentstat=1; run; Thesecondmodeltookabouttwiceaslongtorunasthe ¯rstmodel,whichisusuallythecaseformodelswithtime- dependentcovariates. 320 RESULTSforStanfordHeartTransplantdata: Naivemodelwith¯xedtransplantindicator: Criterion Covariates Covariates ModelChi-Square -2LOGL 718.896 674.699 44.198with1DF(p=0.0001) Score . . 68.194with1DF(p=0.0001) Wald . . 51.720with1DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio TSTAT 1-1.999356 0.27801 51.72039 0.0001 0.135 Modelwithtime-dependenttransplantindicator: Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1330.220 1312.710 17.510with1DF(p=0.0001) Score . . 17.740with1DF(p=0.0001) Wald . . 17.151with1DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio TSTAT 1-0.965605 0.23316 17.15084 0.0001 0.381 321 RecidivismExample: Hazardforarrestwithinoneyearofreleasefromprison: Modelwithoutemploymentstatus Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1350.751 1317.496 33.266with7DF(p=0.0001) Score . . 33.529with7DF(p=0.0001) Wald . . 32.113with7DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio FIN 1-0.379422 0.1914 3.931 0.0474 0.684 AGE 1-0.057438 0.0220 6.817 0.0090 0.944 RACE 10.313900 0.3080 1.039 0.3081 1.369 WEXP 1-0.149796 0.2122 0.498 0.4803 0.861 MAR 1-0.433704 0.3819 1.290 0.2561 0.648 PARO 1-0.084871 0.1958 0.188 0.6646 0.919 PRIO 10.091497 0.0287 10.200 0.0014 1.096 Whataretheimportantpredictorsofrecidivism? 322 RecidivismExample:(cont'd) Now,weusetheindicators ofemploymentstatusforeachof the52weeksinthestudy,recorded asemp1-emp52 . Wecan¯tthemodelin2di®erentways: procphregdata=recid; modelweek*arrest(0)=fin ageracewexpmarparroprioemployed /ties=efron; arrayemp(*)emp1-emp52; doi=1to52; ifweek=ithenemployed=emp(i); end; run; ***ashortcut; procphregdata=recid; modelweek*arrest(0)=fin ageracewexpmarparroprioemployed /ties=efron; arrayemp(*)emp1-emp52; employed=emp(week); run; Thesecondwaytakes23%lesstimethanthe¯rst way,buttheresultsarethesame. 323 RecidivismExample:Output ModelWITHemploymentastime-dependentcovariate Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio FIN 1-0.356722 0.1911 3.484 0.0620 0.700 AGE 1-0.046342 0.0217 4.545 0.0330 0.955 RACE 10.338658 0.3096 1.197 0.2740 1.403 WEXP 1-0.025553 0.2114 0.015 0.9038 0.975 MAR 1-0.293747 0.3830 0.488 0.4431 0.745 PARO 1-0.064206 0.1947 0.109 0.7416 0.938 PRIO 10.085139 0.0290 8.644 0.0033 1.089 EMPLOYED 1-1.328321 0.2507 28.070 0.0001 0.265 Iscurrentemploymentimportant? Dotheothercovariateschangemuch? Canyouthinkofanyproblemwithusingcurrent employmentasapredictor? 324 Anotheroptionforassessingimpactofemploy- ment Allisonsuggests usingtheemploymentstatusofthepast weekratherthanthecurrentweek,asfollows: procphregdata=recid; whereweek>1; modelweek*arrest(0)=fin ageracewexpmarparroprioemployed /ties=efron; arrayemp(*)emp1-emp52; employed=emp(week-1); run; Thecoe±cientforemplo yedchangesfrom-1.33 to-0.79,sotheriskratioisabout0.45insteadof 0.27.Itisstillhighlysigni¯cantwithÂ2=13:1. Doesthismodelimprovethecausalinterpreta- tion? Otheroptionsfortime-dep endentcovariates: ²multiplelagsofemploymentstatus(week-1,week-2,etc.) ²cumulativeemploymentexperience(proportionofweeks worked) 325 Somecautionarynotes ²Time-varyingcovariatesmustbecarefully constructed toensureinterpretabilit y ²Thereisnopointaddingatime-varyingcovariatewhose valuechangesthesameasstudytime.....youwillget thesameanswerasusinga¯xedcovariatemeasured at studyentry.Forexample, supposewewanttostudythe e®ectofageontimetodeath. Wecould 1.useageatstartofthestudyasa¯xedcovariate 2.ageasatimevaryingcovariate However,theresultswillbethesame!Why? 326 Usingtime-varyingcovariatestoassessmodel¯t Supposewehavejust¯tthefollowingmodel: ¸(t;Z)=¸0(t)exp(¯1Z1+¯2Z2+:::¯pZp) E.g.,thenursinghomedatawithgender,maritalstatusand health. Supposewewanttotesttheproportionalit yassumption on health(Zp) Createanewvariable: Zp+1(t)=Zp¤°(t) where°(t)isaknownfunction oftime,suchas °(t)=t orlog(t) ore¡½t orIft>t¤g ThentestingH0:¯p+1=0isatestfornon-prop ortionalit y 327 Illustration:ColonCancerdata ***modelwithout time*covariate interaction; procphregdata=surv; modelsurvtime*censs(1) =trtmstagen; Modelwithouttime*stage interaction EventandCensored Values Percent Total Event Censored Censored 274 218 56 20.44 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1959.927 1939.654 20.273with2DF(p=0.0001) Score . . 18.762with2DF(p=0.0001) Wald . . 18.017with2DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio TRTM 10.016675 0.13650 0.01492 0.9028 1.017 STAGEN 1-0.701408 0.16539 17.98448 0.0001 0.496 328 ***modelWITHtime*covariate interaction; procphregdata=surv ; modelsurvtime*censs(1) =trtmstagentstage; tstage=stagen*exp(-survtime/1000); ModelWITHtime*stage interaction Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1959.927 1902.374 57.553with3DF(p=0.0001) Score . . 35.960with3DF(p=0.0001) Wald . . 19.319with3DF(p=0.0002) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr>Risk Variable DFEstimate Error Chi-Square Chi-Square Ratio TRTM 10.008309 0.13654 0.00370 0.9515 1.008 STAGEN 11.402244 0.45524 9.48774 0.0021 4.064 TSTAGE 1-8.322371 2.04554 16.55310 0.0001 0.000 LikeCoxandOakes,wecanrunafewdi®erentmodels 329 Time-varyingcovariatesinStata CreateadatasetwithanIDcolumn,andonelineperperson foreachdi®erentvalueofthetimevaryingcovariate. .infileidtimestatusgroupzusingcox4_stata.dat or .inputidtimestatusgroup z 13 110 25 010 35 111 46 110 56 010 58 011 64 001 75 000 77 101 88 000 95 000 99 101 103 000 1010 001 .end .stsettimestatus .coxtimegroupz,dead(status) tvid(id) ------------------------------------------------------------------- -------- --- time| status|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- -------- --- group|1.826757 1.228625 1.487 0.137 -.5813045 4.234819 z|.7059632 1.206304 0.585 0.558 -1.65835 3.070276 ------------------------------------------------------------------- -------- --- 330 Time-varyingcovariatesinSplus Createadatasetwithstartandstopvaluesoftime: idstartstopstatusgroup z 103110 205010 305111 406110 506010 568011 604001 705000 757101 808000 905000 959101 1003000 10310 001 331 ThentheSpluscommands andresultsare: Commands: y_read.table("cox4_splus.dat",header=T) agreg(y$start,y$stop,y$status,cbind(y$group,y$z)) Results: AliveDeadDeleted 95 0 coefexp(coef) se(coef) zp [1,]1.827 6.21 1.231.4870.137 [2,]0.706 2.03 1.210.5850.558 exp(coef) exp(-coef) lower.95upper.95 [1,] 6.21 0.161 0.559 69.0 [2,] 2.03 0.494 0.190 21.5 Likelihood ratiotest=2.73on2df,p=0.256 Efficient scoretest=2.73on2df,p=0.256 332 PiecewiseCoxModel:(Collett,Chapter10) Atimedependentcovariatecanbeusedtocreateapiecewise PHcoxmodel.Supposeweareinterested incomparing two treatmen ts,and: ²HR=µ1duringtheinterval(0;t1) ²HR=µ2duringtheinterval(t1;t2) ²HR=µ3duringtheinterval(t2;1) De¯nethefollowingcovariates: ²X-treatmen tindicator (X=0!standard,X=1!newtreatmen t) ²Z2-indicator ofchangeinHRduring2ndinterval Z2(t)=8 >< >:1ift2(t1;t2)andX=1 0otherwise ²Z3-indicator ofchangeinHRduring3rdinterval Z3(t)=8 >< >:1ift2(t2;1)andX=1 0otherwise Themodelforthehazardforindividualiis: ¸i(t)=¸0(t)expf¯1xi+¯2z2i(t)+¯3z3i(t)g Whataretheloghazardratiosforanindividual onthenew treatmen trelativetooneonthestandard treatmen t? 333 Timevarying(ortime-dependent)covariates CaseStudyofMACDiseaseTrial ACTG196wasarandomizedclinicaltrialtostudythee®ects ofcombinationregimensonpreventionofMAC(mycobacterium aviumcomplex)disease,whichisoneofthemostcommonoppor- tunisticinfectionsinAIDSpatientsandisassociatedwithhighmor- talityandmorbidity. Thetreatmentregimenswere: ²clarithromycin(new) ²rifabutin(standard) ²clarithromycinplusrifabutin ThistrialenrolledpatientsbetweenApril1993andFebruary1994, andfollowedpatientsthroughAugust1995.InFebruaryof1994,the dosageofrifabutinwasreducedfrom3capsulesperday(450mg) to2capsulesperday(300mg)duetoconcernoveruveitis,an adverseexperienceresultinginin°ammationoftheuvealtractin theeyes(about3-4%ofpatientsreporteduveitis).Allpatientswere toreducetheirdosagebyMarch8,1994.However,somepatients hadalreadydiscontinuedthetreatment,died,ordiscontinuedthe study. Themainintent-to-treatanalysiscomparedthe3treatmentarms withoutadjustingforthischangeindosage. Othersupportinganalysesattemptedtountanglethee®ectofthis \studywidedosereduction" (SWDR). 334 ProportiononeachtreatmentarmwithSWDR Treatment bystudywidedosereduction TABLEOFTRTMTBYSWDRSTAT TRTMT SWDRSTAT(Study WideDoseReduction Status) Frequency| RowPct|No |Yes |Total ---------+--------+--------+ R |125|266|391 |31.97|68.03| ---------+--------+--------+ C+R |170|219|389 |43.70|56.30| ---------+--------+--------+ C |124|274|398 |31.16|68.84| ---------+--------+--------+ Total 419 759 1178 STATISTICS FORTABLEOFTRTMTBYSWDRSTAT Statistic DFValue Prob ------------------------------------------------------ Chi-Square 216.820 0.001 Likelihood RatioChi-Square 216.610 0.001 Mantel-Haenszel Chi-Square 10.067 0.795 PhiCoefficient 0.119 Contingency Coefficient 0.119 Cramer's V 0.119 SampleSize=1178 335 OriginalLogranktestComparing 3TreatmentArms (Howwouldyougetpairwisetests?) Dependent Variable: MACTIME TimetoMACdisease (days) Censoring Variable: MACSTAT MACstatus(1=yes,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 1178 121 1057 89.73 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1525.932 15.133with2DF(p=0.0005) Score . . 15.890with2DF(p=0.0004) Wald . . 15.209with2DF(p=0.0005) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.231842 0.25748 0.81074 0.3679 1.261 RIF 10.826883 0.23601 12.27480 0.0005 2.286 Variable Label CLARI 1=Clarithromycin arm,0otherwise RIF 1=Rifabutin arm,0otherwise LinearHypotheses Testing Wald Pr> Label Chi-Square DFChi-Square TEST_TRT 15.2094 2 0.0005 336 Kaplan-Meier SurvivalPlot EstimatedProbabilitiesofRemainingMAC-freeSurvival Distribution Function 0.00.10.20.30.40.50.60.70.80.91.0 Time to MAC disease (days)0100200300400500600700800900 STRATA:TRTMT=Clar + Rif TRTMT=Clarithro TRTMT=Rifabutin %ps(mactrt.ps,mode=replace); proclifetest data=weighted noprint outsurv=survres graphics nocensplots=(s); timemactime*macstat(0); stratatrtmt; title'TimetoMACbyTreatment Regimen'; formattrtmttrtfmt.; run; 337 Howwelldoesthismodel¯t? Let'stakealookattheresidualplots... First,thedeviance residuals: D e v i a n c e R e s i d u a l-101234 1=Rifabutin arm, 0 otherwise-1 0 1D e v i a n c e R e s i d u a l-101234 1=Clarithromycin arm, 0 otherwise-1 0 1 D e v i a n c e R e s i d u a l-101234 Linear Predictor0.00.10.20.30.40.50.60.70.80.9 Plottingdevianceresidualsvsbinarycovariates isnotveryuseful. 338 Howaboutthegeneralizedresiduals? (Aretheylikeasamplefromacensored unitexponential?) (i.e., is slope=1, intercept=0) LLS -8-7-6-5-4-3-2-1 Log(generalized residual)-8 -7 -6 -5 -4 -3 -2 -1 intercept=0.056 slope=1.028 (basedon¯ttingaregression linetoresiduals) 339 Wecanalsolookatthelogcumulativehazard plots(i.e.,log[¡log(^S)])versuslogtimetoseewhether thelinesareparallelforthethreetreatmentgroups. Plot of log-log KM versus log-time MAC Prophylaxis Therapy Rifabutin ClarithroClar + RifL n [ - l n ( S ) ] -6-5-4-3-2-1 Log(Time to MAC)3 4 5 6 (Ihavejoinedtheindividual pointsusingi=joininthesym- bolstatemen t,tomakethemeasiertosee.) 340 Shouldn't weadjustforBaselineCD4count? Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1488.737 52.328with3DF(p=0.0001) Score . . 43.477with3DF(p=0.0001) Wald . . 43.680with3DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.198798 0.25747 0.59619 0.4400 1.220 RIF 10.837240 0.23598 12.58738 0.0004 2.310 CD4 1-0.019641 0.00367 28.59491 0.0001 0.981 Analysis ofMaximum Likelihood Estimates Variable Label CLARI 1=Clarithromycin arm,0otherwise RIF 1=Rifabutin arm,0otherwise CD4 CD4CellCount IsCD4countaconfounder? (Ananalysisstrati¯edbyCD4categorygavealmostidenticalre- sults.OtherimportantcovariatesincludedCTG(clinicaltrials group)andKarnofskystatus). 341 Whatdothedevianceresidualslooklikeversus acontinuouscovariate,likeCD4? D e v i a n c e R e s i d u a l-2-101234 CD4 Cell Count0 100 200 300 Wemightwanttoconsidersomekindoftransformation ofCD4 count(likelogorsquareroot).Ifwedon'tfeelcomfortablewith thelinearityofCD4count,wecanalsodichotomizeit(CD4CAT). 342 Anotherwayofcheckingtheproportionalityas- sumptionisbyusingtheWeightedSchoenfeld residualplotsforeachcovariate RawCD4count logCD4count -0.10-0.050.000.050.100.150.200.25 Time to MAC disease (days)0100200300400500600700800900-2-1012 Time to MAC disease (days)0 100200300400500600700800900 SquarerootCD4count -0.8-0.6-0.4-0.20.00.20.40.60.81.01.2 Time to MAC disease (days)0100200300400500600700800900 343 Sofar,thegraphical techniqueshavenotindicated anyma- jordeparture fromproportional hazards. However,wecan testthisformally bycreating atimedependentcovariatefor rifabutin andclarithrom ycin: riftd=rif*((mactime-365)/30); claritd=clari*((mactime-365)/30); Eventhoughthedosereduction wasonlyforrifabutin, pa- tientsonall3armshadtohavethedosereduction ...they justtook2capsules oftheirplacebo,anddidn'tknowwhether itwasplacebooractivedrug. Ihavecenteredthetime-dep endentcovariatesat365days (oneyear),sothattheHRforrifaloneandclarialonewill applyatoneyear.ThenIhavedividedby30,sothatthe resulting HRcanbeinterpreted asthechangeforeachmonth awayfrom365days. Question:Canwedothiswithinadatastepus- ingtheabovestatements,ordothesestatements needtobegiveninthePROCPHREGproce- dure? 344 Time-dependentcovariatesforclariandrif Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1525.837 15.227with4DF(p=0.0043) Score . . 16.033with4DF(p=0.0030) Wald . . 15.327with4DF(p=0.0041) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.229811 0.25809 0.79287 0.3732 1.258 RIF 10.823227 0.23624 12.14274 0.0005 2.278 CLARITD 10.003065 0.04073 0.00566 0.9400 1.003 RIFTD 10.010627 0.03765 0.07965 0.7778 1.011 Analysis ofMaximum Likelihood Estimates Variable Label CLARI 1=Clarithromycin arm,0otherwise RIF 1=Rifabutin arm,0otherwise Neithertime-dependentcovariatewassigni¯cant. 345 Thisanalysis alsoindicated thattherearenomajorde- partures fromproportionalhazardsforthethreetreatmen t arms. However,itmaystillbethecasethathavingthestudy-wide dosereduction hadsomerelationship withMACdisease. Wecanassessthisbycreating atimedependentvariablefor theSWDR. We'lllookatthefollowingmodels: (1)SWDRST ATasasimpleindicator (2)SWDRST ATandSWDRTD,with swdrtd=swdrstat*((mactime-365)/30) (3)SWDRastimedependentcovariate 346 Naivemodelwith¯xedSWDRindicator (SWDRST AT): Dependent Variable: MACTIME TimetoMACdisease (days) Censoring Variable: MACSTAT MACstatus(1=yes,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1495.857 45.208with3DF(p=0.0001) Score . . 51.497with3DF(p=0.0001) Wald . . 48.749with3DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.449936 0.26142 2.96236 0.0852 1.568 RIF 11.006639 0.23852 17.81114 0.0001 2.736 SWDRSTAT 1-1.125032 0.19283 34.04055 0.0001 0.325 Analysis ofMaximum Likelihood Estimates Variable Label CLARI 1=Clarithromycin arm,0otherwise RIF 1=Rifabutin arm,0otherwise SWDRSTAT StudyWideDoseReduction Status Reduction ofdosagefrom450mgto300mgappearstobe protective,whichseemscounter-intuitive 347 PredictedBaselineSurvivalCurves: Another waytoseethisisthrough thepredicted baseline survivalcurves.Thetwolinesareforthosenotonrifabutin, whilethex'sand+'sareforthoseonrifabutin. Ineachcase, thehigherline(betterprognosis) ofthepairisforthosewho didhavetheSWDR. SWDR/Rifabutin Status No SWDR/no RIF SWDR/no RIF No SWDR/RIF SWDR/RIFP r ( s u r v i v a l ) 0.00.20.40.60.81.0 Time to MAC (days)0 100200300400500600700800 348 Testforproportionality: procphregdata=weighted; modelmactime*macstat(0) =claririfswdrstat swdrtd; ***createtimebycovariate interaction forswdrstatus; swdrtd=swdrstat*((mactime-365)/30); test_trt: testclari,rif; title'Testoftreatment Differences'; title2'andtestofproportionality att=365days'; Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1492.692 48.372with4DF(p=0.0001) Score . . 55.174with4DF(p=0.0001) Wald . . 50.719with4DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.430051 0.26126 2.70947 0.0998 1.537 RIF 11.005416 0.23845 17.77884 0.0001 2.733 SWDRSTAT 1-1.126498 0.19752 32.52551 0.0001 0.324 SWDRTD 10.055550 0.03201 3.01112 0.0827 1.057 Variable Label CLARI 1=Clarithromycin arm,0otherwise RIF 1=Rifabutin arm,0otherwise SWDRSTAT StudyWideDoseReduction Status SWDRTD swdrstat*((mactime-365)/30) 349 InterpretationofHazardRatios ¯swdrstat=¡1:1265 ¯swdrtd=0:0556 TimeTime Hazard (months)(days)calculation Ratio 6182.5exp[¡1:1265+(¡6:08)(0:0556)]0.231 12365exp[¡1:1265+(0)(0:0556)]0.324 18547.5exp[¡1:1265+(6:08)(0:0556)]0.454 24730exp[¡1:1265+(12:17)(0:0556)]0.637 30912.5exp[¡1:1265+(18:25)(0:0556)]0.893 361095exp[¡1:1265+(24:33)(0:0556)]1.253 HR=exp[¯swdrstat+¯swdrtdÃmactime¡365) 30! ] Intheearlyperiodafterrandomization totreatmen t,reduc- tionofrandomized dosagefrom450mgto300mgisassoci- atedwithadecreased riskofMACdisease.Aftertakingthe higherdosageforabout32months,dropping tothelower dosagehasnoimpact,andasthetreatmen ttimeincreases beyond32months,alowerdosagetendstobeassociated withincreased riskofMAC. 350 3di®erentwaystocodeSWDRastime-dependent covariate procphregdata=weighted; modelmactime*macstat(0) =claririfswdr; if(swdrtime>=mactime) thenswdr=0; elsedo; ifswdrstat=1 thenswdr=1; elseswdr=0; end; test_trt: testclari,rif; title2'I.Time-dependent indicator ofdosereduction'; procphregdata=weighted; modelmactime*macstat(0) =claririfswdr; ifswdrstat=0 or(swdrtime>=mactime) thenswdr=0; elseswdr=1; test_trt: testclari,rif; title2'II.Time-dependent indicator ofdosereduction'; procphregdata=weighted; modelmactime*macstat(0) =claririfswdr; ifswdrstat=1 and(swdrtime<mactime) thenswdr=1; elseswdr=0; test_trt: testclari,rif; title2'III.Time-dependent indicator ofdosereduction'; 351 Outputisthesameforall3cases: Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 1178 121 1057 89.73 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1517.426 23.639with3DF(p=0.0001) Score . . 24.844with3DF(p=0.0001) Wald . . 24.142with3DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.328849 0.26017 1.59762 0.2062 1.389 RIF 10.905299 0.23775 14.49956 0.0001 2.473 SWDR 1-0.648887 0.21518 9.09389 0.0026 0.523 SWDRisstillprotective?Doesthismakesenseintuitively? Whatothermethodscanweusetoaccountforchangein dosage? 352 Weightedadjusteddose(WAD)analyses Totrytogetabetterideaofthee®ectofchanging doses ofrifabutin onthehazardforMACdisease,Icreatedthe followingweighteddoseofrandomized rifabutin: ²Betweenrandomization dateandSWDRdate =)#Daysat450mg ²BetweenSWDRdateando®-study date =)#Daysat300mg ²Betweenrandomization dateandO®-study date =)#TotalDays ²Weightedrandomized dose rifwadr =(days450 +days300)/totdays ²Transformed tonumberofcapsules perday; rifwadr=rifwadr/150; ²Alsocalculated weighteddosewhileontreatmentby startingwithontreatmen tdate,stopping witho®-treatmen t date,anddividing bythetotaldaysonstudy. 353 Weightedadjusteddose(WAD)analyses Randomizedassignmenttorifabutin Dependent Variable: MACTIME TimetoMACdisease (days) Censoring Variable: MACSTAT MACstatus(1=yes,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 1178 121 1057 89.73 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1493.476 47.588with3DF(p=0.0001) Score . . 52.770with3DF(p=0.0001) Wald . . 50.295with3DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.453283 0.26119 3.01179 0.0827 1.573 RIF 11.004846 0.23826 17.78681 0.0001 2.731 RIFWADR 11.530462 0.25681 35.51502 0.0001 4.620 Foreachadditional capsuleofrifabutin speci¯edasran- domizedtreatment,theHRforMACincreased by4.6 times 354 Weightedadjusteddose(WAD)analyses Actualdosageofrifabutinduringthestudy Dependent Variable: MACTIME TimetoMACdisease (days) Censoring Variable: MACSTAT MACstatus(1=yes,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1489.993 51.071with3DF(p=0.0001) Score . . 55.942with3DF(p=0.0001) Wald . . 53.477with3DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.489583 0.26256 3.47693 0.0622 1.632 RIF 11.019675 0.23873 18.24291 0.0001 2.772 RIFWAD 1-0.664689 0.10686 38.69332 0.0001 0.514 Here,highervaluesofRIFWADprobably re°ectthatthe patientwasabletostayontreatmentlonger,whichwas protective.TheSWDRvariableisalsocapturing whethera patienthadbeenabletotoleratethetreatmentlongenough tohavethechancetohavetheprotocol-mandated dose reduction. 355 Whathappensifweaddtreatmentdiscontinua- tionasatimedependentcovariate? Dependent Variable: MACTIME TimetoMACdisease (days) Censoring Variable: MACSTAT MACstatus(1=yes,0=censored) Censoring Value(s): 0 TiesHandling: BRESLOW Summary oftheNumberof EventandCensored Values Percent Total Event Censored Censored 1178 121 1057 89.73 Testing GlobalNullHypothesis: BETA=0 Without With Criterion Covariates Covariates ModelChi-Square -2LOGL 1541.064 1501.595 39.469with4DF(p=0.0001) Score . . 42.817with4DF(p=0.0001) Wald . . 41.027with4DF(p=0.0001) Analysis ofMaximum Likelihood Estimates Parameter Standard Wald Pr> Risk Variable DF Estimate Error Chi-Square Chi-Square Ratio CLARI 10.420447 0.26111 2.59284 0.1073 1.523 RIF 10.984114 0.23847 17.02975 0.0001 2.675 SWDR 1-0.139245 0.23909 0.33919 0.5603 0.870 RXSTOP 10.902592 0.21792 17.15473 0.0001 2.466 SWDRisnolongersigni¯cant! 356 Lastofall,acomparisonofsomeofthesemodels: AIC Modelterms q¡2logLCriterion Clari, Rif 21525.931531.93 Clari, Rif,Cd4ca t 31497.571506.57 Clari, Rif,Cd4 31488.741497.74 Clari, Rif,Cd4ca t,Ctg,Karnof 51482.671497.67 Clari, Rif,Swdrst at 31495.861504.86 Clari, Rif,Rifwadr 31493.481502.48 Clari, Rif,Swdrst at,Rifwadr41493.441505.44 Clari, Rif,Rifwad 31489.991498.99 Modelswithtime-dependentcovariates Clari, Rif,Claritd, Riftd 41525.841537.84 Clari, Rif,Swdrst at,Swdrtd41492.691504.69 Clari, Rif,Swdr 31517.431526.43 Clari, Rif,Swdr, Rxstop 41501.601513.60 Clari, Rif,Cd4ca t,Karnof, Rxstop51461.901476.90 Clari, Rif,Cd4ca t,Karnof, Rifwad51448.141463.14 357 Parametric SurvivalAnalysis Sofar,wehavefocusedprimarily onnonparametric and semi-parametric approachestosurvivalanalysis, withheavy emphasis ontheCoxproportionalhazardsmodel: ¸(t;Z)=¸0(t)exp(¯Z) Weusedthefollowingestimating approach: ²Weestimated¸0(t)nonparametrically ,usingtheKaplan- Meierestimator, orusingtheKalb°eisc h/Prenticeesti- matorunderthePHassumption ²Weestimated¯byassuming alinearmodelbetweenthe logHRandcovariates,underthePHmodel Bothestimates werebasedonmaximumlikelihoodtheory. 358 Thereareseveralreasonswhyweshouldconsider someal- ternativeapproachesbasedonparametric models: ²Theassumption ofproportionalhazardsmightnotbe appropriate (basedonmajordepartures) ²Ifaparametric modelactually holds,thenwewould probably gaine±ciency ²Wemaywanttohandlenon-standard situations like {intervalcensoring {incorporatingpopulation mortality ²Wemaywanttomakesomeconnections withotherfa- miliarapproaches(e.g.useofthePoissonlikelihood) ²Wemaywanttoobtainsomeestimates foruseindesign- ingafuturesurvivalstudy. 359 Asimplestart:ExponentialRegression ²Observeddata: (Xi;±i;Zi)forindividuali, Zi=(Zi1;Zi2;:::;Zip)representsasetofpcovariates. ²Rightcensoring: AssumethatXi=min(Ti;Ui) ²Survivaldistribution: AssumeTifollowsanexpo- nentialdistribution withaparameter¸thatdependson Zi,say¸i=ª(Zi).Thenwecanwrite: Ti»exponential (ª(Zi)) First,let'sreviewsomefactsabouttheexponentialdistribu- tion(fromour¯rstsurvivallecture): f(t)=¸e¡¸tfort¸0 S(t)=P(T¸t)=Z1 tf(u)du=e¡¸t F(t)=P(T<t)=1¡e¡¸t ¸(t)=f(t) S(t)=¸constanthazard! ¤(t)=Zt 0¸(u)du=Zt 0¸du=¸t 360 Now,wesaythat¸isaconstantovertimet,butwewant toletitdependonthecovariatevalues,sowearesetting ¸i=ª(Zi) Thehazardratewouldtherefore bethesameforanytwo individuals withthesamecovariatevalues. Although therearemanypossiblechoicesforª,onesimple andnaturalchoiceis: ª(Zi)=exp[¯0+Zi1¯1+Zi2¯2+:::+Zip¯p] WHY? ²ensuresapositivehazard ²foranindividual withZ=0,thehazardise¯0. Themodeliscalledexponentialregression becauseof thenaturalgeneralization fromregularlinearregression 361 Exponentialregressionforthe2-samplecase: ²AssumewehaveonlyasinglecovariateZ=Z, i.e.,p=1. HazardRate: ª(Zi)=exp(¯0+Zi¯1) ²De¯ne:Zi=0ifindividualiisingroup0 Zi=1ifindividualiisingroup1 ²Whatisthehazardforgroup0? ²Whatisthehazardforgroup1? ²Whatisthehazardratioofgroup1togroup 0? ²Whatistheinterpretationof¯1? 362 LikelihoodforExponentialModel Undertheassumption ofrightcensored data,eachperson hasoneoftwopossiblecontributions tothelikelihood: (a)theyhaveaneventatXi(±i=1))contribution is Li=S(Xi) |{z}¢¸(Xi) |{z}=e¡¸Xi¸ survivetoXifailatXi (b)theyarecensored atXi(±i=0))contribution is Li=S(Xi) |{z}=e¡¸Xi survivetoXi Thelikelihoodistheproductoveralloftheindividuals: L=Y iLi =Y iµ ¸e¡¸Xi¶±i |{z}µ e¡¸Xi¶(1¡±i) |{z} eventscensorings =Y i¸±iµ e¡¸Xi¶ 363 MaximumLikelihoodforExponential Howdoweusethelikelihood? ²¯rsttakethelog ²thentakethepartialderivativewithrespectto¯ ²thensettozeroandsolveforc¯ ²thisgivesusthemaximumlikelihoodestimators Thelog-likelihoodis: logL=log2 4Y i¸±iµ e¡¸Xi¶3 5 =X i[±ilog(¸)¡¸Xi] =X i[±ilog(¸)]¡X i¸Xi Forthecaseofexponentialregression, wenowsubstitute the hazard¸=ª(Zi)intheabovelog-likelihood: logL=X i[±ilog(ª(Zi))]¡X iª(Zi)Xi(1) 364 GeneralFormofLog-likelihood forRightCensored Data Ingeneral,wheneverwehaverightcensored data,thelikeli- hoodandcorrespondingloglikelihoodwillhavethefollowing forms: L=Y i[¸i(Xi)]±iSi(Xi) logL=X i[±ilog(¸i(Xi))]¡X i¤i(Xi) where ²¸i(Xi)isthehazardfortheindividualiwhofailsatXi ²¤i(Xi)isthecumulativehazardforanindividual attheir failureorcensoring time Forexample, seethederivationofthelikelihoodforaCox modelonp.11-13ofLecture4notes.Westartedwiththe likelihoodabove,thensubstituted thespeci¯cformsfor¸(Xi) underthePHassumption. 365 Consider ourmodelforthehazardrate: ¸=ª(Zi)=exp[¯0+Zi1¯1+Zi2¯2+:::+Zip¯p] Wecanwritethisusingvectornotation, asfollows: LetZi=(1;Zi1;:::Zip)T and¯=(¯0;¯1;:::¯p) (Since¯0istheintercept(i.e.,theloghazardrateforthe baseline group),weputa\1"asthe¯rstterminthevector Zi.) Then,wecanwritethehazardas: ª(Zi)=exp[¯Zi] Nowwecansubstitute ª(Zi)=exp[¯Zi]inthelog-likelihood shownin(1): logL=nX i=1±i(¯Zi)¡nX i=1Xiexp(¯Zi) 366 ScoreEquations Takingthederivativewithrespectto¯0,thescoreequation is: @logL @¯0=nX i=1[±i¡Xiexp(¯Zi)] For¯k,k=1;:::p,theequations are: @logL @¯k=nX i=1[±iZik¡XiZikexp(¯Zi)] =nX i=1Zik[±i¡Xiexp(¯Zi)] To¯ndtheMLE's,wesettheaboveequations to0and solve(simultaneously). Theequations aboveimplythat theMLE'sareobtained bysettingtheweightednumberof failures(P iZik±i)equaltotheweightedcumulativehazard (P iZik¤(Xi)). 367 To¯ndthevarianceoftheMLE's,weneedtotakethesecond derivatives: ¡@2logL @¯k@¯j=nX i=1ZikZijXiexp(¯Zi) Somealgebra(seeCoxandOakessection6.2)revealsthat Var(c¯)=I(¯)¡1=· Z(I¡¦)ZT¸¡1 where ²Z=(Z1;:::;Zn)isa(p+1)£nmatrix (pcovariatesplusthe\1"fortheintercept¯0) ²¦=diag(¼1;:::;¼n)(thismeansthat¦isadiagonal matrix,withtheterms¼1;:::;¼nonthediagonal) ²¼iistheprobabilit ythatthei-thpersoniscensored, so (1¡¼i)istheprobabilit ythattheyfailed. ²Note:Theinformation I(¯)(inverseofthevariance) isproportionaltothenumberoffailures,notthesample size.Thiswillbeimportantwhenwetalkaboutstudy design. 368 TheSingleSampleProblem(Zi=1foreveryone): First,whatistheMLEof¯0? Weset@logL @¯0=Pn i=1[±i¡Xiexp(¯0Zi)]equalto0andsolve: )nX i=1±i=nX i=1[Xiexp(¯0)] d=exp(¯0)nX i=1Xi exp(d¯0)=d Pni=1Xi ^¸=d t wheredisthetotalnumberofdeaths(orevents),andt= PXiisthetotalperson-time contributed byallindividuals. Ifd=tistheMLEfor¸,whatdoesthisimply abouttheMLEof¯0? 369 Usingtheprevious formulaVar(^¯)=· Z(I¡¦)ZT¸¡1, whatisthevarianceofc¯0?: Withsomematrixalgebra, youcanshowthatitis: Var(c¯0)=1 Pni=1(1¡¼i)=1 d Whatabout^¸=e^¯0? Bythedeltamethod, Var(^¸)=^¸2Var(c¯0) =? 370 TheTwo-Sample Problem: ZiSubjectsEventsFollow-up Group0:Zi=0n0d0t0=Pn0i=1Xi Group1:Zi=1n1d1t1=Pn1i=1Xi Thelog-likelihood: logL=nX i=1±i(¯0+¯1Zi)¡nX i=1Xiexp(¯0+¯1Zi) so@logL @¯0=nX i=1[±i¡Xiexp(¯0+¯1Zi)] =(d0+d1)¡(t0e¯0+t1e¯0+¯1) @logL @¯1=nX i=1Zi[±i¡Xiexp(¯0+¯1Zi)] =d1¡t1e¯0+¯1 Thisimplies: ^¸1=e^¯0+^¯1=? ^¸0=e^¯0=? ^¯0=? ^¯1=? 371 ImportantResult: Themaximumlikelihoodestimates (MLE's)ofthehazardratesunder theexponentialmodelarethenum- berofeventsdividedbytheperson- yearsoffollow-up! (thisresultwillbereliedonheavilywhenwedis- cussstudydesign) 372 ExponentialRegression: MeansandMedians MeanSurvivalTime Fortheexponentialdistribution, E(T)=1=¸. ²ControlGroup: T0=1=^¸0=1=exp(^¯0) ²TreatmentGroup: T1=1=^¸1=1=exp(^¯0+^¯1) MedianSurvivalTime ThisisthevalueMatwhichS(t)=e¡¸t=0:5,soM= median=¡log(0:5) ¸ ²ControlGroup: ^M0=¡log(0:5) ^¸0=¡log(0:5) exp(^¯0) ²TreatmentGroup: ^M1=¡log(0:5) ^¸1=¡log(0:5) exp(^¯0+^¯1) 373 ExponentialRegression: VarianceEstimatesandTestStatistics Wecanalsocalculate thevariances oftheMLE'sassimple functions ofthenumberoffailures: var(^¯0)=1 d0 var(^¯1)=1 d0+1 d1 Soourteststatistics areformedas: FortestingHo:¯0=0: Â2 w=µ^¯0¶2 var(^¯0) =[log(d0=t0)]2 1=d0 FortestingHo:¯1=0: Â2 w=µ^¯1¶2 var(^¯1) =· log(d1=t1 d0=t0)¸2 1 d0+1 d1 Howwouldweformcon¯dence intervalsforthehazard ratio? 374 TheLikelihoodRatioTestStatistic: (Analternativ etotheWaldtest) Alikelihoodratiotestisbasedon2timesthelogoftheratio ofthelikelihoodsunderthenullandalternativ e.Wereject H0if2log(LR)>Â2 1;0:05,where LR=L(H1) L(H0)=L(b¸0;b¸1) L(b¸) Forasampleofnindependentexponentialrandomvariables withparameter¸,theLikelihoodis: L=nY i=1[¸±iexp(¡¸xi)] =¸dexp(¡¸Xxi) =¸dexp(¡¸n¹x) wheredisthenumberofdeathsorfailures. Thelog-likelihoodis `=dlog(¸)¡¸n¹x andtheMLEis b¸=d=(n¹x) 375 2-SampleCase:LRtestcalculations Data: Group0:d0failuresamongthen0females meanfailuretimeis¹x0=(Pn0iXi)=n0 Group1:d1failuresamongthen1males meanfailuretimeis¹x1=(Pn1iXi)=n1 Underthealternativehypothesis: L=¸d11exp(¡¸1n1¹x1)£¸d00exp(¡¸0n0¹x0) log(L)=d1log(¸1)¡¸1n1¹x1+d0log(¸0)¡¸0n0¹x0 TheMLE'sare: b¸1=d1=(n1¹x1)formales b¸0=d0=(n0¹x0)forfemales Underthenullhypothesis: L=¸d1+d0exp[¡¸(n1¹x1+n0¹x0)] log(L)=(d1+d0)log(¸)¡¸[n1¹x1+n0¹x0] ThecorrespondingMLEis b¸=(d1+d0)=[n1¹x1+n0¹x0] 376 Alikelihoodratiotestcanbeconstructed bytakingtwicethe di®erence ofthelog-likelihoodsunderthealternativ eandthe nullhypotheses: ¡22 4(d0+d1)log0 @d0+d1 t0+t11 A¡d1log[d1=t1]¡d0log[d0=t0]3 5 Nursinghomeexample: Forthefemales: ²n0=1173 ²d0=902 ²t0=310754 ²¹x0=265 Forthemales: ²n1=418 ²d1=367 ²t1=75457 ²¹x1=181 Plugging thesevaluesin,wegetaLRteststatisticof64.20. 377 HandCalculationsusingeventsandfollow-up: Byaddingup\los"formalestogett1andforfemalesto gett0,Iobtained: ²d0=902(females) d1=367(males) ²t0=310754(femalefollow-up) t1=75457(malefollow-up) ²ThisyieldsanestimatedlogHR: ^¯1=log2 4d1=t1 d0=t03 5=log2 4367=75457 902=3107543 5=log(1:6756)=0:5162 ²Theestimatedstandarderroris: r var(^¯1)=vuut1 d1+1 d0=vuut1 902+1 367=0:06192 ²SotheWaldtestbecomes: Â2 W=^¯2 1 var(^¯1)=(0:51619)2 0:061915=69:51 ²Wecanalsocalculate^¯0=log(d0=t0)=¡5:842, alongwithitsstandarderrorse(^¯0)=q (1=d0)=0:0333 378 ExponentialRegressioninSTATA .usenurshome .stsetlosfail .streggender, dist(exp) nohr failure _d:fail analysis time_t:los Iteration 0:loglikelihood =-3352.5765 Iteration 1:loglikelihood =-3321.966 Iteration 2:loglikelihood =-3320.4792 Iteration 3:loglikelihood =-3320.4766 Iteration 4:loglikelihood =-3320.4766 Exponential regression --logrelative-hazard form No.ofsubjects = 1591 Numberofobs=1591 No.offailures = 1269 Timeatrisk = 386211 LRchi2(1) =64.20 Loglikelihood =-3320.4766 Prob>chi2 =0.0000 ------------------------------------------------------------------- ------ _t|Coef.Std.Err. zP>|z| [95%Conf.Interval] ---------|--------------------------------------------------------- ----- gender|.516186 .0619148 8.337 0.000 .3948352 .6375368 _cons|-5.842142 .0332964 -175.459 0.000 -5.907402 -5.776883 ------------------------------------------------------------------- ------ SinceZ=8:337,thechi-squaretestisZ2=69:51. 379 ExponentialRegressioninSAS-proclifereg procformat; valuecensfmt 1='Censored' 0='Dead'; valuegrpfmt 0='Group 0(F)' 1='Group 1(M)'; Title'Exponential HazardModelforNursing HomePatients'; datamorris; infile'ch12.dat'; inputlosagetrtgendermarstat hltstat cens; datamorris2; setmorris; iflos=0thendelete; procfreqdata=morris2; tablecens*gender/ norownocolnopercent; formatcenscensfmt. gendergrpfmt.; proclifereg data=pop covoutoutest=survres; modellos*censor(1)=gender /dist=exponential; run; RESUL TS: TABLEOFCENSBYGENDER CENS GENDER Frequency|Group 0|Group1|Total |(F) |(M) | ---------+--------+--------+ Event |902|367|1269 ---------+--------+--------+ Censored |271|51|322 ---------+--------+--------+ Total 1173 418 1591 380 PROCLIFEREGRESULTS: Exponential HazardModelforNursing HomePatients Lifereg Procedure DataSet =WORK.MORRIS2 Dependent Variable=Log(LOS) Censoring Variable=CENS Censoring Value(s)= 1 Noncensored Values= 1269RightCensored Values= 322 LeftCensored Values= 0Interval Censored Values= 0 LogLikelihood forEXPONENT -3320.476626 Lifereg Procedure Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value INTERCPT 15.84213388 0.033296 307860.0001Intercept GENDER 1-0.5161878 0.061915 69.50734 0.0001 SCALE 0 1 0 Extreme valuescale Notethattheestimatesfor¯0and¯1aboveare theoppositesofwhatwecalculated.I'llexplain whytheoutputhasthisformwhenwegetto AFTmodels. 381 TheWeibullRegressionModel Atthebeginning ofthecourse,wesawthatthesurvivorship function foraWeibullrandomvariableis: S(t)=exp[¡¸(t·)] andthehazardfunction is: ¸(t)=·¸t(·¡1) TheWeibullregression modelassumes thatforsomeone with covariatesZi,thesurvivorshipfunction is S(t;Zi)=exp[¡ª(Zi)(t·)] whereª(Zi)isde¯nedasinexponentialregression tobe: ª(Zi)=exp[¯0+Zi1¯1+Zi2¯2+:::Zip¯p] Forthe2-sample problem, wehave: ª(Zi)=exp[¯0+Zi1¯1] 382 WeibullMLEsforthe2-sampleproblem: Log-likelihood: logL=nX i=1±ilogh ·exp(¯0+¯1Zi)X·¡1 ii ¡nX i=1X· iexp(¯0+¯1Zi) )exp(^¯0)=d0=t0· exp(^¯0+^¯1)=d1=t1· wheretj·=njX i=1X^· iamongnjsubjects ^¸0(t)=^·exp(^¯0)t^·¡1 ^¸1(t)=^·exp(^¯0+^¯1)t^·¡1 dHR=^¸1(t)=^¸0(t)=exp(^¯1) =exp0 B@d1=t1· d0=t0·1 CA 383 WeibullRegression: MeansandMedians MeanSurvivalTime FortheWeibulldistribution, E(T)=¸(¡1=·)¡[(1=·)+1]. ²ControlGroup: T0=^¸(¡1=^·) 0¡[(1=^·)+1] ²TreatmentGroup: T1=^¸(¡1=^·) 1¡[(1=^·)+1] MedianSurvivalTime FortheWeibulldistribution, M=median=·¡log(0:5) ¸¸1=· ²ControlGroup: ^M0=2 64¡log(0:5) ^¸03 751=^· ²TreatmentGroup: ^M1=2 64¡log(0:5) ^¸13 751=^· where^¸0=exp(^¯0)and^¸1=exp(^¯0+^¯1). 384 Note:thesymbol¡isthe\gamma" function. Ifxisan integer,then ¡(x)=(x¡1)! Incaseswherexisnotaninteger,thisfunction hastobe evaluatednumerically . TheWeibullregression modelisveryeasyto¯t: ²Insas:usemodeloptiondist=weibull withinthe proclifereg procedure ²Instata:Justspecifydist(weibull) instead ofdist(exp) withinthestregcommand Note:togetmoreinformation onthesemodelingprocedures, usetheonlinehelpfacilities. Forexample, inStata,you cantype: .helpstreg 385 WeibullinStata: .streggender, dist(weibull) nohr failure _d:fail analysis time_t:los Fitting constant-only model: Iteration 0:loglikelihood =-3352.5765 Iteration 1:loglikelihood =-3074.978 Iteration 2:loglikelihood =-3066.1526 Iteration 3:loglikelihood =-3066.143 Iteration 4:loglikelihood =-3066.143 Fitting fullmodel: Iteration 0:loglikelihood =-3066.143 Iteration 1:loglikelihood =-3045.8152 Iteration 2:loglikelihood =-3045.2772 Iteration 3:loglikelihood =-3045.2768 Iteration 4:loglikelihood =-3045.2768 Weibull regression --logrelative-hazard form No.ofsubjects = 1591 Numberofobs=1591 No.offailures = 1269 Timeatrisk = 386211 LRchi2(1) =41.73 Loglikelihood =-3045.2768 Prob>chi2 =0.0000 ------------------------------------------------------------------- ----- _t|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----- gender|.4138082 .0621021 6.663 0.000.2920903 .5355261 _cons|-3.536982 .0891809 -39.661 0.000-3.711773 -3.362191 ---------+--------------------------------------------------------- ----- /ln_p|-.4870456 .0232089 -20.985 0.00-.5325343 -.4415569 ------------------------------------------------------------------- ----- p|.614439 .0142605 .5871152 .6430345 1/p|1.627501 .0377726 1.555127 1.703243 ------------------------------------------------------------------- ----- 386 WeibullinSAS proclifereg data=morris2 covoutoutest=survres; modellos*censor(1)=gender /dist=weibull; run; DataSet =WORK.MORRIS2 Dependent Variable=Log(LOS) Censoring Variable=CENS Censoring Value(s)= 1 Noncensored Values= 1269 RightCensored Values= 322 LeftCensored Values= 0 Interval Censored Values= 0 LogLikelihood forWEIBULL -3045.276811 Lifereg Procedure Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value INTERCPT 15.75644118 0.0542 11280.04 0.0001Intercept GENDER 1-0.6734732 0.101067 44.40415 0.0001 SCALE 11.62750085 0.037773 Extreme valuescale 387 InSAS,boththeexponentialandWeibullarespecialcases ofthegeneralclassofacceleratedlifemodelsandthe parameter interpretations followfromthisapproach. Totranslate theoutputofSAS(orStatausingtheereg command) forWeibullregression, wehavetotakethenega- tiveofthenumbersintheoutput,dividedbythe\scale" parameter (¾,or1=·). ²^¯0=¡intercpt=scale ²^¯1=¡covariate=scale Thenwecalculate theestimated HRasexp(^¯1). TheMLE'sare: ²^¯0=¡intercpt=scale=¡5:756=1:627=¡3:537 ²^¯1=¡covariate=scale=0:6735=1:625=0:414 andtheestimated HRisdHR=exp(^¯1)=exp(0:414)= 1:513. 388 WeibullRegression: VarianceEstimatesandTestStatistics Itisnotsoeasytogetvarianceestimates fromtheoutput ofproclifereg inSASorweibullinstata,atleastfor theparameters we'reinterested in. Thevariances dependoninvertinga(3£3)matrixcorre- spondingtotheparameters ¯0,¯1,and·.TheMLEfor^· hastobeobtained numerically (i.e.,noclosedform),sothe standard errorsalsohavetobeobtained bycomputer. Mainobjective:toobtains:e:(^¯1),sothatwecanform testsandcon¯dence intervalsforthehazardratio. Theoutputgivesus^¯¤ 1ands:e:(^¯¤ 1),where^¯1=¡^¯¤ 1=^¾.If ¾wasaconstant,thenwecouldjustcompute var(^¯1)=1 ^¾2var(^¯¤ 1) but¾isalsoarandomvariable! Instead, youneedtouse anapproximation forthevarianceofaratiooftworandom variables: var(^¯1)=1 ^¾4· ^¾2var(^¯¤ 1)+(^¯¤ 1)2var(^¾)¡2^¯¤ 1^¾cov(^¯¤ 1;^¾)¸ whereyougetvar(^¯¤ 1)andvar(^¾)bysquaring thestandard errorsofthecovariatetermandscaleterm,respec- tively,fromtheproclifereg orweibulloutput. 389 ComparisonofExponentialwithKaplan-Meier WecanseehowwelltheExponentialmodel¯tsbycompar- ingthesurvivalestimates formalesandfemalesunderthe exponentialmodel,i.e.,P(T¸t)=e(¡^¸zt),totheKaplan- Meiersurvivalestimates: S u r v i v a l 0.00.10.20.30.40.50.60.70.80.91.0 Length of Stay (days)010020030040050060070080090010001100 390 ComparisonofWeibullwithKaplan-Meier WecanseehowwelltheWeibullmodel¯tsbycomparing thesurvivalestimates,P(T¸t)=e(¡^¸zt^·),totheKaplan- Meiersurvivalestimates. S u r v i v a l 0.00.10.20.30.40.50.60.70.80.91.0 Length of Stay (days)010020030040050060070080090010001100 Whichdoyouthink¯tsbest? 391 Otherusefulplotsforevaluating¯ttoexponen- tialandWeibullmodels ²¡log(^S(t))vst ²log[¡log(^S(t))]vslog(t) Whyaretheseuseful? IfTisexponential,thenS(t)=exp(¡¸t)) solog(S(t))=¡¸t and ¤(t)=¸t astraightlineintwithslope¸andintercept=0 IfTisWeibull,thenS(t)=exp(¡(¸t)·) solog(S(t))=¡¸t· then ¤(t)=¸t· and log(¡log(S(t)))=log(¸)+·¤log(t) astraightlineinlog(t)withslope·andinterceptlog(¸). 392 Sowecancalculate ourestimated ¤(t)andplotitversust, andifitseemstoformastraightline,thentheexponential distribution isprobably appropriate forourdataset. Plotsfornursinghomedata:^¤(t)vstNegative Log SDF 0.00.20.40.60.81.01.21.41.61.82.0 LOS0100200300400500600700800900100011001200 393 Orwecanplotlog^¤(t)versuslog(t),andifitseemsto formastraightline,thentheWeibulldistribution isprobably appropriate forourdataset. Plotsfornursinghomedata:log[¡log(^S(t))]vslog(t)Log Negative Log SDF -0.5-0.4-0.3-0.2-0.10.00.10.20.30.40.50.60.7 Log of LOS4.504.755.005.255.505.756.006.256.506.757.007.25 394 ComparisonofMethods fortheTwo-sampleproblem: Data: ZiSubjectsEventsFollow-up Group0:Zi=0n0d0t0=Pn0i=1Xi Group1:Zi=1n1d1t1=Pn1i=1Xi InGeneral: ¸z(t)=¸(t;Z=z)forz=0or1: ThehazardratedependsonthevalueofthecovariateZ. Inthiscase,weareassuming thatweonlyhaveasingle covariate,anditisbinary(Z=1orZ=0) 395 MODELS ExponentialRegression: ¸z(t)=exp(¯0+¯1Z) )¸0=exp(¯0) ¸1=exp(¯0+¯1) HR=exp(¯1) WeibullRegression: ¸z(t)=·exp(¯0+¯1Z)t·¡1 )¸0=·exp(¯0)t·¡1 ¸1=·exp(¯0+¯1)t·¡1 HR=exp(¯1) ProportionalHazardsModel: ¸z(t)=¸0(t)exp(¯1) )¸0=¸0(t) ¸1=¸0(t)exp(¯1) HR=exp(¯1) 396 Remarks ²ExponentialmodelisaspecialcaseoftheWeibullmodel with·=1(note:Collettuses°insteadof·) ²ExponentialandWeibullmodelsarebothspecialcases oftheCoxPHmodel. Howcanyoushowthis? ²IfeithertheexponentialmodelortheWeibullmodelis valid,thenthesemodelswilltendtobemoree±cient thanPH(smaller s.e.'sofestimates). Thisisbecause theyassumeaparticular formfor¸0(t),ratherthanes- timating itateverydeathtime. 397 FortheExponentialmodel,thehazards areconstantover time,giventhevalueofthecovariateZi: Zi=0)^¸0=exp(^¯0) Zi=1)^¸0=exp(^¯0+^¯1) FortheWeibullmodel,wehavetoestimate thehazardasa function oftime,giventheestimates of¯0;¯1and·: Zi=0)^¸0(t)=^·exp(^¯0)t^·¡1 Zi=1)^¸1(t)=^·exp(^¯0+^¯1)t^·¡1 However,theratioofthehazardsisstilljustexp(^¯1),since theothertermscancelout. 398 Here'swhattheestimatedhazardslooklikefor thenursinghomedata: Exponential Hazard: Female Exponential Hazard: Male Weibull Hazard: Female Weibull Hazard: MaleH a z a r d R a t e 0.0000.0050.0100.0150.0200.0250.030 Length of stay (days)01002003004005006007008009001000 399 ComparisonwithProportionalHazardsModel .stcoxgender, nohr failure _d:fail analysis time_t:los Iteration 0:loglikelihood =-8556.5713 Iteration 1:loglikelihood =-8537.8013 Iteration 2:loglikelihood =-8537.5605 Iteration 3:loglikelihood =-8537.5604 Refining estimates: Iteration 0:loglikelihood =-8537.5604 Coxregression --Breslow methodforties No.ofsubjects = 1591 Numberofobs=1591 No.offailures = 1269 Timeatrisk = 386211 LRchi2(1) =38.02 Loglikelihood =-8537.5604 Prob>chi2 =0.0000 ------------------------------------------------------------------- ---- _t| _d|Coef. Std.Err. zP>|z|[95%Conf.Interval] ---------+--------------------------------------------------------- ---- gender|.3943588 .0621004 6.350 0.000.2726441 .5160734 ------------------------------------------------------------------- ---- ForthePHmodel,^¯1=0:394anddHR=e0:394=1:483. 400 ComparisonwiththeLogrankandWilcoxonTests .ststestgender failure _d:fail analysis time_t:los Log-rank testforequality ofsurvivor functions ------------------------------------------------ |Events gender|observed expected -------+------------------------- 0| 902 995.40 1| 367 273.60 -------+------------------------- Total|1269 1269.00 chi2(1) =41.08 Pr>chi2 =0.0000 .ststestgender, wilcoxon failure _d:fail analysis time_t:los Wilcoxon (Breslow) testforequality ofsurvivor functions ---------------------------------------------------------- |Events Sumof gender|observed expected ranks -------+-------------------------------------- 0| 902 995.40 -99257 1| 367 273.60 99257 -------+-------------------------------------- Total|1269 1269.00 0 chi2(1) =41.47 Pr>chi2 =0.0000 401 ComparisonofHazardRatiosandTestStatistics fore®ectofGender Wald Model/Metho d¸0¸1HRlog(HR) se(logHR)Statistic Exponential0.00290.00491.6760.5162 0.0619 69.507 Weibull t=50 0.00400.00601.5130.4138 0.0636 42.381 t=100 0.00300.00461.513 t=500 0.00160.00251.513 Logrank 41.085 Wilco xon 41.468 CoxPH Ties=Breslo w 1.4830.3944 0.0621 40.327 Ties=Discrete 1.4870.3969 0.0623 40.565 Ties=Efron 1.4860.3958 0.0621 40.616 Ties=Exact 1.4860.3958 0.0621 40.617 Score(Discrete) 41.085 402 ComparisonofMeanandMedianSurvival TimesbyGender MeanSurvivalMedianSurvival Model/MethodFemaleMaleFemaleMale Exponential 344.5205.6238.8142.5 Weibull 461.6235.4174.288.8 Kaplan-Meier 318.6200.714470 CoxPH 13172 (Kalb°eisch/Prentice) 403 TheAcceleratedFailureTimeModel Thegeneralformofanaccelerated failuretime(AFT)model is: log(Ti)=¯AFTZi+¾² where ²log(Ti)isthelogofasurvivaltime ²¯AFTisthevectorofAFTmodelparameters corre- spondingtothecovariatevectorZi ²²isarandom\error"term ²¾isascalefactor Inotherwords,wecanmodelthelog-survival timesasalinearfunctionofthecovariates. proclifereg inSASandthestregcommand instata (without theexponentialorweibulloption)allusethis\log- linear"modelformulationfor¯ttingparametric models. 404 Bychoosingdi®erentdistributions for²,wecanobtaindif- ferentparametric distributions: ²Exponential ²Weibull ²Gamma ²Log-logistic ²Normal ²Lognormal Wecancompare thepredicted survivalunderanyofthese parametric distributions totheKMestimated survivaltosee whichoneseemsto¯tbest. Oncewedecideonacertainclassofmodel(say,Gamma), wecanevaluatethecontributions ofcovariatesby¯nding theMLE's,andconstructing Wald,Score,orLRtestsofthe covariatee®ects. 405 WecanmotivatetheAFTmodelby¯rstdemonstrating the followingtworelationships: ²1.FortheExponentialModel: IfthefailuretimesTi=T(Zi)followanexponential distribution, i.e.,Si(t)=e¡¸itwith¸i=exp(¯Zi), then log(Ti)=¡¯Zi+² where²followsanextremevaluedistribution (whichjust meansthate²followsaunitexponentialdistribution). ²2.FortheWeibullModel: IfthefailuretimesTi=T(Zi)followaWeibulldistri- bution,i.e.,Si(t)=e¸it·with¸i=exp(¯Zi),then log(Ti)=¡¾¯Zi+¾² where²againfollowsanextreme valuedistribution, and ¾=1=·. Inotherwords,boththeExponentialandWeibullmodelcan bewrittenintheformofalog-linear modelforthesurvival times,ifwechoosetherightdistribution for². 406 Thelog-linear formfortheexponentialcanbederivedby: (1)Creating anewvariableT0=TZ£exp(¯Zi) (2)TakingthelogofTZ,yieldinglog(TZ)=logà T0 exp(¯Zi)! Step(1):Foranexponentialmodel,recallthat: Si(t)=Pr(TZ¸t)=e¡¸t;with¸=exp(¯Zi) ItfollowsthatT0»exp(1): S0(t)=Pr(T0¸t)=Pr(TZ¢exp(¯Z)¸t) =Pr(TZ¸texp(¡¯Z)) =exp[¡¸texp(¡¯Z)] =exp[¡exp(¯Z)texp(¡¯Z)] =exp(¡t) Step(2):Nowtakethelogofthesurvivaltime: log(TZ)=log0 B@T0 exp(¯Zi)1 CA =log(T0)¡log(exp(¯Zi)) =¡¯Zi+log(T0) =¡¯Zi+² where²=log(T0)followstheextremevaluedistribution. 407 RelationshipbetweenExponentialandWeibull IfTZhasaWeibulldistribution, i.e.,S(t)=e¡¸t· with¸=exp(¯Zi),thenyoucanshowthatthenewvariable T¤ Z=T· Z followsanexponentialdistribution withparameter exp(¯Zi). Basedontheprevious page,wecantherefore write: log(T¤)=¡¯Z+² (where²hasanextreme valuedistribution.) Butsincelog(T¤)=log(T·)=·£log(T),wecanwrite: log(T)=log(T¤)=· =(1=·)(¡¯Zi+²) =¡¾¯Zi+¾² where¾=1=·. 408 Thismotivatesthefollowinggeneralde¯nition ofthe AcceleratedFailureTimeModelby: log(Ti)=¯AFTZi+¾² where²isarandom\error"term,¾isascalefactor,Yis thelogofasurvivalrandomvariable,and ¯AFT=¡¾¯e where¯ecamefromthehazard¸=exp(¯Z). Thede¯ningfeatureofanAFTmodelis: S(t;Z)=Si(t)=S0(Át) Thatis,thee®ectofcovariatesistoaccelerate (stretch)ordecelerate (shrink)thetime-scale. E®ectofAFTonhazard: ¸i(t)=Á¸0(Át) 409 OnewaytointerprettheAFTmodelisviaitse®ecton mediansurvivaltimes.IfSi(t)=0:5,thenS0(Át)=0:5. Thismeans: Mi=ÁM0 Interpretation: ²ForÁ<1,thereisanacceleration oftheendpoint (ifM0=2yrsincontrolandÁ=0:5,thenMi=1yr. ²ForÁ>1,thereisastretchingordelayinendpoint ²Ingeneral, thelifetimeofindividualiisÁtimeswhat theywouldhaveexperienced inthereference group SinceÁmustbepositiveandafunction ofthecovariates,we modelÁ=exp(¯Zi). 410 WhendoesProportionalhazards=AFT? According totheproportionalhazardsmodel: S(t)=S0(t)exp(¯Zi) andaccording totheaccelerated failuretimemodel: S(t)=S0(texp(¯Zi)) SayTi»Weibull(¸;·).Then¸(t)=¸·t(·¡1) UndertheAFTmodel: ¸i(t)=Á¸0(Át) =e¯Zi¸0(e¯Zit) =e¯Zi¸0·Ã e¯Zit!(·¡1) =à e¯Zi!· ¸0·t(·¡1) =à e¯Zi!· ¸0(t) ButthislooksjustlikethePHmodel: ¸i(t)=exp(¯¤Zi)¸0(t) ItturnsoutthattheWeibulldistribution (andexponential, sincethisisjustaspecialcaseofaWeibullwith·=1) istheonlyoneforwhichtheaccelerated failuretimeand proportionalhazardsmodelscoincide. 411 SpecialcasesofAFTmodels ²Exponentialregression: ¾=1,²followingtheextreme valuedistribution. ²Weibullregression:¾arbitrary ,²followingtheextreme valuedistribution. ²Lognormal regression: ¾arbitrary ,²followingthenor- maldistribution. Examplesinstata:Usingthestregcommand, one hasthefollowingoptionsofdistributions forthelog-surviv al times: .stregtrt,dist(lognormal) ²exponential ²weibull ²gompertz ²lognormal ²loglogistic ²gamma 412 .streggender, dist(exponential) nohr ------------------------------------------------------------------- ----------- _t|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- gender|.516186 .0619148 8.337 0.000 .3948352 .6375368 ------------------------------------------------------------------- ----------- .streggender, dist(weibull) nohr ------------------------------------------------------------------- ----------- _t|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- gender|.4138082 .0621021 6.663 0.000 .2920903 .5355261 1/p|1.627501 .0377726 1.555127 1.703243 ------------------------------------------------------------------- ----------- .streggender, dist(lognormal) ------------------------------------------------------------------- ----------- _t|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- gender|-.6743434 .1127352 -5.982 0.000 -.8953002 -.4533866 _cons|4.957636 .0588939 84.179 0.000 4.842206 5.073066 sigma|1.94718 .040584 1.86924 2.028371 ------------------------------------------------------------------- ----------- .streggender, dist(gamma) ------------------------------------------------------------------- ----------- _t|Coef. Std.Err. zP>|z| [95%Conf.Interval] ---------+--------------------------------------------------------- ----------- gender|-.6508469 .1147116 -5.674 0.000 -.8756774 -.4260163 _cons|4.788114 .1020906 46.901 0.000 4.58802 4.988208 sigma|1.97998 .0429379 1.897586 2.065951 ------------------------------------------------------------------- ----------- 413 Thisgivesagoodideaofthesensitivit yofthetestofgender tothechoiceofmodel.Itisalsoeasytogetpredicted sur- vivalcurvesunderanyoftheparametric modelsusingthe following: .streggender, dist(gamma) .stcurv, survival Theoptions hazard andcumhaz canalsobesubstituted forsurvivalabovetoobtainplots. 414 AFTmodelsinSAS proclifereg data=pop covout outest=survres; modellos*censor(1)=gender /dist=exponential; modellos*censor(1)=gender /dist=weibull; modellos*censor(1)=gender /dist=gamma; modellos*censor(1)=gender /dist=normal; Otheroptionsarelognormal, logistic,andlog-logistic. The defaultistomodellogofresponse.Canspecify"NOLOG" fornolog-transformation. Inthiscase,"normal" isthesame as"lognormal." 415 Lifereg Procedure LogLikelihood forEXPONENT -3320.476626 Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value INTERCPT 15.84213388 0.033296 307860.0001Intercept GENDER 1-0.5161878 0.061915 69.50734 0.0001 SCALE 0 1 0 Extreme valuescale Lagrange Multiplier ChiSquare forScale337.5998 Pr>Chiis0.0001. LogLikelihood forWEIBULL -3045.276811 Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value INTERCPT 15.75644118 0.0542 11280.04 0.0001Intercept GENDER 1-0.6734732 0.101067 44.40415 0.0001 SCALE 11.62750085 0.037773 Extreme valuescale LogLikelihood forGAMMA-2970.388508 Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value INTERCPT 14.78811071 0.104333 2106.114 0.0001Intercept GENDER 1-0.6508468 0.114748 32.17096 0.0001 SCALE 11.97998063 0.043107 Gammascaleparameter SHAPE 1-0.1906006 0.094752 Gammashapeparameter LogLikelihood forNORMAL-9593.512838 Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value INTERCPT 1303.824624 9.919629 938.1129 0.0001Intercept GENDER 1-107.09585 18.97784 31.84577 0.0001 SCALE 1330.093584 6.918237 Normalscaleparameter 416 Designing aSurvivalStudy Wewillfocusonthepoweroftestsbasedontheexponential distribution andthelogranktest. ²Asinstandard designs,thepowerdependson {TheTypeIerror(signi¯cance level) {Thedi®erence ofinterest,¢,underHa. ²Anotabledi®erence fromtheusualscenarioisthatpower dependsonthenumberoffailures ,notthetotal samplesize. ²Inpractice, designing asurvivalstudyinvolvesdeciding howmanypatientsorindividuals toenter,aswellashow longtheyshouldbefollowed. ²Designsmaybe¯xedsamplesizeorsequential (Moreonthislater!) References: Collett Chapter 12 PocockChapter 9ofClinicalTrials Williams Chapter 10ofAIDSClinicalTrials (eds.FinkelsteinandSchoenfeld) 417 Reviewofpowercalculationsfor2-samplenormal Supposewehavethefollowingdata: Group1:(Y11;:::Y1n1) Group0:(Y01;:::Y0n0) andmakethefollowingassumptions: Y1j»N(¹1;¾2)Y0j»N(¹0;¾2) Ourobjectiveistotest: H0:¹1=¹0)H0:4=0where4=¹1¡¹0 Thestandard testisbasedontheZstatistic: Z=Y1¡Y0r s2(1 n1+1 n0) wheres2isthepooledsamplevariance(weareassuming equalvariances here).Thisteststatistic followsaN(0;1) distribution underH0. Ifthesamplesizesareequalinthetwoarms,n0=n1=n=2, (whichwillmaximize thepower),thenwehavethesimpler form: Z=Y1¡Y0s s2(1 n=2+1 n=2)=Y1¡Y0 2s=pn 418 Thestepstofollowincalculating thesamplesizeare: (1)Determine thecriticalvalue,c,forrejecting thenull whenitistrue. (2)Calculate theprobabilit yofrejecting thenullwhenthe alternativ eistrue,substituting cfromabove. (3)Rewritetheexpression intermsofthesamplesizefora givenpower. Step(1): Setthesigni¯cance level,®,equaltotheprobabilit yofre- jectingthenullhypothesiswhenitistrue: ®=Pr(jY1¡Y0j>cjH0) =Pr0 B@jY1¡Y0j 2s=pn>c 2s=pnjH01 CA =Pr0 B@jZj>c 2s=pn1 CA=2¢©0 B@c 2s=pn1 CA soz1¡®=2=c 2s=pn orc=z1¡®=22spn Notethatz°isthevaluesuchthat©(z°)=Pr(Z<z°)= °. 419 Step(2): Calculate theprobabilit yofrejecting thenullwhenHais true.Startoutbywritingdowntheprobabilit yofaTypeII error: ¯=Pr(acceptH0jHa) so1¡¯=Pr(rejectH0jHa) =Pr(jY1¡Y0j>cjHa) =Pr0 B@jY1¡Y0j¡¢ 2s=pn>c¡¢ 2s=pnjHa1 CA =Pr0 B@Z>c¡¢ 2s=pn1 CA sowegetz¯=¡z1¡¯=c¡¢ 2s=pn NowwesubstitutecfromStep(1): ¡z1¡¯=z1¡®=22s=pn¡¢ 2s=pn =z1¡®=2¡¢ 2s=pn 420 Step(3): Nowrewritetheequation intermsofsamplesizeforagiven power,1¡¯,andsigni¯cance level,®: z1¡®=2+z1¡¯=¢ 2s=pn =¢pn 2s =)n=(z1¡®=2+z1¡¯)24s2 ¢2 Notes: Thepowerisanincreasing function ofthestandardized dif- ference: ¹T(4)=4 2s=pn Thisisjustthenumberofstandard errorsbetweenthetwo means,undertheassumption ofequalvariances. 1.Asnincreases, thepowerincreases. 2.For¯xedn,thepowerincreases with4. 3.For¯xednand4,thepowerdecreases withs. 4.Assigning equalnumbersofpatientstothetwogroups (n1=n0=n=2)isbestintermsofmaximizing power. 421 AnExample: n=µ z1¡® 2+z1¡¯¶24s2 42 Saywewanttoderivethetotalsamplesizerequired toyield 90%powerfordetecting adi®erence of0.5standard devia- tionsbetweenmeans,basedonatwo-sided0.05leveltest. ®=0.05 z1¡® 2=1.96 ¯=0.10 z1¡¯=z0:90=1.28 n=(1:96+1:28)24s2 42¼42s2 42 Fora0.5standard deviation di®erence, ¢=s=0:5,so n¼42 (0:5)2=168 Ifyouendupwithn<30,thenyoushouldbeusingthe t-distribution ratherthanthenormaltocalculate critical values,andthentheprocessisiterative. 422 SurvivalStudies:ComparingProportionsofEvents Insomecases,thesamplesizeforasurvivaltrialisbased onacrudecomparison oftheproportionofeventsatsome ¯xedpointintime. Inthiscase,wecanapplytheresultsjustshowntogetsample sizes,basedonthenormalapproximation tothebinomial: De¯ne: Pcprobabilit yofeventincontrolarmbytimet Peprobabilit yofeventin\experimental"armbytimet Thenumberofpatientsrequired pertreatmen tarmbased onachi-square testcomparing binomial proportionsis: N=fz1¡® 2q 2P(1¡P)+z1¡¯q Pe(1¡Pe)+Pc(1¡Pc)g2 (Pc¡Pe)2 whereP=(Pe+Pc)=2 (Thislooksslightlydi®erentbecausethevarianceisnotthe sameunderHoandHa,aswasthecaseinthenormalpre- viousexample.) 423 Notesoncomparingproportionsoffailures: ²Useofchi-square testisbestwhen0:2<Pe;Pc<0:8 ²Shouldhave¸15patientsineachcellofthe(2x2)table ²Forsmallersamplesizes,useFisher'sexacttesttomo- tivatepowercalculations ²E±ciency vslogranktestisnear100%forstudieswith shortdurations relativetothemedianeventtime Whatdoesthismeanintermsoftheevent rates?Highorlow? ²Calculation ofsamplesizeforcomparing proportionsof- tenprovidesanupperboundtothosebasedoncompar- isonofsurvivaldistributions 424 Samplesizebasedonthelogranktest Recap: Consider atwogroupsurvivalproblem, withequal numbersofindividuals inthetwogroups(sayn0ingroup0 andn1ingroup1).Let¿1;:::;¿KrepresenttheKordered, distinctfailuretimes,andatthej-theventtime: Die/Fail Group Yes NoTotal 0d0jr0j¡d0jr0j 1d1jr1j¡d1jr1j Totaldjrj¡djrj whered0jandd1jarethenumberofdeaths(events)ingroup 0and1,respectively,atthej-theventtime,andr0jandr1j arethecorrespondingnumbersatrisk. Thelogranktestis:(z-statistic version) ZLR=PK j=1(d1j¡ej) s PKj=1vj withej=djr1j=rj vj=r1jr0jdj(rj¡dj)=[r2 j(rj¡1)] 425 Distributionofthelogrankstatistic Supposethatthehazardratesinthetwogroupsare¸0(t) and¸1(t),withhazardratio µ=e¯=¸1(t) ¸0(t) andsupposeweareinterestedintestingHo:¯=ln(µ)=0 (whichisequivalenttotestingHo:µ=1.) [Note:wewilluseln(µ)ratherthan¯inthefollowing,sothatthere isnoconfusionwiththeTypeIIerrorrate] Itispossibletoshowthat ²iftherearenoties,and ²weare\near"H0: then: ²E(d1j¡ejjd1j;d0j;r1j;r0j)¼ln(µ)=4 ²vj¼1=4 So,atapointln(µ)inthealternativ e,weget: ZLR¼PK j=1ln(µ)=4 rPKj=11=4=dln(µ)=4 r d=4=p dln(µ) 2 andZLR»N(ln(µ)p d=2;1) 426 HeuristicProof: E(d1jjd1j;d0j;r1j;r0j)=Pr(d1j=1jdj=1;r1j;r0j) =r1j¸0µ r1j¸0µ+r0j¸0 =r1jµ r1jµ+r0j =r1j r1j+r0j+ln(µ)2 64r1jr0j (r1j+r0j)23 75 Butej=r1j=(r1j+r0j),so: E(d1jjd1j;d0j;r1j;r0j)¡ej=ln(µ)2 64r1jr0j (r1j+r0j)23 75 Ifn0=n1,thennearH0:,r1j¼r0j,hence, E(d1jjd1j;d0j;r1j;r0j)¡ej=ln(µ)=4 Similarly ,withnoties,wehave vj=r1jr0j=r2 j¼1=4 427 Thiscanalsobederivedviathepartiallikelihood: Wecanwritethepartiallikelihoodas: l(¯)=log2 664nY j=10 B@e¯Zj P `2R(¿j)e¯Z`1 CA±j3 775 =nX j=1±j2 64¯Zj¡log0 B@X `2R(¿j)e¯Z`1 CA3 75 andthenthe\score"(partialderivativeoflog-likelihood)becomes: U(¯)=@ @¯`(¯) =nX j=1±j2 64Zj¡P `2R(¿j)Z`e¯Z` P `2R(¿j)e¯Z`3 75 Wecanwritethe\information" (minussecondpartialderivative ofthelog-likelihood)as: ¡@2 @¯2`(¯)=nX j=1±j2 64P `2R(¿j)e¯Z`P `2R(¿j)Z`e¯Z`¡(P `2R(¿j)Z`e¯Z`)2 P `2R(¿j)e¯Z`3 75 Thelogrankstatistic(withnoties)isequivalenttothescorestatistic fortesting¯=0: ZLRU(0)p I(0) ByaTaylorseriesexpansion: U(0)»=U(¯)¡¯@U @¯(0) E[U(0)]»=¯d=4andI(0)»=d=4 428 PoweroftheLogrankTest Usingasimilarargumen ttobefore,thepowerofthelogrank test(basedonatwo-sided®leveltest)isapproximately: Power(µ)¼1¡©· z1¡® 2¡ln(µ)p d=2¸ Note:Powerdependsonlyondandµ! Wecaneasilysolvefortherequired numberofeventsto achieveacertainpowerataspeci¯edvalueofµ: Toyieldpower(µ)=1¡¯,wewantdsothat 1¡¯=1¡©µ z1¡® 2¡ln(µ)p d=2¶ )z¯=z1¡® 2¡ln(µ)p d=2 )d=4µ z1¡® 2¡z¯¶2 [ln(µ)]2 ord=4µ z1¡® 2+z1¡¯¶2 [ln(µ)]2 429 Example: Saywewereplanning a2-armstudy,andwantedtobeable todetectahazardratioof1.5with90%powerata2-sided signi¯cance levelof®=0:05. Required numberofevents: d=4µ z1¡® 2+z1¡¯¶2 [ln(µ)]2 =4(1:96+1:282)2 [ln(1:5)]2 ¼42 0:1644=256 #EventsrequiredforvariousHazardRatios HazardPower Ratio 80% 90% 1.5 191 256 2.0 66 88 2.5 38 50 3.0 26 35 Moststudiesaredesigned todetectahazardratioof1.5-2.0. 430 PracticalConsiderations ²Howdowedecideonµ? ²Howdowetranslate numbersoffailurestonumbersof patients? Hazardratiosfortheexponentialdistribution Thehazardratiofromtwoexponentialdistributions canbe easilytranslated intomoreintuitivelyinterpretable quanti- ties: Median: IfTi»exp(¸i),then Median(Ti)=¡ln(0:5)=¸i Itfollowsthat Median(T1) Median(T0)=¸0 ¸1=e¡¯=1 µ Hence,doubling themediansurvivalofatreatedcompared toacontrolgroupwillcorrespondtohalvingthehazard. 431 R-yearsurvivalrates SupposetheR-yearsurvivalrateingroup1isS1(R)andin group0isS0(R).Undertheexponentialmodel: Si(R)=exp(¡¸iR) Hence, ln(S1(R)) ln(S0(R))=¡¸1R ¡¸0R=¸1 ¸0=e¯=µ Hence,doubling thehazardratefromgroup1togroup0will correspondtodoubling thelogoftheR-yearsurvivalrate. NotethatthisresultdoesnotdependonR!. Example: Supposethe5-yearsurvivalrateontreatmen tA is20%andwewant90%powertodetectanimprovementof thatrateto30%.Thecorrespondinghazardratiooftreated tocontrolis: ln(0:3) ln(0:2)=¡1:204 ¡1:609=0:748 Fromourprevious formula,thenumberofevents(deaths) neededtodetectthisimprovementwith90%power,based ona2-sided5%leveltestis: d=4(1:96+1:282)2 [ln(0:748)]2=499 432 TranslatingtoNumberofEnrolledPatients First,supposethatwewillenterNpatientsintoourstudy attime0,andwillthencontinuethestudyforFunitsof time. UnderH0,theprobabilit ythatanindividual willfailduring thestudyis: Pr(fail)=ZF 0¸0e¡¸0tdt =1¡e¡¸0F Hence,ifourcalculations sayweneeddfailures, thento decidehowmanypatientstoenter,wesimplysolve d=(N=2)(1¡e¡¸0F)+(N=2)(1¡e¡¸1F) Tosolvetheaboveequation forN,weneedtosupplyvalues ofFandd.Inotherwords,herewearealreadydeciding whatHRwewanttodetect(withwhatpower,etc),andfor howlongwearegoingtofollowpatients.Whatwegetis thetotalnumberofpatientsweneedtoenrollinorderto observethedesirednumberofeventsinFunitsoffollow-up time. 433 Example: Supposewewanttodetecta50%improvement inthemediansurvivalfrom12monthsto18monthswith 80%powerat®=0:05,andweplanonfollowingpatients for3years(36months). Wecanusethetwomedians tocalculate boththeparameters ¸0and¸1andthehazardratio,µ: Median(Ti)=¡ln(0:5)=¸i so¸1=¡ln(0:5) M1=0:6931 18=0:0385 ¸0=¡ln(0:5) M0=0:6931 12=0:0578 µ=¸1 ¸0=0:0385 0:0578=12 18=0:667 andfromourprevious table,#eventsrequired isd=191 (sameforµ=1:5asitisfor1/1.5=0.667). Soweneedtosolve: 191=(N=2)(1¡e¡0:0578¤36)+(N=2)(1¡e¡0:0385¤36) =(N=2)(0:875)+(N=2)(0:7500)=(N=2)(1:625) )N=235 (forpractical reasons, wewouldprobably roundupto236 andrandomize 118patientstoeachtreatmen tarm) 434 Amorerealisticaccrualpattern Inreality,noteveryonewillenterthestudyonthesameday. Instead, theaccrualwilloccurina\staggered" mannerover aperiodoftime. Thestandardassumption: Supposeindividuals enterthestudyuniformly overanac- crualperiodlastingAunitsoftime,andthataftertheac- crualperiod,follow-upwillcontinueforanotherFunitsof time. TotranslatedtoN,weneedtocalculate theprobabilit ythat apatientfailsunderthisaccrualandfollow-upscenario. Pr(fail)=ZA 0Pr(failjenterata)f(a)da =1¡RA 0S(a+F)da A(2) Thensolve:d=(N=2)Pr(fail;¸0)+(N=2)Pr(fail;¸1) =(N=2)Pc+(N=2)Pe =(N=2)(Pc+Pe) IfwenowsolveforN(substituting informulaford),weget: N=2d (Pc+Pe) N=8µ z1¡® 2+z1¡¯¶2 [ln(µ)]21 (Pc+Pe) 435 HowcanwegetPcandPefrom(2)? Ifweassumethattheexponentialdistribution holds,then wecansolve(2)toobtain: Pi=1¡exp(¡¸iF)(1¡exp(¡¸iA)) ¸iA(3) (fori=c;e) Freedman suggested anapproximation forPcandPe,by computing theprobabilit yofaneventatthemedianduration offollow-up,(A=2+F): Pi=Pr(fail;¸i)=1¡exp[¡¸i(A=2+F)](4) Heshowedthatthisapproximation worksprettywellforthe exponentialdistribution (i.e.,itgivesvaluescloseto(3)). 436 Analternativeformulation Rubenstein,Gail,andSantner(1981)suggestthefollowing approachforcalculating thetotalsamplesizethatmustbe enrolled: N=2µ z1¡® 2+z1¡¯¶2 [ln(µ)]22 41 Pc+1 Pe3 5 wherePcandPearetheexpectedproportionofpatientsor individuals whowillfail(haveanevent)onthecontroland treatmen tarms. Howdowecalculate (estimate)PcandPe? ²usingthegeneralformulaforadistribution Sgivenin (2) ²usingtheexactformulaforanexponentialdistribution givenin(3) ²usingtheapproximation givenby(4) Note:alloftheseformulascanbemodi¯edforunequal assignmen ttotreatmen t(orexposure)groupsbychanging (N=2)intheformulasonp.17-19to(qc¤N)and(qe¤N), whereqcandqearetheproportionsassigned tothecontrol andexposedgroups,respectively. 437 Freedman'sApproach(1982) Freedman's approachisbasedonthelogrankstatisticunder theassumption ofproportionalhazards, butdoesnotrequire theassumption ofexponentialsurvivaldistributions. Totalnumberofevents: d=µ z1¡® 2+z1¡¯¶20 @µ+1 µ¡11 A2 Totalsamplesize: N=2µ z1¡® 2+z1¡¯¶2 Pe+Pc0 @µ+1 µ¡11 A2 wherePeandPcareestimated using(4). Thisapproximation dependsontheassumption ofacon- stantratiobetweenthenumberofpatientsatriskinthetwo treatmen tgroupspriortoeacheventtime=)r0j¼r1j(as showninthe\heuristic proof").Whenthisassumption is notsatis¯ed, therequired samplesizestendtobeoveresti- mated. Q.Whenwouldthisassumptionnotbesatis¯ed? A.Whenthesmallestdetectabledi®erenceislarge. 438 Someexamplesofstudydesign ExampleI: Aclinicaltrialinesophageal cancerwillrandomize patients toradiotherap yalone(RxA)versusradiotherap ypluschemother- apy(RxB).Thegoalofthestudyistocompare thetwo treatmen tswithrespecttosurvival,andweplantousethe logranktest.Fromhistorical data,weknowthatthemedian survivalonRXAforthisdiseaseisaround9months.We want90%powertodetectanimprovementinthismedian to18months.Paststudieshavebeenabletoaccrueap- proximately 50patientsperyear.Chooseasuitable study design. 439 ExampleII: Aclinicaltrialinearlystagebreastcancerwillrandomize patientsaftertheirsurgerytoTamoxifen(ARMA)versus observationonly(ARMB).Thegoalofthestudyistocom- parethetwotreatmen tswithrespecttotimetorelapse,and thelogranktestwillbeusedintheanalysis. Fromhistorical data,weknowthatafter¯veyears,65%ofthepatientswill stillbediseasefree.Wewouldliketohave90%powerto detectanimprovementinthisdiseasefreerateto75%.Past studieshavebeenabletoaccrueapproximately 200patients peryear.Chooseasuitablestudydesign. 440 ExampleIII: Someinvestigators intheenvironmen talhealthdepartmen t wanttoconduct astudytoassessthee®ectsofexposureto tolueneontimetopregnancy .Theywillconduct acohort studyinvolvingwomenwhoworkinachemical factoryin China. Itisestimated that20%ofthewomenwillhave workplace exposuretotoluene. Furthermore, itisknown thatamongunexposedwomen,80%willbecomepregnant withinayear.Theinvestigators willbeabletoenroll200 womenperyearintothestudy,andplananadditional year offollow-upattheendofaccrual. Assuming theyhave2 yearsaccrual, whatreduction inthe1-yearpregnancy rate forexposedwomenwilltheybeabletodetectwith85% power?Whatiftheyhave3yearsofaccrual? 441 Otherimportantissues: Theapproachesjustdescribedaddressthebasicquestion of calculating asamplesizeforstudywithasurvivalendpoint. Theseapproachesoftenneedtobemodi¯edslightlytoad- dressthefollowingcomplications: ²Losstofollow-up ²Non-compliance (orcross-overs) ²Strati¯cation ²Sequentialmonitoring ²Equivalencehypotheses Nextwesummarize someofthemainpoints. 442 Losstofollow-up Ifsomepatientsarelosttofollowup(asopposedtocensored attheendofthetrialwithouttheevent),thepowerwillbe decreased. Therearetwomainapproachesfordealingwiththis: ²Simplein°ationmethod-If`*100%ofpatients areanticipated tobelosttofollowup,calculate target samplesizetobe N¤=0 @1 1¡`1 A¢N Example: SayyoucalculateN=200,andanticipate lossesof20%.Thesimplein°ation methodwouldgive youatargetsamplesizeofN¤=(1=0:8)¤200=250. Warning: peopleoftenmakethemistakeofjustin- °atingtheoriginalsamplesizeby`*100%,whichwould havegivenN¤=240fortheexample above. ²Exponentiallossassumption -theaboveapproach assumes thatlossescontributeNOinformation. Butwe actuallyhaveinformation onthemupuntilthetimethat theyarelost.Incorporatethisbyassuming thattimeto lossalsofollowsanexponentialdistribution, andmodify PeandPc. 443 Noncompliance Ifsomepatientsdon'ttaketheirassigned treatmen ts,the powerwillbedecreased. Thisissuehastwosides: ²Drop-outs(de)-patientswhocannottolerate the medication stoptakingit;theirhazardratewouldbe- comethesameastheplacebogroup(ifincluded instudy) atthatpoint. ²Drop-ins(dc)-patientsassigned tolesse®ectivether- apymaynotgetrelieffromsymptoms andseekother therapy,orrequesttocross-over. Aconservativeremedy{adjustPeandPcasfollows: P¤ e=Pe(1¡de)+Pcde P¤ c=Pc(1¡dc)+Pedc 444 DesignStrategy: 1.Decideon ²TypeIerror(signi¯cance level) ²clinically importantdi®erence (intermsofHR) ²desiredpower 2.Determine thenumberoffailuresneeded 3.Basedonpastexperience ²decideonareasonable distribution forthecontrols (usually exponential) ²estimate anticipated accrualperunittime ²estimate expectedrateoflosstofollowup VarythevaluesofAandFuntilyougetsomething prac- ticallyfeasiblethatgivestherightnumberoffailures. 4.Consider noncompliance, sequentialmonitoring, andother issuesimpacting samplesize 445 Included onthenextseveralpagesisaSASprogram tocal- culatesamplesizesforsurvivalstudies.Itusesseveralofthe approacheswe'vediscussed, including: ²Rubenstein,GailandSantner(RGS,1981) ²Freedman (1982) ²LachinandFoulkes(1986) Acopyofthisprogram isshownonthenextseveralpages. Theprogram requiresentryof: ²Signi¯cance level(alpha) ²Power ²Sides(1forone-sided test,2fortwo-sidedtest) ²Accrualperiod ²Followupperiod ²Yearlyrateoflosstofollow-up ²Proportionrandomized toexperimentaltreatmen tarm ²Oneofthefollowing: {Yearlyeventrateoncontrolandexperimentaltreat- mentarms {Yearlyeventrateoncontrolarm,andthehazard ratio {Mediantimetoeventoncontrolandexperimental treatmen tarms 446 TheSASprogram rgsnew.sas datargs; ******************************************************************* ****; ***enterthefollowing information inthisblock; alpha=0.05; /*significance level*/ sides=2; /*one-sided ortwo-sided test*/ power=0.90; /*Desired power*/ accrual =2; /*Accrual periodinyears*/ fu=1.5; /*Followupafterlastpatient isaccrued */ loss=0.0; /*yearlyrateofloss*/ qe=0.5; /*proportion randomized toexperimental arm*/ ***eitherenterthemediantimetoeventinyearsoncontrol; ***ortheyearlyeventrate-leavetheothervaluemissing; medianc =0.75; /*mediantimetoeventoncontrol arm*/ probc=.; /*yearlyeventrateincontrol arm*/ ***eitherentertheyearlyeventrateintheexperimental arm; ***orthehazardratioforcontrol vsexperimental ; ***orthemediantimetoeventonexperimental arminyears ; ***leavetheothervaluesmissing (.); mediane =1.5; /*mediantimetoeventonexperimental */ probe=.; /*yearlyeventrateinexperimental arm*/ rr=.; /*hazardratio*/ ******************************************************************* ****; beta=1-power; qc=1-qe; zalpha=probit(1-alpha/sides); zbeta=probit(1-beta); ***calculate yearlyeventrateinbotharmsusingmedians, ifsupplied; ifmedianc^=. thendo; hazc=-log(0.5)/medianc; probc=1-exp(-hazc); end; ifmediane^=. thendo; haze=-log(0.5)/mediane; probe=1-exp(-haze); end; hazc=-log(1-probc); 447 ***calculate hazardinexperimental group,usingyearlyeventrate; ***orhazardratio; ifprobe^=. thenhaze=-log(1-probe); ifrr^=.thenhaze=hazc/rr; ifprobe^=. andhaze^=. thendo; put"**************************************************************"; put"WARNING: bothyearlyeventrateandhazardratio(HR)have"; put" beenspecified. Calculations willusetheHR"; put"**************************************************************" /; end; ***calculate mediansurvival timesifnotsupplied; medianc=-log(0.5)/hazc; mediane=-log(0.5)/haze; hazl=-log(1-loss); avghaz=qc*hazc +qe*haze; rr=hazc/haze; log_rr=log(hazc/haze); totloss=(accrual*0.5 +fu)*loss; ***compute expected probability ofdeath(event) duringtrial; ***givenstaggered accrual butNOloss; pc0loss =1-((exp(-hazc*fu)-exp(-hazc*(accrual+fu)))/(hazc*accrual) ); pe0loss =1-((exp(-haze*fu)-exp(-haze*(accrual+fu)))/(haze*accrual) ); ***compute expected probability ofeventduringtrial; ***givenstaggered accrual ANDloss; pc=(1-(exp(-(hazc+hazl)*fu)-exp(-(hazc+hazl)*(fu+accrual))) /((hazc+hazl)*accrual))*(hazc/(hazc+hazl)); pe=(1-(exp(-(haze+hazl)*fu)-exp(-(haze+hazl)*(fu+accrual))) /((haze+hazl)*accrual))*(haze/(haze+hazl)); pbar=(1-(exp(-(avghaz+hazl)*fu)-exp(-(avghaz+hazl)*(fu+accrual)) ) /((avghaz+hazl)*accrual))*(avghaz/(avghaz+hazl)); ***compute totalsamplesizeassuming loss; N=int(((zalpha+zbeta)**2)/(log_rr**2)*(1/(qc*pc)+1/(qe*pe) ))+1; ***Compute samplesizeusingmethodofFreedman (1982); N_FRD=int((2*(((rr+1)/(rr-1))**2)*(zalpha+zbeta)**2)/(2*(q e*pe+qc*pc)))+1; 448 ***Compute samplesizeusingmethodofLachinandFoulkes (1986); ***withratesunderH0givenbypooledhazard; N_LF=int((zalpha*sqrt((avghaz**2)*(1/pbar)*(1/qc +1/qe))+ zbeta*sqrt((hazc**2)*(1/(qc*pc)) +(haze**2)*(1/(qe*pe))))**2/ ((hazc-haze)**2)) +1; ***compute totalsamplesizeassuming noloss; n_0loss =int(((zalpha+zbeta)**2)/(log_rr**2)* (1/(qc*pc0loss)+1/(qe*pe0loss))) +1; ***compute samplesizeusingsimpleinflation methodforloss; naive=int(n_0loss/(1-totloss)) +1; ***verifythatactualpowerissameasdesired power; newpower =probnorm(sqrt(((N)*(log_rr**2))/(1/(qc*pc) +1/(qe*pe))) -zalpha); ifabs(newpower-power)>0.001 thendo; put'***WARNING: actualpowerisnotequaltodesired power'; put'Desired power:'power'Actualpower:'newpower; end; ******************************************************************* ****; ***compute numberofeventsexpected duringtrial; ******************************************************************* ****; ***Compute expected numberundernull; n_evth0 =int(n*pbar) +1; r=qc/qe; ***Rubinstein, GailandSantner (1981)method-simpleapproximation; n_evtrgs =int((((r+1)**2)/r)*((zalpha +zbeta)**2)/(log_rr**2)) +1; ***Freedman (1982); n_evtfrd =int((((rr+1)/(rr-1))**2) *(zalpha+zbeta)**2)+1; ***Usingbacktracking methodofLachinandFoulkes (1986); n_evt_c =int(N*qc*pc) +1; n_evt_e =int(N*qe*pe) +1; n_evtlf =n_evt_c +n_evt_e; 449 labelsides='Sides' alpha='Alpha' power='Power' beta='Beta' zalpha='Z(alpha)' zbeta='Z(beta)' accrual='Accrual (yrs)' fu='Follow-up (yrs)' loss='Yearly Loss' totloss='Total Loss' probc='Yearly eventrate:control' probe='Yearly eventrate:active' rr='Hazard ratio' log_rr='Log(HR)' N='Total Samplesize(RGS)' N_FRD='Total Samplesize(Freedman)' N_LF='Total Samplesize(L&F)' n_0loss='Sample size(noloss)' pc='Pr(event), control' pe='Pr(event), active' pc0loss='Pr(event| noloss),control' pe0loss='Pr(event| noloss),active' n_evth0='# events(Ho-pooled)' n_evtrgs='# events(RGS)' n_evtfrd='# events(Freedman)' n_evtlf='# events(L&F)' medianc='Median survival, control' mediane='Median survival, active' naive='Sample size(naiveloss)'; procprintdata=rgs labelnoobs; title'Sample size&expected eventsforcomparing twosurvival distributions'; title2'UsingmethodofRubinstein, GailandSanter(RGS,1981)'; title3'Freedman (1982), orLachinandFoulkes (L&F,1986)'; varsidesalphapoweraccrual fulosstotloss probcprobemedianc mediane pcpepc0loss pe0loss rrlog_rr n_evth0 n_evtrgs n_evtfrd n_evtlf NN_FRDN_LFn_0loss naive; formatpowerlosstotloss f4.2medianc mediane f5.3 probcproberrlog_rrpcpepc0loss pe0loss f6.4; 450 BacktoExampleI: Aclinicaltrialinesophageal cancerwillrandomize patients toradiotherap yalone(RxA)versusradiotherap ypluschemother- apy(RxB).Thegoalofthestudyistocompare thetwo treatmen tswithrespecttosurvival,andweplantousethe logranktest.Fromhistorical data,weknowthatthemedian survivalonRxAforthisdiseaseisaround9months.We want90%powertodetectanimprovementinthismedian to18months.Paststudieshavebeenabletoaccrueap- proximately 50patientsperyear.Chooseasuitable study design. First,let'swritedownwhatweknow: ²desiredsigni¯cance levelnotstated,souse®=0:05 (assume atwo-sidedtest) ²assumeequalrandomization totreatmen tarms (unlessotherwise stated) ²desiredpoweris90% ²mediansurvivaloncontrolis9months)M0=9 ²wanttodetectimprovementto18monthsonRxB) M1=18 ²Maximumaccrualperyearis50patients 451 Wehavealloftheinformation weneedtoruntheprogram, excepttheaccrualandfollowuptimes.Weneedtousetrial anderrortogetthese. NumberofTotalTotal Accrual Follow-up EventsSample Study PeriodPeriodRequired SizeDuration 1 2.5 88 106 3.5 2 1.5 88 115 3.5 2.5 1 88 122 3.5 3 0.5 88 133 3.5 3 1 88 117 4 Shownonthenextpageistheoutputfromrgsnew.sas us- ingAccrual=2, Follow-up=1.5. I'vegiventheRGSnumbers above. Whichoftheabovearefeasibledesigns? 452 Samplesize&expected eventsforcomparing twosurvival distributions UsingmethodofRubinstein, GailandSanter(RGS,1981) Freedman (1982), orLachinandFoulkes (L&F,1986) Yearly event Accrual Follow-up Yearly Total rate: Sides Alpha Power (yrs) (yrs) Loss Loss control 20.050.90 2 1.5 0.00 0.00 0.6031 Yearly event Median Median Pr(event| rate: survival, survival, Pr(event), Pr(event), noloss), active control active control active control 0.3700 0.750 1.500 0.8860 0.6737 0.8860 Pr(event| noloss), Hazard #events #events #events active ratio Log(HR) (Ho-pooled) (RGS) (Freedman) 0.6737 2.0000 0.6931 94 88 95 Total Sample Total Sample Total Sample size #events Sample size Sample size(no(naive (L&F) size(RGS) (Freedman) size(L&F) loss) loss) 90 115 122 121 115 115 453 Howdowepickfromthefeasibledesigns? The¯rst4designsallhave31/2yearstotalduration, since thefollow-upperiodstartsafterthelastpatienthasbeen accrued. Theshorterthefollow-upperiodgiventhis¯xed studyduration, themorepatientswehavetoenroll. Insomecases,itwillbemuchmorecost-e®ectiv etoenroll fewerpatientsandfollowthemforlonger.Thiscorresponds tocaseswheretheinitialcostperpatientisveryhigh. Inothercases(wheretheinitialcostperpatientislower),it willbebettertoenrollmorepatients.Themedianfollow-up forthe¯rst4designsare3,2.5,2.25,and2years,respec- tively.Thetotalcostoftreatmen tcouldbeestimated by multiplying thenumberofpatientsbythemedianfollow-up time. Someprefertokeeptheaccrualperiodasshortaspossi- ble,givenhowmanypatientscanfeasiblybeenrolled. This willtendtogivethesmallest numberofpatientsamongthe feasibledesigns. Whichdesignwouldthiscorrespondto? Another issuetothinkaboutiswhetherthebackground con- ditionsofthediseasearechangingrapidly(likeAIDS)orare fairlystable(likemanytypesofcancer). Fortheformersitu- ation,itwouldbebesttohaveastudywithashortduration sotheresultswillhavemoreinterpretation. 454 Usingtheinformation given,therearealotofotherquanti- tieswecancalculate: ²Thehazardratioofcontroltotreated is: median(Rx B) median(Rx A)=18 9=2 ²Thehazardratesforthetwotreatmen tarmsare: forRxA:¸0=¡log(0:5) median(Rx A)=¡log(0:5) 9=0:0770 forRxB:¸1=¡log(0:5) median(Rx B)=¡log(0:5) 18=0:0385 ²Theyearlyprobabilityofaneventis: forRxA:Pr(T<1j¸0)=1¡e(¡¸0¤t) =1¡e(¡0:0770¤12)=0:603 forRxB:Pr(T<1j¸1)=1¡e(¡¸1¤t) =1¡e(¡0:0385¤12)=0:370 Whatwouldhappenaboveifweusedtimet inyears(i.e.,t=1)insteadofmonths? Whatwouldhappenifwecalculatedboththe hazardrateandyearlyeventprobabilityusing timeinyears? 455 Basedonadesignwith2.5yearsaccrualand 1yearfollow-up: ²Themedianfollow-uptime medianFU=A=2+F =30=2+12=27months ²Theprobabilit yofaneventduringtheentirestudyis: (usingtheapproximation innotes) forRxA:Pc=1¡exp(¡¸0¤[A=2+F]) =1¡exp(¡0:0770¤27)=0:875 forRxB:Pe=1¡exp(¡¸0¤[A=2+F]) =1¡exp(¡0:0385¤27)=0:646 (theabovenumbersdi®erfromwhatyou'dgetinthe printoutfromtheprogram, sinceitcalculates theexact probabilit yundertheexponentialdistribution, instead ofusingtheapproximation) Inthecalculations above,allofthe\time"periodswerein termsofmonths.Youhavetoremembertokeepthescale thesamethroughout. Tousetheprogram, youneedtotranslate thetimescalein termsofyears.Soamedianof18monthssurvivalwouldbe enteredasmedian=1.5. 456 Whathappensifweaddlosstofollow-up? RequiredsamplesizeforA=2.5,FU=1year YearlyNumberofTotalTotal LosstoEventsSample Study Follow-upRequired SizeDuration 0 88 122 3.5 5% 88 128 3.5 10% 88 133 3.5 20% 88 147 3.5 457 SequentialDesignandAnalysisofsurvivalstud- ies Inclinicaltrialsandotherstudies,itisoftendesirable to conduct interimanalyses ofastudywhileitisstillongoing. Rationale: ²ethical: ifonetreatmen tissubstantiallyworsethan another, thenitiswrongtocontinuetogivetheinferior treatmen ttopatients. ²timelyreporting: ifthehypothesisofinteresthas beenclearlyestablished halfwaythroughthestudy,then scienceandthepublicmaybene¯tfromearlyreporting. WARNING!! Unplanned interimanalyses canseriously in°atethetrue typeIerrorofatrial.Ifinterimanalyses aretobeperformed, itisESSENTIAL tocarefully plantheseinadvance,andto adjustalltestsappropriately sothethetypeIerrorisofthe desiredsize. 458 HowdoesthetypeIerrorbecomein°ated? Consider atwogroupstudycomparing treatmen tsAandB. Supposethedataarenormally distributed (sayXi»N(¹A;¾2) ingroupA,andsimilarly forgroupB),sothatthenullhy- pothesisofinterestis H0:¹A=¹B Itisnottoohardto¯gureouthowthetypeIerrorcanget in°atedifanaiveapproachisused. SupposeweplantodoKinterimanalyses, andthatexactly mindividuals willentereachtreatmen tbetweeneachanal- ysis.Theteststatistic atthekthanalysiswillbe Zk=Pk i=1Pm j=1(XAij¡XBij)=km r 2¾=km=Pk i=1di=k r 2¾=km wherediisthedi®erence betweenthetwogroupmeansat theithanalysis, di=XAi¡XBi andXAiandXBiarethemeansingroupsAandBofthe mindividuals whoenteredinthei-thtimeperiod. 459 (naive)Interimmonitoringprocedure: ²Allowmpatientstoenteroneachtreatmen tarm (totalof2madditional patients) ²CalculateZkbasedonthecurrentdata ²RejectthenullhypothesisifjZkj>z1¡®=2,where®is thedesiredtypeIerror. TheoveralltypeIerrorrateforthestudyis: Pr(jZ1j>z1¡®=2orjZ2j>z1¡®=2...orjZKj>z1¡®=2) Ifthetestateachinterimanalysis isperformed atlevel®, thenclearlythisprobabilit ywillexceed®.Thetablebelow showstheTypeIerrorrateifeachtestisdoneat®=0:05 forvariousvaluesofK: Numberofinterimanalyses(K) 123451025 5%8.3%10.7%12.6%14.2%19.3%26.6% (fromLee,StatisticalMethodsforSurvivalData,Table 12.9) 460 Forsurvivaldata,thecalculations becomeMUCHmorecom- plicated sincethedatacollected withineachtimeinterval continuestochangeastimegoeson! WhatcanwedotoprotectagainstthistypeIerrorin°ation? PocockApproach: Pickasmallersigni¯cance level(say®0)touseateachinterim analysissothattheoveralltypeIerrorstaysatlevel®. Aproblem withthePocockmethodisthateventhevery lastanalysis hastobeperformed atlevel®0.Thistendsto beveryconservativeatthe¯nalanalysis. O'BrienandFlemingApproach: Apreferable approachwouldbetovarythealphalevelsused foreachoftheKinterimanalyses, andtrytokeepthevery lastone\close"tothedesiredoverallsigni¯cance level.The O'Brien-Fleming approachdoesthat. 461 Commentsandnotes: ²Thereareseveralotherapproachesavailableforsequen- tialdesignandanalysis. TheO'BrienandFleming approachisprobably themostpopularinpractice. ²Therearemanyvariations onthethemeofsequential design.Thetypewehavediscussed hereiscalledGroup sequentialanalysis . {Thereareotherapproachesthatrequirecontinuous analysisaftereachnewindividual entersthestudy! {Therearealsoapproacheswheretherandomization itselfismodi¯edasthetrialproceeds.E.g.Ze- len's\Playthewinnerrule"(NewEngland Journalof Medicine 300,1979,page1242)andWare's\ECMO" study(Statistical Science, 4,1989,page298) ²Somedesignsallowforearlystopping intheabsenceofa su±cienttreatmen te®ectasthetrialprogresses. These proceduresarereferredtoas\stochasticcurtailmen t"or \conditional power"calculations. 462 ²Designing agroupsequentialtrialforsurvivaldatare- quiressophisticated andhighlyspecialized software.EaSt, apackagefromCYTEL SOFTWAREthatdoesstan- dard(¯xed)survivaldesigns, aswellassequentialde- signs. ²Many\non-statistical" issuesenterdecisions aboutwhether ornottostopatrialearly ²P-valuesbasedonanalyses ofstudieswithsequential designsaredi±culttointerpret. ²Onceyoudo5interimanalyses, thenaddingmoremakes littledi®erence. Someclinicaltrialsgroups(HSPHAIDS group)havelargerandomized PhaseIIIstudiesmon- itoredatleastonceperyear(forsafetyreasons), and moststudieshave1-3interimlooks. ²Goingfroma¯xedtoagroupsequentialdesignadds onlyabout3-4%totherequired maximumsamplesize. Thisisagoodruleofthumbtouseincalculating the samplesizewhenyouplanondoinginterimmonitoring. 463 CompetingRisksandMultipleFailureTimes Sofar,we'vebeenactingasiftherewasonlyoneendpoint ofinterest,andthatcensoring duetodeath(orsomeother event)wasindependentoftheeventofinterest. However,inmanycontextsitislikelythatthetimetocen- soringissomehowcorrelated withthetimetotheeventof interest.Ingeneral, weoftenhaveseveraldi®erenttypes offailure(death,relapse,opportunistic infection, etc)which arerelated(i.e.,dependentor\competing"risks). Examples: ²Afterabonemarrowtransplan tation,patientsarefol- lowedtoevaluate\leukemia-fr eesurvival",sotheend- pointistimetoleukemiarelapseordeath.Thisendpoint conistsoftwotypesoffailures(competingrisks): {leukemiarelapse {non-relapse deaths ²Incardiovascularstudies,deathsfromothercauses(such ascancer)areconsidered competingrisks. ²Inactuarial analyses, welookattimetodeath,butwant toprovideseparate estimates ofhazardsforeachcause ofdeath(multipledecremen tlifetables). 464 Anotherexample: FortheMACstudy,theanalyses you havebeendoingoftimetoMACassumethatthecensoring timeisindependent. Recall: T=timetoeventofinterest(MAC) U=timetocensoring (death,losstoFU) X=min(T;U) ±=I(T·U) ObservableData:(X;±) Whatarethepossiblitieshere? ²(1)FailureTandcensoringUareindependent ²(2)FailureTandcensoringUaredependent 465 Case(1):Independentfailuretimes (thisincludes thecaseofindependentcensoring) BOTTOMLINE)NOPROBLEM Nonparametric estimation: Inthiscase,wecanusetheKaplan-Meier estimator toesti- mateST(t)=P(T>t). Parametricestimation: Ifweknowthejointdistribution of(T;U)hasacertain parametric form(exponential,Weibull,log-logistic), thenwe canusethelikelihoodfor(X;±)togetparameter estimates ofthemarginal distribution ofST(t). Semi-parametric estimation: WecanapplytheCoxregression modeltoassessthee®ects ofcovariatesonthemarginal hazard. 466 Case(2):Dependentfailuretimes BOTTOMLINE)BIGPROBLEM Tsiatis(1975)showedthatST(t)=P(T¸t)(i.e.,thesur- vivalfunction fortheeventTofinterest)cannotbe\identi- ¯ed"fromdataoftheform(X;±)foreachsubject. Infact,observing (X;±)doesnotprovideenoughinforma- tiontoestimate thejointdistribution of(T;U)sothatwe canevencheckwhether theassumption ofindependence is valid. Whenisitreasonabletoassumeindependentrisks? ²whencensoring occursbecausethestudyends,orbe- causethesubjectmovestoadi®erentstate ²andthereisnotrendovertimeinhealthstatusofen- rollingpatients InthecaseofourMACstudy,thefactthatsomeone dies mayre°ectthattheywouldhavebeenatgreaterriskofMAC iftheyhadnotdiedthansomeone elsewhoremained alive atthatpoint. Theassumption ofindependence meansthatthehazardfor someone whoiscensored attimetisexactlythesameasthat forsomeone withthesamecovariateswhoisalsoatriskat timet. 467 Whatistheimpactofdependentcompetingrisks? SludandByar(1988)showthatdependentcausesofdeath canpotentiallymakeriskfactorsappearprotectiv e: Ifwehave T=deathfromcauseofinterest andU=censoring, fromdeathduetoothercause andasinglebinarycovariateZ Z=8 >< >:1ifriskfactorispresent 0otherwise andwecalculate theKaplan-Meier survivalestimates ^S1(t) forZ=1and^S0(t)forZ=0assuming independentcen- soring,thenwecould(intheirhypothetical example) endup reversingthesignofthesurvivalfunctions: Trueorderingbetweensurvivaldistributions: S1(t)<S0(t)forallt KaplanMeierestimatesofsurvivaldistributions: ^S1(t)>^S0(t)forallt 468 Whatcanwedoifwesuspectdependentrisks? Alotofpeoplehavetriedtotacklethisproblem! References AlyEAA,KocharSC,andMcKeagueIW(1994).Sometestsfor comparingcumulativeincidencefunctionsandcause-speci¯chazard rates. JASA89,994-999. BenichouJandGailMH(1990).Estimatesofabsolutecausespeci¯c riskincohortstudies. Biometrics 46,813-826. (*)BoothALandSatchellSE(1995).ThehazardsofdoingaPhD:an analysisofcompletionandwithdrawalratesofBritishPhDstudents inthe1980's. JRSS-A,297-318. FarewellVT(1979).AnapplicationofCox'sproportionalhazard modeltomultipleinfectiondata.Applie dStatistics28,136-143. Gail,M(1982).Competingrisks. Encyclop ediaofStatistic alScienc es 2,75-81. LinDY,RobinsJM,andWeiLJ(1996).Comparingtwofailuretime distributionsinthepresenceofdependentcensoring. Biometrika 83,381-393. LunnMandMcNeilD(1995).ApplyingCoxregressiontocompeting risks. Biometrics 51,524-532 MoeschbergerMLandKleinJP(1988).Boundsonnetsurvivalprob- abilitiesfordependentcompetingrisks. Biometrics 44,529-538. (*)MoeschbergerMLandKleinJP(1995).Statisticalmethodsforde- pendentcompetingrisks. Lifetime DataAnalysis1,193-204. 469 PepeMS(1991).Inferenceforeventswithdependentrisksinmultiple endpointstudies. JASA86,770-778. (*)PepeMSandMoriM(1993).Kaplan-Meier, marginalorcondi- tionalprobabilitycurvesinsummarizingcompetingrisksfailure timedata? Statistics inMedicine12,737-751. PrenticeRL,Kalb°eischJD,PetersonAV,FlournoyN,FarewellVT, andBreslowNE(1978).Theanalysisoffailuretimesinthepresence ofcompetingrisks. Biometrics 34,541-554. SludEV,ByarDP,andSchatzkinA(1988).Dependentcompeting risksandthelatentfailuremodel.Biometrics 44,1203-1205. SludEandByarD(1988).Howdependentcausesofdeathcanmake riskfactorsappearprotective.Biometrics 44,265-269. Tsiatis,A.(1975).Anonidenti¯abilityaspectoftheproblemofcom- petingrisks. ProceedingsoftheNational Academy ofScienc es72, 20-22. 470 Therehasbeenalivelydebateintheliterature aboutthebestwaytoattackthisproblem.The twosidesarebasicallydividedaboutwhichtype ofmodeltouse: ²basedoncause-speci¯chazard functions (observ- ables) ²basedonlatentvariable models(unobserv ables) The¯rstapproachfocusesonwhattheobservedsurvivalis duetoacertaincauseoffailure,acknowledingthatthereare othertypesoffailuresoperatingatthesametime. Thesecondapproachattempts toestimate whatthesurvival associatedwithacertainfailuretypewouldhavebeen,ifthe othertypesoffailureshadbeenremoved. 471 GeneralCaseofMultipleFailureTypes Ingeneral, saywehavemdi®erenttypesoffailure(say, causesofdeath),andtherespectivetimestofailureare: T1;T2;T3;¢¢¢;Tm andweobserveT=min(T1;T2;:::;Tm) Wecanwritethecause-speci¯chazardfunction forthej-th failuretypeas: ¸j(t)=lim ¢t!01 ¢tPr(t·T<t+¢t;J=jjT¸t) Theoverallhazardofdeathisthesumoverthefailuretypes: ¸(t)=mX j=1¸j(t) where¸(t)=lim ¢t!01 ¢tPr(t·T<t+¢tjT¸t) Q.Canweestimatethesequantities?...evenif therisksaredependent? A.Yes,Prentice(1978)showsthatprobabilities thatcanbeexpressedasafunctionofthecause- speci¯chazardscanbeestimated. 472 Forexample, estimable quantitiesinclude: (a)Theoverallsurvivalprobability3: ST(t)=P(T¸t)=exp" ¡Zt 0¸(u)du# =exp2 64¡Zt 0X j¸j(u)du3 75 (b)Conditionalprobabilityoffailingfromcause jintheinterval(¿i¡1;¿i] Q(i;j)=[ST(¿i¡1)]¡1Z¿ ¿¡1¸j(u)ST(u)du (c)Conditionalprobabilityofsurviving ithinter- val ½i=1¡mX j=1Q(i;j) 3Note:previously Isaidyoucouldn'testimateST(t),butthatwaswhenTwas thetimetoeventofinterest(possiblyunobservable)andUwasthepossiblycor- relatedcensoring time.Here,ST(t)isthesurvivaldistribution fortheminimum ofallfailures,whichcanalwaysbeobserved 473 Estimators: (a)TheMLEofQ(i;j)issimply ^Q(i;j)=dij ri i.e.,thenumberoffailures(deaths) duetocausejduring thei-thintervalamongtherisubjectsatriskoffailure atthebeginning oftheinterval. (b)TheMLEof½iis:: ^½i=ri¡Pm j=1dij ri=1¡Pm j=1dij ri (c)TheMLEofST(t)isbasedon½i: ^ST(¿i)=iY k=1½k 474 Sowhatcan'tweestimate? Compare thecause-speci¯chazardfunction : ¸j(t)=lim ¢t!01 ¢tPr(t·T<t+¢t;J=jjT¸t) withthemarginalhazardfunction : ¸j(t)=lim ¢t!01 ¢tPr(t·Tj<t+¢tjTj¸t) Wecangetestimates ofthecause-speci¯chazardfunction, sincewecanestimateST(t)=P(T¸t)evenifthefail- uretimesaredependent.(Inotherwords,wecanobserve whether eachpatientisstillaliveornot) Butunfortunately ,wecan'testimate themarginal hazard function whentherisksaredependent,sincewecan'testi- mateSj(t)=P(Tj¸t).(wecan'ttellwhentheywould havehadeventTjiftheyhaveadi®erentevent¯rst) Thisisthemaintrickyissueofcompetingrisksanalyses. 475 Backtooriginalquestion... Whatcanwedoifwesuspectdependentrisks? Ex.Saywehavetwotypesoffailures,T1andT2,andwe thinktheyaredependent.However,weareinterestedinthe ¯rsttypeoffailureT1,andviewthecompetingriskT2as censoring (likeinbonemarrowtransplan texample). Methodsofsummarizing datawithcompetingrisks: (1)Summarize thecause-speci¯chazard rateovertime (2)UsetheKaplan-Meier estimate anyway,^ST1(t) (3)ReportthecomplementtotheKM,1¡^ST1(t) (4)Usecumulativeincidence curves(crudeincidence curve) (5)Usetheconditionalprobabilityfunction (6)Giveupperandlowerboundsforthetruemarginal survivalfunction, intheabsenceofthecompetingrisk PepeandMorireviewthe¯rst5oftheseoptions, andrec- ommendagainstoption(2),butnotethatthisisoftenwhat peopleendupdoing. 476 Tomaketheexamplemoreconcrete: Sayweareinterested intimetoMACordeath,whichever occurs¯rst.Wede¯ne: T=timetoMACordeath andU=censoring (assumed independent) andthetypeoffailureisdenoted byj j=8 >< >:MifeventisMAC Difeventisdeathfromothercauses Inthealternativ e\latentvariable"framework,wewouldde- ¯ne TM=timetoMAC andTD=timetoDeath although wemightnotbeabletoobserveTMifTDoccurred ¯rst. 477 Methodsforcompetingrisks: (1)Summarizingthecause-speci¯chazardover time Asmentionedabove,thisisoneofthequantitiesthatwe canestimate. Ourbasicestimator duringtimeintervaliis ^Q(i;j)=dij ri. Sowecanplot^¸j(t)overtime,andgetsomeinsightasto biological phenomona involved. However,ifyouremembersomeoftheplotsIshowedyouof hazardsovertime,theytendedtobehighlyvariable.Several contributions havebeenmadetowards\smoothing"outthe inherentvariabilityintheestimates ofcause-speci¯chazards. ²Efron(1988) ²Ramlau-Hansen (1983) ²TannerandWong(1983) Drawback:thehazardfunctions alonedonotgiveoverall e®ectofacovariateonsurvival. Example: Ifthehazardfunctions for¸j(t)fortwotreat- mentscross,wecan'tsaywhichtreatmen tleadstolower overalleventrate. 478 FigurefromPepeandMoriforLeukemiaData KernelEstimatesofCause-Speci¯cHazards 479 Methodsforcompetingrisks: (2)ApplyingKaplan-Meier tocause-speci¯c¸j's SayweevaluatetheMACsurvivaldistribution bytreating ²allMACcasesas\events" ²anydeathswithoutMACas\censorings" andconstruct theKaplan-Meier survivalcurve. Whatareweestimating? S¤ M(t)=exp" ¡Zt 0¸M(u)du# where¸M(t)isthecause-speci¯chazardforMAC: ¸M(t)=lim ¢t!01 ¢tPr(t·T<t+¢t;j=MjT¸t) i.e.,theconditional probabilit ythatMACoccursinashort periodoftime,giventhatthesubjectisaliveandMAC-free. Theinterpretation oftheKaplan-Meier curveisasthe \exponentialofthenegativecumulativecause- speci¯chazard". Clinicians (andothers!) havedi±cultyunderstanding this function, sinceithasnodirectclinicalinterpretation. 480 FigurefromPepeandMoriforLeukemiaData KaplanMeierwithCauseSpeci¯cHazards Theythoughtthiswassuchabadidea,thatthey didnotincludeanyplotofthis! 481 Methodsforcompetingrisks: (3)UsingthecomplementoftheKaplan-Meier Another function usedfairlyofteninthecompetingrisks areaissometimes referred toasthepureprobability function : 1¡S¤ j(t)=1¡exp" ¡Zt 0¸j(u)du# InourMACexample, 1¡S¤ M(t)couldbeinterprested asthe predictiv eprobabilit yofMACbytimetiftheriskofdeath couldberemoved. Ifweweredesigning anewstudyforamiracledrugthat seemedsopowerfulthatitwouldnotonlyreduceMACbut preventalldeathfromothercausesinHIV-infected patients, thenwecoulduseestimates 1¡^S¤ M(t)tohelpdesignour newstudy. Thiswouldbeprettyoptimistic, andtherehasbeenalotof workonthestrict(anduntestable )assumptions required to interprettheKMcurveinthismanner. PepeandMoricontendthatthisfunction isirrelevantfor summarizing datafromacompetingrisksstudy. 482 FigurefromPepeandMoriforLeukemiaData ComplementKaplan-Meier Functions 483 Methodsforcompetingrisks: (4)UsingCumulativeIncidenceCurves Thishasalsobeentermedthe\crudeincidence curve",and estimates themarginalprobabilityofaneventinthe settingwhereothercompetingrisksareacknowledged toex- ist. DescriptionofMethod:themethodisdescribedin moredetailinKalb°eisc handPrentice(p.169). Testsforcovariates: Testsforcomparing cumulative incidence curvesamongtreatmen tgroups(orsomeotherco- variate)havebeendevelopedbyBobGray(1988). They aresimilartologranktestsinthattheyare\linearrank" statistics. Ifwewereabletofollowupallsubjectstotimet,thenthe cumulativeincidence curveswouldrelfectwhatproportion ofthetotalstudypopulation havehadtheparticular event (i.e.,MAC)bytimet. 484 FigurefromPepeandMoriforLeukemiaData CumulativeIncidenceCurves 485 Methodsforcompetingrisks: (5)ConditionalProbabilityCurves Thishasthesame°avorasthecomplemen tKM,butamore naturalinterpretation. PepeandMoride¯netheconditional probabilit yfunction as: dCPM(t)=P(TM·tjTD¸t) =^PM(t) 1¡^PD(t) where^PM(t)=Zt 0^ST(u)dNM(u) Y(u) and^PD(t)=Zt 0^ST(u)dND(u) Y(u) Intheabove,^ST(u)istheKMestimate oftheoverallmac- freesurvivaldistribution, andthetermsNM(u),ND(u),and Y(u)re°ectthe\countingprocess"forthenumberofsub- jectswithMAC,death,andatriskattimeu,respectively. Intheabsenceofcensoring, theinterpretation isthepropor- tionofpatientswhodevelopMACamongthosewhodonot dieofothercauses. Tests:PepeandMorialsopresenttests,whicharesums overtimeofweighteddi®erences betweentheconditional (ormarginal) probabilities fortwogroups. 486 FigurefromPepeandMoriforLeukemiaData ConditionalProbabilityandMarginalCurves 487 Methodsforcompetingrisks: (6)BoundsonNetSurvivalCurves Asnotedpreviously ,wecannotestimateSj(t)=P(Tj¸t) ifthefailuretimesaredependent(eg,wecan'testimate the survivalfunction forMACiftimetoMACiscorrelated with timetodeathwithoutMAC). However,wemaybeabletosaysomething abouttherange ofSj(t)by¯ndingupperandlowerboundsthatcontain Sj(t). ²Peterson (1976)obtained generalboundsbasedonthe minimal andmaximal dependencestructure for(TM;TD). Theboundsallowanypossibledependence structure, butcanbeverywide. ²SludandRubenstein (1983)obtained tighterbounds onSj(t)byusingadditional information, butrequirethe usertospecifyreasonable boundsonafunction½.Once ½issupplied, themarginal distribution Sj;½(t)canbe obtained. ²KleinandMoeshberger(1988)usetheframework ofClaytonandOakesforbivariatesurvivaltoobtain tighterboundsthanthoseofPeterson. Again,theuser hastosupplyboundsonafunctionµ,andoncethisis given^Sj;µ(t)canbeobtained. 488 FigurefromKleinandMoeschberger(1988) BoundsonNetSurvivalCurves 489 Onelastexample:PromotionofFacultyatHSPH Iwasaskedtoanalyzetheschool'sdatafrom1980-1995 on promotion ofFacultyfromAssistantProfessor toAssociate Professor toassesswhether thereweredi®erences between malesandfemalesandamongacademic areas(social,labo- ratory,quantitative). Problem: Wouldyouthinkthat\censoring" (someone leavingtheir tenuretrackpositionpriortogettingpromoted) isindepen- dentoftheprobabilit yofpromotion? Iconsidered 3approachesforaccountingforcensoring: MethodI:assumes thosewhodeparted wouldNOThave beenpromoted MethodII:assumes thosewhodeparted wouldhavebeen promoted atthesamerateasthosewhostayed MethodIII:assumes 50%ofthosewhodeparted wouldnot havebeenpromoted, andtheother50%would havebeenpromoted atthesamerateasthose whostayed WhichMethodcorrespondsto\non-informativ e" (independent)censoring? 490 Results:Cumulativeprobabilitiesofpromotion E®ectofGenderonPromotion OverallMalesFemales MethodI:0.6310.719 0.451 MethodII:0.9331.000 0.674 MethodIII:0.7360.825 0.531 E®ectofAcademicAreaonPromotion OverallQuantitativeSocialLaboratory MethodI:0.631 0.703 0.238 0.701 MethodII:0.933 0.950 0.389 1.000 MethodIII:0.736 0.803 0.287 0.801 491