survival analysis_Ibrahim_J
PDF · 491 pages · 2.3 MB
Open PDF file
Lecture notes for a survival analysis course, apparently by or associated with Ibrahim (per the file name), kept in the Probability and Statistics folder. The opening sections cover time-to-event data, right, left and interval censoring, independent versus informative censoring, and Type I/II/III censoring. They also present example clinical datasets (nursing home stays, MAC prevention trial, UMARU drug study, halibut survival) and define the density, survivor, hazard and cumulative hazard functions. Only the first part of a long document was read.
AI-written summary; may contain errors. This description is approximate.
Extracted text (machine-read; may contain errors)
SurvivalAnalysis: Introduction
SurvivalAnalysis typicallyfocusesontimetoeventdata.
Inthemostgeneralsense,itconsistsoftechniquesforpositive-
valuedrandomvariables, suchas
²timetodeath
²timetoonset(orrelapse)ofadisease
²lengthofstayinahospital
²duration ofastrike
²moneypaidbyhealthinsurance
²viralloadmeasuremen ts
²timeto¯nishing adoctoraldissertation!
Kindsofsurvivalstudiesinclude:
²clinicaltrials
²prospectivecohortstudies
²retrospectivecohortstudies
²retrospectivecorrelativ estudies
Typically,survivaldataarenotfullyobserved,butrather
arecensored.
1
Inthiscourse,wewill:
²describesurvivaldata
²comparesurvivalofseveralgroups
²explainsurvivalwithcovariates
²designstudieswithsurvivalendpoints
Someknowledgeofdiscrete datamethodswillbeuseful,
sinceanalysis ofthe\timetoevent"usesinformation from
thediscrete(i.e.,binary)outcome ofwhether theeventoc-
curredornot.
Someusefulreferences:
²Collett:ModellingSurvivalDatainMedicalResearch
²CoxandOakes:AnalysisofSurvivalData
²Kalb°eisc handPrentice:TheStatisticalAnalysisof
FailureTimeData
²Lee:StatisticalMethodsforSurvivalDataAnalysis
²Fleming &Harrington: Counting ProcessesandSur-
vivalAnalysis
²Hosmer&Lemesho w:AppliedSurvivalAnalysis
²Kleinbaum:SurvivalAnalysis:Aself-learningtext
2
²Klein&Moeschberger:SurvivalAnalysis:Techniques
forcensoredandtruncateddata
²Cantor:Extending SASSurvivalAnalysisTechniques
forMedicalResearch
²Allison:SurvivalAnalysisUsingtheSASSystem
²Jennison &Turnbull:GroupSequentialMethodswith
ApplicationstoClinicalTrials
²Ibrahim, Chen,&Sinha:Bayesian SurvivalAnalysis
3
SomeDe¯nitionsandnotation
Failuretimerandomvariables arealwaysnon-negative.
Thatis,ifwedenotethefailuretimebyT,thenT¸0.
Tcaneitherbediscrete (takinga¯nitesetofvalues,e.g.
a1;a2;:::;an)orcontinuous(de¯ned on(0;1)).
ArandomvariableXiscalledacensoredfailuretime
randomvariable ifX=min(T;U),whereUisanon-
negativecensoring variable.
Inordertode¯neafailuretimerandomvariable,
weneed:
(1)anunambiguoustimeorigin
(e.g.randomization toclinicaltrial,purchaseofcar)
(2)atimescale
(e.g.realtime(days,years),mileageofacar)
(3)de¯nition oftheevent
(e.g.death,needanewcartransmission)
4
Illustrationofsurvivaldata
X
X
y
y
X
y
X
y
study
opensstudy
closes
y=censored observation
X=event
5
Theillustration ofsurvivaldataontheprevious pageshows
severalfeatures whicharetypicallyencounteredinanalysis
ofsurvivaldata:
²individuals donotallenterthestudyatthesametime
²whenthestudyends,someindividuals stillhaven'thad
theeventyet
²otherindividuals dropoutorgetlostinthemiddleof
thestudy,andallweknowaboutthemisthelasttime
theywerestill\free"oftheevent
The¯rstfeatureisreferredtoas\staggeredentry"
Thelasttwofeatures relateto\censoring" ofthefailure
timeevents.
6
Typesofcensoring:
²Right-censoring :
onlyther.v.Xi=min(Ti;Ui)isobserveddueto
{losstofollow-up
{drop-out
{studytermination
Wecallthisright-censoring becausethetrueunobserv ed
eventistotherightofourcensoring time;i.e.,allwe
knowisthattheeventhasnothappenedattheendof
follow-up.
Inaddition toobservingXi,wealsogettoseethefail-
ureindicator :
±i=8
><
>:1ifTi·Ui
0ifTi>Ui
Somesoftwarepackagesinsteadassumewehavea
censoringindicator :
ci=8
><
>:0ifTi·Ui
1ifTi>Ui
Right-censoring isthemostcommon typeofcensoring
assumption wewilldealwithinsurvivalanalysis.
7
²Left-censoring
CanonlyobserveYi=max(Ti;Ui)andthefailureindi-
cators:
±i=8
><
>:1ifUi·Ti
0ifUi>Ti
e.g.(Miller)studyofageatwhichAfricanchildrenlearn
atask.Somealreadyknew(left-censored), somelearned
duringstudy(exact),somehadnotyetlearnedbyend
ofstudy(right-censored).
²Interval-censoring
Observe(Li;Ri)whereTi2(Li;Ri)
Ex.1:Timetoprostate cancer,observelongitudinal
PSAmeasuremen ts
Ex.2:Timetoundetectable viralloadinAIDSstudies,
basedonmeasuremen tsofviralloadtakenateachclinic
visit
Ex.3:Detectrecurrence ofcoloncanceraftersurgery.
Followpatientsevery3monthsafterresection ofprimary
tumor.
8
Independentvsinformativecensoring
²Wesaycensoring isindependent(non-informativ e)if
UiisindependentofTi.
{Ex.1IfUiistheplanned endofthestudy(say,2
yearsafterthestudyopens),thenitisusuallyinde-
pendentoftheeventtimes.
{Ex.2IfUiisthetimethatapatientdropsout
ofthestudybecausehe/shegotmuchsickerand/or
hadtodiscontinuetakingthestudytreatmen t,then
UiandTiareprobably notindependent.
AnindividualcensoredatUshouldberepre-
sentativeofallsubjectswhosurvivetoU.
Thismeansthatcensoring atUcoulddependonprog-
nosticcharacteristics measured atbaseline, butthatamong
allthosewiththesamebaselinecharacteristics, theprob-
abilityofcensoring priortoorattimeUshouldbethe
same.
²Censoring isconsideredinformativeifthedistribu-
tionofUicontainsanyinformation abouttheparameters
characterizing thedistribution ofTi.
9
Supposewehaveasampleofobservationsonnpeople:
(T1;U1);(T2;U2);:::;(Tn;Un)
Therearethreemaintypesof(right)censoring times:
²TypeI:AlltheUi'sarethesame
e.g.animalstudies,allanimalssacri¯ced after2years
²TypeII:Ui=T(r),thetimeoftherthfailure.
e.g.animalstudies,stopwhen4/6havetumors
²TypeIII:theUi'sarerandomvariables,±i'sarefailure
indicators:
±i=8
><
>:1ifTi·Ui
0ifTi>Ui
TypeIandTypeIIarecalledsinglycensored data,
TypeIIIiscalledrandomly censored (orsometimes pro-
gressively censored).
10
Someexampledatasets:
ExampleA.Durationofnursinghomestay
(Morrisetal.,CaseStudiesinBiometry ,Ch12)
TheNational CenterforHealthServices Researchstudied
36for-pro¯t nursinghomestoassessthee®ectsofdi®erent
¯nancial incentivesonlengthofstay.\Treated"nursing
homesreceivedhigherperdiemsforMedicaid patients,and
bonusesforimprovingapatient'shealthandsendingthem
home.
Studyincluded 1601patientsadmitted betweenMay1,1981
andApril30,1982.
Variablesinclude:
LOS-Lengthofstayofaresident(indays)
AGE-Ageofaresident
RX-Nursinghomeassignmen t(1:bonuses,0:nobonuses)
GENDER -Gender(1:male, 0:female)
MARRIED -(1:married, 0:notmarried)
HEALTH-healthstatus(2:second best,5:worst)
CENSOR -Censoring indicator (1:censored, 0:discharged)
Firstfewlinesofdata:
378610020
617710040
11
ExampleB.Fecundability
Womenwhohadrecentlygivenbirthwereaskedtorecall
howlongittookthemtobecomepregnant,andwhether or
nottheysmokedduringthattime.Theoutcome ofinter-
est(summarized below)istimetopregnancy (measured in
menstrual cycles).
19subjectswerenotabletogetpregnantafter12months.
CycleSmokersNon-smok ers
129 198
216 107
317 55
4 4 38
5 3 18
6 9 22
7 4 7
8 5 9
9 1 5
10 1 3
11 1 6
12 3 6
12+ 7 12
12
ExampleC:MACPreventionClinicalTrial
ACTG196wasarandomized clinicaltrialtostudythee®ects
ofcombination regimens onpreventionofMAC(mycobac-
teriumaviumcomplex),oneofthemostcommon oppor-
tunisticinfections inAIDSpatients.
Thetreatmentregimens were:
²clarithrom ycin(new)
²rifabutin (standard)
²clarithrom ycinplusrifabutin
Othercharacteristics oftrial:
²PatientsenrolledbetweenApril1993andFebruary1994
²Follow-upendedAugust1995
²InFebruary1994,rifabutin dosagewasreduced from3
pills/day(450mg) to2pills/day(300mg) duetoconcern
overuveitis1
Themainintent-to-treat analysiscompared the3treatmen t
armswithoutadjusting forthischangeindosage.
1Uveitisisanadverseexperienceresultinginin°ammationofthe
uvealtractintheeyes(about3-4%ofpatientsreporteduveitis).
13
ExampleD:HMOStudyofHIV-relatedSurvival
Thisishypothetical datausedbyHosmer&Lemesho w(de-
scribedonpages2-17)containing100observationsonHIV+
subjectsbelonging toanHealthMaintenanceOrganization
(HMO). TheHMOwantstoevaluatethesurvivaltimeof
thesesubjects.Inthishypothetical dataset, subjectswere
enrolled fromJanuary1,1989untilDecember31,1991.
StudyfollowupthenendedonDecember31,1995.
Variables:
ID SubjectID(1-100)
TIME Survivaltimeinmonths
ENTDATEEntrydate
ENDDATEDatefollow-upendedduetodeathorcensoring
CENSOR DeathIndicator (1=death, 0=censor)
AGE Ageofsubjectinyears
DRUG HistoryofIVDrugUse(0=no,1=y es)
ThisdatasetisusedbyHosmer &Lemesho wtomotivate
someconcepts insurvivalanalysisinChap.1oftheirbook.
14
ExampleE:UMARUImpactStudy(UIS)
ThisdatasetcomesfromtheUniversityofMassachusetts
AIDSResearchUnit(UMARU)IMPACTStudy,a5-year
collaborativeresearchprojectcomprised oftwoconcurren t
randomized trialsofresidentialtreatmen tfordrugabuse.
(1)ProgramA:Randomized 444subjectstoa3-or6-
monthprogram ofhealtheducation andrelapsepreven-
tion.Clientsweretaughttorecognize \high-risk" situ-
ationsthataretriggerstorelapse,andtaughtskillsto
copewiththesesituations withoutusingdrugs.
(2)ProgramB:Randomized 184participan tstoa6-or
12-monthprogram withhighlystructured life-styleina
communallivingsetting.
Variables:
ID SubjectID(1-628)
AGEAgeinyears
BECKTOTABeckDepressionScore
HERCOCHeroinorCocaineUsepriortoentry
IVHX IVDruguseatAdmission
NDRUGTXNumberpreviousdrugtreatments
RACESubject'sRace(0=White,1=Other)
TREATTreatmentAssignment(0=short,1=long)
SITE TreatmentProgram(0=A,1=B)
LOT LengthofTreatment(days)
TIME TimetoReturntoDrugUse(days)
CENSOR IndicatorofDrugUseRelapse(1=yes,0=censored)
15
ExampleF:AtlanticHalibutSurvivalTimes
Oneconservationmeasure suggested fortrawl¯shingisa
minimumsizelimitforhalibut(32inches).However,thissize
limitwouldonlybee®ectiveifcaptured ¯shbelowthelimit
surviveduntilthetimeoftheirrelease.Anexperimentwas
conducted toevaluatethesurvivalratesofhalibutcaughtby
trawlsorlonglines, andtoassessotherfactorswhichmight
contributetosurvival(duration oftrawling,maximumdepth
¯shed,sizeof¯sh,andhandling time).
AnarticlebySmith,WaiwoodandNeilson,SurvivalAnaly-
sisforSizeRegulationofAtlanticHalibutinCaseStudies
inBiometry compares parametric survivalmodelstosemi-
parametric survivalmodelsinevaluating thisdata.
Survival TowDi® Length HandlingTotal
Obs Time CensoringDurationinofFishTime log(catch)
#(min) Indicator (min.) Depth(cm) (min.) ln(weight)
100353.0 1 301539 5 5.685
109111.0 1 100 544 29 8.690
11364.0 0 100 1053 4 5.323
116500.0 1 100 1044 4 5.323
....
16
MoreDe¯nitionsandNotation
Thereareseveralequivalentwaystocharacterize theprob-
abilitydistribution ofasurvivalrandomvariable. Someof
thesearefamiliar; othersarespecialtosurvivalanalysis. We
willfocusonthefollowingterms:
²Thedensityfunctionf(t)
²ThesurvivorfunctionS(t)
²Thehazardfunction¸(t)
²Thecumulativehazardfunction ¤(t)
²Densityfunction(orProbabilityMassFunc-
tion)fordiscreter.v.'s
SupposethatTtakesvaluesina1;a2;:::;an.
f(t)=Pr(T=t)
=8
>><
>>:fjift=aj;j=1;2;:::;n
0ift6=aj;j=1;2;:::;n
²DensityFunctionforcontinuousr.v.'s
f(t)=lim
¢t!01
¢tPr(t·T·t+¢t)
17
²SurvivorshipFunction :S(t)=P(T¸t).
Inothersettings, thecumulativedistribution function,
F(t)=P(T·t),isofinterest.Insurvivalanalysis, our
interesttendstofocusonthesurvivalfunction,S(t).
Foracontinuousrandomvariable:
S(t)=Z1
tf(u)du
Foradiscreterandomvariable:
S(t)=X
u¸tf(u)
=X
aj¸tf(aj)
=X
aj¸tfj
Notes:
²Fromthede¯nition ofS(t)foracontinuousvariable,
S(t)=1¡F(t)aslongasF(t)isabsolutely continuous
w.r.ttheLebesguemeasure. [Thatis,F(t)hasadensity
function.]
²Foradiscretevariable,wehavetodecidewhattodoif
aneventoccursexactlyattimet;i.e.,doesthatbecome
partofF(t)orS(t)?
²Togetaroundthisproblem, severalbooksde¯ne
S(t)=Pr(T>t),orelsede¯neF(t)=Pr(T<t)
(eg.Collett)
18
²HazardFunction¸(t)
Sometimes calledaninstantane ousfailurerate,the
forceofmortality ,ortheage-speci¯cfailurerate.
{Continuousrandomvariables:
¸(t)=lim
¢t!01
¢tPr(t·T<t+¢tjT¸t)
=lim
¢t!01
¢tPr([t·T<t+¢t]T[T¸t])
Pr(T¸t)
=lim
¢t!01
¢tPr(t·T<t+¢t)
Pr(T¸t)
=f(t)
S(t)
{Discreterandomvariables:
¸(aj)´¸j=Pr(T=ajjT¸aj)
=P(T=aj)
P(T¸aj)
=f(aj)
S(aj)
=f(t)
P
k:ak¸ajf(ak)
19
²CumulativeHazardFunction ¤(t)
{Continuousrandomvariables:
¤(t)=Zt
0¸(u)du
{Discreterandomvariables:
¤(t)=X
k:ak<t¸k
20
RelationshipbetweenS(t)and¸(t)
We'vealreadyshownthat,foracontinuousr.v.
¸(t)=f(t)
S(t)
Foraleft-continuoussurvivorfunctionS(t),wecanshow:
f(t)=¡S0(t)orS0(t)=¡f(t)
Wecanusethisrelationship toshowthat:
¡d
dt[logS(t)]=¡0
B@1
S(t)1
CAS0(t)
=¡¡f(t)
S(t)
=f(t)
S(t)
Soanotherwaytowrite¸(t)isasfollows:
¸(t)=¡d
dt[logS(t)]
21
RelationshipbetweenS(t)and¤(t):
²Continuouscase:
¤(t)=Zt
0¸(u)du
=Zt
0f(u)
S(u)du
=Zt
0¡d
dulogS(u)du
=¡logS(t)+logS(0)
)S(t)=e¡¤(t)
²Discretecase:
Supposethataj<t·aj+1.Then
S(t)=P(T¸a1;T¸a2;:::;T¸aj+1)
=P(T¸a1)P(T¸a2jT¸a1)¢¢¢P(T¸aj+1jT¸aj)
=(1¡¸1)£¢¢¢£(1¡¸j)
=Y
k:ak<t(1¡¸k)
Coxde¯nes¤(t)=P
k:ak<tlog(1¡¸k)sothatS(t)=
e¡¤(t)inthediscretecase,aswell.
22
MeasuringCentralTendencyinSurvival
²Meansurvival-callthis¹
¹=Z1
0uf(u)duforcontinuousT
=nX
j=1ajfjfordiscreteT
²Mediansurvival-callthis¿,isde¯nedby
S(¿)=0:5
Similarly ,anyotherpercentilecouldbede¯ned.
Inpractice, wedon'tusuallyhitthemediansurvival
atexactlyoneofthefailuretimes.Inthiscase,the
estimated mediansurvivalisthesmallesttime¿such
that
^S(¿)·0:5
23
Somehazardshapesseeninapplications:
²increasing
e.g.agingafter65
²decreasing
e.g.survivalaftersurgery
²bathtub
e.g.age-speci¯cmortality
²constant
e.g.survivalofpatientswithadvancedchronicdisease
24
Estimatingthesurvivalorhazardfunction
Wecanestimate thesurvival(orhazard) function intwo
ways:
²byspecifying aparametric modelfor¸(t)basedona
particular densityfunctionf(t)
²bydevelopinganempirical estimate ofthesurvivalfunc-
tion(i.e.,non-parametric estimation)
Ifnocensoring:
Theempirical estimate ofthesurvivalfunction, ~S(t),isthe
proportionofindividuals witheventtimesgreaterthant.
Withcensoring:
Iftherearecensored observations,then~S(t)isnotagood
estimate ofthetrueS(t),soothernon-parametric methods
mustbeusedtoaccountforcensoring (life-table methods,
Kaplan-Meier estimator)
25
SomeParametricSurvivalDistributions
²TheExponentialdistribution (1parameter)
f(t)=¸e¡¸tfort¸0
S(t)=Z1
tf(u)du
=e¡¸t
¸(t)=f(t)
S(t)
=¸constanthazard!
¤(t)=Zt
0¸(u)du
=Zt
0¸du
=¸t
Check:DoesS(t)=e¡¤(t)?
median: solve0:5=S(¿)=e¡¸¿:
)¿=¡log(0:5)
¸
mean:
Z1
0u¸e¡¸udu=1
¸
26
²TheWeibulldistribution (2parameters)
Generalizes exponential:
S(t)=e¡¸t·
f(t)=¡d
dtS(t)=·¸t·¡1e¡¸t·
¸(t)=·¸t·¡1
¤(t)=Zt
0¸(u)du=¸t·
¸-thescaleparameter
·-theshapeparameter
TheWeibulldistribution isconvenientbecauseofitssim-
pleform.Itincludes severalhazardshapes:
·=1!constanthazard
0<·<1!decreasing hazard
·>1!increasing hazard
27
²Rayleighdistribution
Another 2-parameter generalization ofexponential:
¸(t)=¸0+¸1t
²compoundexponential
T»exp(¸);¸»g
f(t)=Z1
0¸e¡¸tg(¸)d¸
²log-normal ,log-logistic :
Possibledistributions forTobtained byspecifying for
logTanyconvenientfamilyofdistributions, e.g.
logT»normal(non-monotone hazard)
logT»logistic
28
Whyuseoneversusanother?
²technicalconvenienceforestimation andinference
²explicitsimpleformsforf(t);S(t),and¸(t).
²qualitativ eshapeofhazardfunction
Onecanusuallydistinguish betweenaone-parameter model
(liketheexponential)andtwo-parameter (likeWeibullor
log-normal) intermsoftheadequacy of¯ttoadataset.
Without alotofdata,itmaybehardtodistinguish between
the¯tsofvarious2-parameter models(i.e.,Weibullvslog-
normal)
29
PlotsofestimatesofS(t)
BasedonKM,exponential,Weibull,andlog-normal
forstudyofprotease inhibitors inAIDSpatients
(ACTG320)
KM Curves for Time to PCP
- 2 Drug Arm
daysProbability
0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00
1KM
Exponential
Weibull
LognormalKM Curves for Time to PCP
- 3 Drug Arm
daysProbability
0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00
1KM
Exponential
Weibull
Lognormal
30
PlotsofestimatesofS(t)
BasedonKM,exponential,Weibull,andlog-normal
forstudyofprotease inhibitors inAIDSpatients
(ACTG320)
KM Curves for Time to MAC
- 2 Drug Arm
daysProbability
0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00
1KM
Exponential
Weibull
LognormalKM Curves for Time to MAC
- 3 Drug Arm
daysProbability
0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00
1KM
Exponential
Weibull
Lognormal
31
PlotsofestimatesofS(t)
BasedonKM,exponential,Weibull,andlog-normal
forstudyofprotease inhibitors inAIDSpatients
(ACTG320)
KM Curves for Time to CMV
- 2 Drug Arm
daysProbability
0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00
1KM
Exponential
Weibull
LognormalKM Curves for Time to CMV
- 3 Drug Arm
daysProbability
0 100 200 300 4000.90 0.92 0.94 0.96 0.98 1.00
1KM
Exponential
Weibull
Lognormal
32
PreviewofComingAttractions
Nextwewilldiscussthemostfamousnon-parametric ap-
proachforestimating thesurvivaldistribution, calledthe
Kaplan-Meier estimator .
Tomotivatethederivationofthisestimator, wewill¯rst
consider asetofsurvivaltimeswherethereisnocensoring.
Thefollowingaretimestorelapse (weeks)for21leukemia
patientsreceiving controltreatmen t(Table1.1ofCox&
Oakes):
1,1,2,2,3,4,4,5,5,8,8,8,8,11,11,12,12,15,17,22,23
Howwouldweestimate S(10),theprobabilit ythatanindi-
vidualsurvivestotime10orlater?
Whatabout~S(8)? Isit12
21or8
21?
33
Let'sconstruct atableof~S(t):
Valuesoft^S(t)
t·121/21=1.000
1<t·219/21=0.905
2<t·317/21=0.809
3<t·4
4<t·5
5<t·8
8<t·11
11<t·12
12<t·15
15<t·17
17<t·22
22<t·23
EmpiricalSurvivalFunction:
Whenthereisnocensoring, thegeneralformulais:
~S(t)=#individualswithT¸t
totalsamplesize
34
Inmostsoftwarepackages,thesurvivalfunction isevaluated
justaftertimet,i.e.,att+.Inthiscase,weonlycountthe
individuals withT>t.
Example forleukemiadata(controlarm):
35
StataCommandsforSurvivalEstimation
.useleukem
.stsetremissstatusiftrt==0 (tokeeponlyuntreated patients)
(21observations deleted)
.stslist
failure _d:status
analysis time_t:remiss
Beg. NetSurvivor Std.
Time Total FailLostFunction Error [95%Conf.Int.]
----------------------------------------------------------- -----------
1 21 200.9048 0.0641 0.6700 0.9753
2 19 200.8095 0.0857 0.5689 0.9239
3 17 100.7619 0.0929 0.5194 0.8933
4 16 200.6667 0.1029 0.4254 0.8250
5 14 200.5714 0.1080 0.3380 0.7492
8 12 400.3810 0.1060 0.1831 0.5778
11 8200.2857 0.0986 0.1166 0.4818
12 6200.1905 0.0857 0.0595 0.3774
15 4100.1429 0.0764 0.0357 0.3212
17 3100.0952 0.0641 0.0163 0.2612
22 2100.0476 0.0465 0.0033 0.1970
23 1100.0000 . . .
----------------------------------------------------------- -----------
.stsgraph
36
SASCommandsforSurvivalEstimation
dataleuk;
inputt;
cards;
1
1
2
2
3
4
4
5
5
8
8
8
8
11
11
12
12
15
17
22
23
;
proclifetest data=leuk;
timet;
run;
37
SASOutputforSurvivalEstimation
TheLIFETEST Procedure
Product-Limit Survival Estimates
Survival
Standard Number Number
tSurvival Failure Error Failed Left
0.0000 1.0000 0 0 0 21
1.0000 . . . 1 20
1.0000 0.9048 0.0952 0.0641 2 19
2.0000 . . . 3 18
2.0000 0.8095 0.1905 0.0857 4 17
3.0000 0.7619 0.2381 0.0929 5 16
4.0000 . . . 6 15
4.0000 0.6667 0.3333 0.1029 7 14
5.0000 . . . 8 13
5.0000 0.5714 0.4286 0.1080 9 12
8.0000 . . . 10 11
8.0000 . . . 11 10
8.0000 . . . 12 9
8.0000 0.3810 0.6190 0.1060 13 8
11.0000 . . . 14 7
11.0000 0.2857 0.7143 0.0986 15 6
12.0000 . . . 16 5
12.0000 0.1905 0.8095 0.0857 17 4
15.0000 0.1429 0.8571 0.0764 18 3
17.0000 0.0952 0.9048 0.0641 19 2
22.0000 0.0476 0.9524 0.0465 20 1
23.0000 01.0000 0 21 0
38
SASOutputforSurvivalEstimation(cont'd)
Summary Statistics forTimeVariable t
Quartile Estimates
Point 95%Confidence Interval
Percent Estimate [Lower Upper)
7512.0000 8.0000 17.0000
50 8.0000 4.0000 11.0000
25 4.0000 2.0000 8.0000
Mean Standard Error
8.6667 1.4114
Summary oftheNumberofCensored andUncensored Values
Percent
TotalFailed Censored Censored
21 21 0 0.00
39
Doesanyonehaveaguessregardinghowtocalcu-
latethestandarderroroftheestimatedsurvival?
^S(8+)=P(T>8)=8
21=0:381
(att=8+,wecountthe4eventsattime=8asalready
havingfailed)
se[^S(8+)]=0:106
40
S-PlusCommandsforSurvivalEstimation
>t_c(1,1,2,2,3,4,4,5,5,8,8,8,8,11,11,12,12,15,17,22,23)
>surv.fit(t,status=rep(1,21))
95percent confidence interval isoftype"log"
timen.riskn.event survival std.dev lower95%CIupper95%CI
121 20.90476190 0.06405645 0.78753505 1.0000000
219 20.80952381 0.08568909 0.65785306 0.9961629
317 10.76190476 0.09294286 0.59988048 0.9676909
416 20.66666667 0.10286890 0.49268063 0.9020944
514 20.57142857 0.10798985 0.39454812 0.8276066
812 40.38095238 0.10597117 0.22084536 0.6571327
11 8 20.28571429 0.09858079 0.14529127 0.5618552
12 6 20.19047619 0.08568909 0.07887014 0.4600116
15 4 10.14285714 0.07636035 0.05010898 0.4072755
17 3 10.09523810 0.06405645 0.02548583 0.3558956
22 2 10.04761905 0.04647143 0.00703223 0.3224544
23 1 10.00000000 NA NA NA
41
Estimating theSurvivalFunction
One-samplenonparametric methods:
Wewillconsider threemethodsforestimating asurvivorship
function
S(t)=Pr(T¸t)
withoutresorting toparametric methods:
(1)Kaplan-Meier
(2)Life-table (Actuarial Estimator)
(3)viatheCumulativehazardestimator
42
(1)TheKaplan-Meier Estimator
TheKaplan-Meier (orKM)estimator isprobably
themostpopularapproach.Itcanbejusti¯ed
fromseveralperspectives:
²productlimitestimator
²likelihoodjusti¯cation
²redistribute totherightestimator
Wewillstartwithanintuitivemotivationbased
onconditional probabilities, thenreviewsomeof
theotherjusti¯cations.
43
Motivation:
First,consider anexample wherethereisnocensoring.
Thefollowingaretimesofremission (weeks)for21leukemia
patientsreceiving controltreatmen t(Table1.1ofCox&
Oakes):
1,1,2,2,3,4,4,5,5,8,8,8,8,11,11,12,12,15,17,22,23
Howwouldweestimate S(10),theprobabilit ythatanindi-
vidualsurvivestotime10orlater?
Whatabout~S(8)?Isit12
21or8
21?
Let'sconstruct atableof~S(t):
Valuesoft^S(t)
t·121/21=1.000
1<t·219/21=0.905
2<t·317/21=0.809
3<t·4
4<t·5
5<t·8
8<t·11
11<t·12
12<t·15
15<t·17
17<t·22
22<t·23
44
EmpiricalSurvivalFunction:
Whenthereisnocensoring, thegeneralformulais:
~S(t)=#individualswithT¸t
totalsamplesize
Example forleukemiadata(controlarm):
45
Whatifthereiscensoring?
Consider thetreatedgroupfromTable1.1ofCoxandOakes:
6+;6;6;6;7;9+;10+;10;11+;13;16;17+
19+;20+;22;23;25+;32+;32+;34+;35+
[Note:timeswith+arerightcensored]
WeknowS(6)=21/21,becauseeveryonesurvivedatleast
untiltime6orgreater. But,wecan'tsayS(7)=17/21,
becausewedon'tknowthestatusofthepersonwhowas
censored attime6.
Ina1958paperintheJournaloftheAmericanStatistical
Association,KaplanandMeierproposedawaytononpara-
metrically estimate S(t),eveninthepresence ofcensoring.
Themethodisbasedontheideasofconditionalproba-
bility.
46
Aquickreviewofconditionalprobability:
ConditionalProbability:SupposeAandBaretwo
events.Then,
P(AjB)=P(A\B)
P(B)
Multiplication lawofprobability:canbeobtained
fromtheaboverelationship, bymultiplying bothsidesby
P(B):
P(A\B)=P(AjB)P(B)
Extensiontomorethan2events:
SupposeA1;A2:::Akarekdi®erentevents.Then,theprob-
abilityofallkeventshappeningtogether canbewrittenas
aproductofconditional probabilities:
P(A1\A2:::\Ak)=P(AkjAk¡1\:::\A1)£
£P(Ak¡1jAk¡2\:::\A1)
:::
£P(A2jA1)
£P(A1)
47
Now,let'sapplytheseideastoestimateS(t):
Supposeak<t·ak+1.Then
S(t)=P(T¸ak+1)
=P(T¸a1;T¸a2;:::;T¸ak+1)
=P(T¸a1)£kY
j=1P(T¸aj+1jT¸aj)
=kY
j=1[1¡P(T=ajjT¸aj)]
=kY
j=1[1¡¸j]
so^S(t)»=kY
j=10
B@1¡dj
rj1
CA
=Y
j:aj<t0
B@1¡dj
rj1
CA
djisthenumberofdeathsataj
rjisthenumberatriskataj
48
IntuitionbehindtheKaplan-Meier Estimator
Thinkofdividing theobservedtimespan ofthestudyintoa
seriesof¯neintervalssothatthereisaseparate intervalfor
eachtimeofdeathorcensoring:
D C CDDD
Usingthelawofconditional probabilit y,
Pr(T¸t)=Y
jPr(survivej-thintervalIjjsurvivedtostartofIj)
wheretheproductistakenoveralltheintervalsincluding or
preceding timet.
49
4possibilities foreachinterval:
(1)Noevents(deathorcensoring) -conditional prob-
abilityofsurviving theintervalis1
(2)Censoring -assumetheysurvivetotheendofthein-
terval,sothattheconditional probabilit yofsurviving
theintervalis1
(3)Death,butnocensoring -conditional probabilit y
ofnotsurviving theintervalis#deaths(d)dividedby#
`atrisk'(r)atthebeginning oftheinterval.Sothecon-
ditionalprobabilit yofsurviving theintervalis1¡(d=r).
(4)Tieddeathsandcensoring -assumecensorings last
totheendoftheinterval,sothatconditional probabilit y
ofsurviving theintervalisstill1¡(d=r)
GeneralFormulaforjthinterval:
Itturnsoutwecanwriteageneralformulafortheconditional
probabilit yofsurviving thej-thintervalthatholdsforall4
cases:
1¡dj
rj
50
Wecouldusethesameapproachbygrouping theeventtimes
intointervals(say,oneintervalforeachmonth),andthen
countingupthenumberofdeaths(events)ineachtoesti-
matetheprobabilit yofsurviving theinterval(thisiscalled
thelifetableestimate ).
However,theassumption thatthosecensored lastuntilthe
endoftheintervalwouldn'tbequiteaccurate, sowewould
endupwithacruderapproximation.
Astheintervalsget¯nerand¯ner,theapproximations made
inestimating theprobabilities ofgettingthrougheachinter-
valbecomesmallerandsmaller, sothattheestimator con-
vergestothetrueS(t).
Thisintuitionclari¯eswhyanalternativ enamefortheKM
istheproductlimitestimator .
51
TheKaplan-Meier estimatorofthesurvivorship
function(orsurvivalprobability)S(t)=Pr(T¸t)
is:
^S(t)=Q
j:¿j<trj¡dj
rj
=Q
j:¿j<t0
@1¡dj
rj1
A
where
²¿1;:::¿KisthesetofKdistinctdeathtimesobservedin
thesample
²djisthenumberofdeathsat¿j
²rjisthenumberofindividuals \atrisk"rightbeforethe
j-thdeathtime(everyonedeadorcensored atorafter
thattime).
²cjisthenumberofcensored observationsbetweenthe
j-thand(j+1)-stdeathtimes.Censorings tiedat¿j
areincluded incj
Note:twousefulformulasare:
(1)rj=rj¡1¡dj¡1¡cj¡1
(2)rj=X
l¸j(cl+dl)
52
CalculatingtheKM-CoxandOakesexample
Makeatablewitharowforeverydeathorcensoring time:
¿jdjcjrj1¡(dj=rj)^S(¿+
j)
6312118
21=0.857
71017
90116
10
11
13
16
17
19
20
22
23
Notethat:
²^S(t+)onlychangesatdeath(failure) times
²^S(t+)is1uptothe¯rstdeathtime
²^S(t+)onlygoesto0ifthelasteventisadeath
53
KMplotfortreatedleukemiapatients
Note:moststatisticalsoftwarepackagessumma-
rizetheKMsurvivalfunctionat¿+
j,i.e.,justaf-
terthetimeofthej-thfailure.
Inotherwords,theyprovide^S(¿+
j).
Whenthereisnocensoring, theempirical survivalestimate
wouldthenbe:
~S(t+)=#individualswithT>t
totalsamplesize
54
OutputfromSTATAKMEstimator:
failure time:weeks
failure/censor: remiss
Beg. NetSurvivor Std.
Time Total FailLostFunction Error [95%Conf.Int.]
----------------------------------------------------------- --------
6 21 310.8571 0.0764 0.6197 0.9516
7 17 100.8067 0.0869 0.5631 0.9228
9 16 010.8067 0.0869 0.5631 0.9228
10 15 110.7529 0.0963 0.5032 0.8894
11 13 010.7529 0.0963 0.5032 0.8894
13 12 100.6902 0.1068 0.4316 0.8491
16 11 100.6275 0.1141 0.3675 0.8049
17 10 010.6275 0.1141 0.3675 0.8049
19 9010.6275 0.1141 0.3675 0.8049
20 8010.6275 0.1141 0.3675 0.8049
22 7100.5378 0.1282 0.2678 0.7468
23 6100.4482 0.1346 0.1881 0.6801
25 5010.4482 0.1346 0.1881 0.6801
32 4020.4482 0.1346 0.1881 0.6801
34 2010.4482 0.1346 0.1881 0.6801
35 1010.4482 0.1346 0.1881 0.6801
55
TwoOtherJusti¯cationsforKMEstimator
I.Likelihood-basedderivation(CoxandOakes)
Foradiscretefailuretimevariable,de¯ne:
djnumberoffailuresataj
rjnumberofindividuals atriskataj
(including thosecensored ataj).
¸jPr(death) inj-thinterval
(conditional onsurvivaltostartofinterval)
Thelikelihoodisthatofgindependentbinomials:
L(¸)=gY
j=1¸dj
j(1¡¸j)rj¡dj
Therefore, themaximumlikelihoodestimator of¸j
is:
^¸j=dj=rj
NowweplugintheMLE'sof¸toestimate S(t):
^S(t)=Y
j:aj<t(1¡^¸j)
=Y
j:aj<t0
B@1¡dj
rj1
CA
56
II.Redistributetotherightjusti¯cation
(Efron,1967)
Intheabsence ofcensoring, ^S(t)isjusttheproportionof
individuals withT¸t.TheideabehindEfron'sapproach
istospreadthecontributions ofcensored observationsout
overallthepossibletimestotheirright.
Algorithm:
²Step(1):arrangethenobservedtimes(deathsorcensor-
ings)inincreasing order.Ifthereareties,putcensored
afterdeaths.
²Step(2):Assignweight(1=n)toeachtime.
²Step(3):Movingfromlefttoright,eachtimeyouen-
counteracensored observation,distribute itsmasstoall
timestoitsright.
²Step(4):Calculate ^Sjbysubtracting the¯nalweight
fortimejfrom^Sj¡1
57
Exampleof\redistributetotheright"algorithm
Consider thefollowingeventtimes:
2,2.5+,3,3,4,4.5+,5,6,7
Thealgorithm goesasfollows:
(Step1) (Step4)
Times Step2Step3aStep3b^S(¿j)
21/9=0.11 0.889
2.5+1/9=0.11 0 0.889
32/9=0.22 0.25 0.635
41/9=0.11 0.13 0.508
4.5+1/9=0.11 0.13 00.508
51/9=0.11 0.13 0.17 0.339
61/9=0.11 0.13 0.17 0.169
71/9=0.11 0.13 0.17 0.000
Thiscomesoutthesameastheproductlimitapproach.
58
PropertiesoftheKMestimator
Inthecaseofnocensoring:
^S(t)=~S(t)=#deathsattorgreater
n
wherenisthenumberofindividuals inthestudy.
Thisisjustlikeanestimated probabilit yfromabinomial
distribution, sowehave:
^S(t)'N(S(t);S(t)[1¡S(t)]=n)
Howdoescensoringa®ectthis?
²^S(t)isstillapproximately normal
²Themeanof^S(t)convergestothetrueS(t)
²Thevarianceisabitmorecomplicated (sincethede-
nominatornincludes somecensored observations).
Oncewegetthevariance,thenwecanconstruct (pointwise)
(1¡®)%con¯dence intervals(NOTbands)about^S(t):
^S(t)§z1¡®=2se[^S(t)]
59
Greenwood'sformula(Collett2.1.3)
WecanthinkoftheKMestimator as
^S(t)=Y
j:¿j<t(1¡^¸j)
where^¸j=dj=rj:
Sincethe^¸j'sarejustbinomial proportions,wecanapply
standard likelihoodtheorytoshowthateach^¸jisapproxi-
matelynormal,withmeanthetrue¸j,and
var(^¸j)¼^¸j(1¡^¸j)
rj
Also,the^¸j'sareindependentinlargeenoughsamples.
Since^S(t)isafunction ofthe¸j's,wecanestimate itsvari-
anceusingthedeltamethod:
Deltamethod:IfYisnormalwithmean¹and
variance¾2,theng(Y)isapproximately normally
distributed withmeang(¹)andvariance[g0(¹)]2¾2.
60
Twospeci¯cexamplesofthedeltamethod:
(A)Z=log(Y)
thenZ»N2
64log(¹);0
@1
¹1
A2
¾23
75
(B)Z=exp(Y)
thenZ»N·
e¹;[e¹]2¾2¸
Theexamples aboveusethefollowingresultsfromcalculus:
d
dxlogu=1
u0
@du
dx1
A
d
dxeu=eu0
@du
dx1
A
61
Greenwood'sformula(continued)
Insteadofdealingwith^S(t)directly,wewilllookatitslog:
log[^S(t)]=X
j:¿j<tlog(1¡^¸j)
Thus,byapproximateindependence ofthe^¸j's,
var(log[^S(t)])=X
j:¿j<tvar[log(1¡^¸j)]
by(A) =X
j:¿j<t0
B@1
1¡^¸j1
CA2
var(^¸j)
=X
j:¿j<t0
B@1
1¡^¸j1
CA2
^¸j(1¡^¸j)=rj
=X
j:¿j<t^¸j
(1¡^¸j)rj
=X
j:¿j<tdj
(rj¡dj)rj
Now,^S(t)=exp[log[^S(t)]].Thusby(B),
var(^S(t))=[^S(t)]2var·
log[^S(t)]¸
Greenwood'sFormula:
var(^S(t))=[^S(t)]2P
j:¿j<tdj
(rj¡dj)rj
62
Backtocon¯denceintervals
Fora95%con¯dence interval,wecoulduse
^S(t)§z1¡®=2se[^S(t)]
wherese[^S(t)]iscalculated usingGreenwood'sformula.
Problem: Thisapproachcanyieldvalues>1or<0.
Betterapproach:Geta95%con¯dence intervalfor
L(t)=log(¡log(S(t)))
Sincethisquantityisunrestricted, thecon¯dence interval
willbeintheproperrangewhenwetransform back.
Toseewhythisworks,notethefollowing:
²Since^S(t)isanestimated probabilit y
0·^S(t)·1
²Takingthelogof^S(t)hasbounds:
¡1·log[^S(t)]·0
²Takingtheopposite:
0·¡log[^S(t)]·1
²Takingthelogagain:
¡1·log·
¡log[^S(t)]¸
·1
Totransform back,reversestepswithS(t)=exp(¡exp(L(t))
63
Log-logApproachforCon¯denceIntervals:
(1)De¯neL(t)=log(¡log(S(t)))
(2)Forma95%con¯dence intervalforL(t)basedon^L(t),
yielding[^L(t)¡A;^L(t)+A]
(3)SinceS(t)=exp(¡exp(L(t)),thecon¯dence bounds
forthe95%CIonS(t)are:
[exp(¡e(^L(t)+A));exp(¡e(^L(t)¡A))]
(notethattheupperandlowerboundsswitch)
(4)Substituting ^L(t)=log(¡log(^S(t)))backintotheabove
bounds,wegetcon¯dence boundsof
([^S(t)]eA;[^S(t)]e¡A)
64
WhatisA?
²Ais1:96se(^L(t))
²Tocalculate this,weneedtocalculate
var(^L(t))=var·
log(¡log(^S(t)))¸
²Fromourprevious calculations, weknow
var(log[^S(t)])=X
j:¿j<tdj
(rj¡dj)rj
²Applying thedeltamethodasinexample (A),weget:
var(^L(t))=var(log(¡log[^S(t)]))
=1
[log^S(t)]2X
j:¿j<tdj
(rj¡dj)rj
²Wetakethesquarerootoftheabovetogetse(^L(t)),
andthenformthecon¯dence intervalsas:
^S(t)e§1:96se(^L(t))
²ThisistheapproachthatStatauses.Splusgivesanop-
tiontocalculate thesebounds(useconf.type=''log-log''
insurv.fit ).
65
SummaryofCon¯denceIntervalsonS(t)
²Calculate ^S(t)§1:96se[^S(t)]wherese[^S(t)]iscalcu-
latedusingGreenwood'sformula,andreplacenegative
lowerboundsby0andupperboundsgreaterthan1by
1.
{Recommended byCollett
{ThisisthedefaultusingSAS
{notverysatisfactory
²Usealogtransformation tostabilize thevarianceand
allowfornon-symmetric con¯dence intervals.Thisis
whatisnormally doneforthecon¯dence intervalofan
estimated oddsratio.
{Usevar[log(^S(t))]=P
j:¿j<tdj
(rj¡dj)rjalreadycalcu-
latedaspartofGreenwood'sformula
{ThisisthedefaultinSplus
²Usethelog-logtransformation justdescribed
{Somewhat complicated, butalwaysyieldsproperbounds
{ThisisthedefaultinStata.
66
SoftwareforKaplan-Meier Curves
²Stata-stsetandstscommands
²SAS-proclifetest
²Splus-surv.¯t(time,status)
DefaultsforCon¯denceIntervalCalculations
²Stata-\log-log")^L(t)§1:96se[^L(t)]
whereL(t)=log[¡log(S(t))]
²SAS-\plain")^S(t)§1:96se[^S(t)]
²Splus-\log")logS(t)§1:96se[log(^S(t))]
butSpluswillalsogiveeitheroftheothertwooptionsif
yourequestthem.
67
StataCommands
Createa¯lecalled\leukemia.dat" withtherawdata,with
acolumnfortreatmen t,weekstorelapse(i.e.,duration of
remission), andrelapsestatus:
.infile trtremissstatususingleukemia.dat
.stsetremissstatus (setsupafailure timedataset,
withfailtime statusinthatorder,
typehelpstsettogetdetails)
.stslist (estimated S(t),se[S(t)], and95%CI)
.stsgraph,saving(kmtrt) (creates aKaplan-Meier plot,and
savestheplotinfilekmtrt.gph,
type``helpgphdot'' togetsome
printing instructions)
.graphusingkmtrt (redisplays thegraphatanylatertime)
IfthedatasethasalreadybeencreatedandloadedintoStata,
thenyoucansubstitute thefollowingcommands forinitial-
izingthedata:
.useleukem (findsStatadataset leukem.dta)
.describe (provides adescription ofthedataset)
.stsetremissstatus (declares datatobefailure type)
.stdes (givesadescription ofthesurvival dataset)
68
STATAOutputforTreatedLeukemiaPatients:
.useleukem
.stsetremissstatusiftrt==1
.stslist
failure time:remiss
failure/censor: status
Beg. NetSurvivor Std.
Time Total FailLostFunction Error [95%Conf.Int.]
----------------------------------------------------------- --------
6 21 310.8571 0.0764 0.6197 0.9516
7 17 100.8067 0.0869 0.5631 0.9228
9 16 010.8067 0.0869 0.5631 0.9228
10 15 110.7529 0.0963 0.5032 0.8894
11 13 010.7529 0.0963 0.5032 0.8894
13 12 100.6902 0.1068 0.4316 0.8491
16 11 100.6275 0.1141 0.3675 0.8049
17 10 010.6275 0.1141 0.3675 0.8049
19 9010.6275 0.1141 0.3675 0.8049
20 8010.6275 0.1141 0.3675 0.8049
22 7100.5378 0.1282 0.2678 0.7468
23 6100.4482 0.1346 0.1881 0.6801
25 5010.4482 0.1346 0.1881 0.6801
32 4020.4482 0.1346 0.1881 0.6801
34 2010.4482 0.1346 0.1881 0.6801
35 1010.4482 0.1346 0.1881 0.6801
69
SASCommandsforKaplanMeierEstimator-
PROCLIFETEST
TheSAScommand fortheKaplan-Meier estimate is:
timefailtime*censor(1);
or timefailtime*failind(0);
The¯rstvariableisthefailuretime,andthesecondisthe
failureorcensoring indicator. Inparenthesesyouneedtoput
thespeci¯cnumericvaluethatcorrespondstocensoring.
Theupperandlowercon¯dence limitson^S(t)areincluded
inthedataset\OUTSUR V"whenspeci¯ed.Theupperand
lowerlimitsarecalled:sdfucl,sdflcl.
dataleukemia;
inputweeksremiss;
labelweeks='Time toRemission (inweeks)'
remiss='Remission indicator (1=yes,0=no)';
cards;
61
61
........... (lineseditedouthere)
340
350
;
proclifetest data=leukemia outsurv=confint;
timeweeks*remiss(0);
title'Leukemia datafromTable1.1ofCoxandOakes';
run;
procprintdata=confint;
title'95%Confidence Intervals forEstimated Survival';
70
OutputfromSASProcLifetest
Note:thisinformation isnotprintedifyouuseNOPRINT.
Leukemia datafromTable1.1ofCoxandOakes
TheLIFETEST Procedure
Product-Limit Survival Estimates
Survival
Standard Number Number
WEEKS Survival Failure Error Failed Left
0.0000 1.0000 0 0 0 21
6.0000 . . . 1 20
6.0000 . . . 2 19
6.0000 0.8571 0.1429 0.0764 3 18
6.0000* . . . 3 17
7.0000 0.8067 0.1933 0.0869 4 16
9.0000* . . . 4 15
10.0000 0.7529 0.2471 0.0963 5 14
10.0000* . . . 5 13
11.0000* . . . 5 12
13.0000 0.6902 0.3098 0.1068 6 11
16.0000 0.6275 0.3725 0.1141 7 10
17.0000* . . . 7 9
19.0000* . . . 7 8
20.0000* . . . 7 7
22.0000 0.5378 0.4622 0.1282 8 6
23.0000 0.4482 0.5518 0.1346 9 5
25.0000* . . . 9 4
32.0000* . . . 9 3
32.0000* . . . 9 2
34.0000* . . . 9 1
35.0000* . . . 9 0
*Censored Observation
71
OutputfromprintingtheCONFINT¯le
95%Confidence Intervals forEstimated Survival
OBS WEEKS _CENSOR_ SURVIVAL SDF_LCL SDF_UCL
1 0 0 1.00000 1.00000 1.00000
2 6 0 0.85714 0.70748 1.00000
3 6 1 0.85714 . .
4 7 0 0.80672 0.63633 0.97711
5 9 1 0.80672 . .
6 10 0 0.75294 0.56410 0.94178
7 10 1 0.75294 . .
8 11 1 0.75294 . .
9 13 0 0.69020 0.48084 0.89955
10 16 0 0.62745 0.40391 0.85099
11 17 1 0.62745 . .
12 19 1 0.62745 . .
13 20 1 0.62745 . .
14 22 0 0.53782 0.28648 0.78915
15 23 0 0.44818 0.18439 0.71197
16 25 1 . . .
17 32 1 . . .
18 32 1 . . .
19 34 1 . . .
20 35 1 . . .
Theoutputdatasetwillhaveoneobservationforeachunique
combination ofweeksandcensor .Itwillalsoaddan
observationforfailuretimeequalto0.
72
SplusCommands
Createa¯lecalled\leukemia.dat" withthevariablesnames
inthe¯rstrow,asfollows:
tc
61
61
etc...
InSplus,type
y_read.table('leukemia.dat',header=T)
surv.fit(y$t,y$c)
plot(surv.fit(y$t,y$c))
(theplotcommand willalsoyield95%con¯dence intervals)
Tospecifythetypeofcon¯dence intervals,usetheconf.type=
optioninthesurv.¯tstatemen ts:e.g.conf.type=\log-log"
orconf.type=\plain"
73
>surv.fit(y$t,y$c)
95percent confidence interval isoftype"log"
timen.riskn.event survival std.dev lower95%CIupper95%CI
621 30.8571429 0.07636035 0.7198171 1.0000000
717 10.8067227 0.08693529 0.6531242 0.9964437
1015 10.7529412 0.09634965 0.5859190 0.9675748
1312 10.6901961 0.10681471 0.5096131 0.9347692
1611 10.6274510 0.11405387 0.4393939 0.8959949
22 7 10.5378151 0.12823375 0.3370366 0.8582008
23 6 10.4481793 0.13459146 0.2487882 0.8073720
>surv.fit(y$t,y$c,conf.type="log-log")
95percent confidence interval isoftype"log-log"
timen.riskn.event survival std.dev lower95%CIupper95%CI
621 30.8571429 0.07636035 0.6197180 0.9515517
717 10.8067227 0.08693529 0.5631466 0.9228090
1015 10.7529412 0.09634965 0.5031995 0.8893618
1312 10.6901961 0.10681471 0.4316102 0.8490660
1611 10.6274510 0.11405387 0.3675109 0.8049122
22 7 10.5378151 0.12823375 0.2677789 0.7467907
23 6 10.4481793 0.13459146 0.1880520 0.6801426
>surv.fit(y$t,y$c,conf.type="plain")
95percent confidence interval isoftype"plain"
timen.riskn.event survival std.dev lower95%CIupper95%CI
621 30.8571429 0.07636035 0.7074793 1.0000000
717 10.8067227 0.08693529 0.6363327 0.9771127
1015 10.7529412 0.09634965 0.5640993 0.9417830
1312 10.6901961 0.10681471 0.4808431 0.8995491
1611 10.6274510 0.11405387 0.4039095 0.8509924
22 7 10.5378151 0.12823375 0.2864816 0.7891487
23 6 10.4481793 0.13459146 0.1843849 0.7119737
74
KMSurvivalEstimateandCon¯denceintervals
(SPlus)
TimeSurvival
0 5 10 15 20 25 30 350.0 0.2 0.4 0.6 0.8 1.0
75
Means,Medians,QuantilesbasedontheKM
²Mean:Pk
j=1¿jPr(T=¿j)
²Median -byde¯nition, thisisthetime,¿,suchthat
S(¿)=0:5.However,inpractice, itisde¯nedasthe
smallest timesuchthat^S(¿)·0:5.Themedianismore
appropriate forcensored survivaldatathanthemean.
Forthetreatedleukemiapatients,we¯nd:
^S(22)=0:5378
^S(23)=0:4482
Themedianisthus23.Thiscanalsobeseenvisuallyon
thegraphtotheleft.
²Lowerquartile(25thpercentile):
thesmallest time(LQ)suchthat^S(LQ)·0:75
²Upperquartile(75thpercentile):
thesmallest time(UQ)suchthat^S(UQ)·0:25
76
The(2)LifetableEstimatorofSurvival:
Wesaidthatwewouldconsider thefollowingthreemethods
forestimating asurvivorshipfunction
S(t)=Pr(T¸t)
withoutresorting toparametric methods:
(1)pKaplan-Meier
(2)=)Life-table (Actuarial Estimator)
(3)=)Cumulativehazardestimator
77
(2)TheLifetableorActuarialEstimator
²oneoftheoldesttechniquesaround
²usedbyactuaries, demographers, etc.
²applieswhenthedataaregrouped
Ourgoalisstilltoestimate thesurvivalfunction, hazard,and
densityfunction, butthisiscomplicated bythefactthatwe
don'tknowexactlywhenduringeachtimeintervalanevent
occurs.
78
Lee(section 4.2)providesagooddescription oflifetable
methods,anddistinguishes severaltypesaccording tothe
datasources:
Popula tionLifeTables
²cohortlifetable-describesthemortalityexperience
frombirthtodeathforaparticular cohortofpeopleborn
ataboutthesametime.Peopleatriskatthestartofthe
intervalarethosewhosurvivedtheprevious interval.
²currentlifetable-constructed from(1)censusinfor-
mationonthenumberofindividuals aliveateachage,
foragivenyearand(2)vitalstatistics onthenumber
ofdeathsorfailuresinagivenyear,byage.Thistype
oflifetable isoftenreportedintermsofahypothetical
cohortof100,000people.
Generally ,censoring isnotanissueforPopulation LifeTa-
bles.
Clinical Lifetables-appliestogroupedsurvivaldata
fromstudiesinpatientswithspeci¯cdiseases. Because pa-
tientscanenterthestudyatdi®erenttimes,orbelostto
follow-up,censoring mustbeallowed.
79
Notation
²thej-thtimeintervalis[tj¡1;tj)
²cj-thenumberofcensorings inthej-thinterval
²dj-thenumberoffailuresinthej-thinterval
²rjisthenumberenteringtheinterval
Example :2418MaleswithAnginaPectoris(Lee,p.91)
Yearafter
Diagnosis jdjcjrjr0
j=rj¡cj=2
[0;1)1456024182418.0
[1;2)22263919621942.5 (1962-39
2)
[2;3)31522216971686.0
[3;4)41712315231511.5
[4;5)51352413291317.0
[5;6)612510711701116.5
[6;7)783133938871.5
etc..
80
Estimatingthesurvivorshipfunction
WecouldapplytheK-Mformuladirectlytothenumbersin
thetableontheprevious page,estimating S(t)as
^S(t)=Y
j:¿j<t0
B@1¡dj
rj1
CA
However,thisapproachisunsatisfactory forgroupeddata....
ittreatstheproblem asthoughitwereindiscretetime,with
eventshappeningonlyat1yr,2yr,etc.Infact,whatwe
aretryingtocalculate hereistheconditional probabilit yof
dyingwithintheinterval,givensurvivaltothebeginning of
it.
Whatshouldwedowiththecensoredpeople?
Wecanassumethatcensoringsoccur:
²atthebeginning ofeachinterval:r0
j=rj¡cj
²attheendofeachinterval:r0
j=rj
²onaveragehalfwaythrough theinterval:
r0
j=rj¡cj=2
Thelastassumption yieldstheActuarial Estimator. Itis
appropriate ifcensorings occuruniformly throughout thein-
terval.
81
Constructingthelifetable
First,someadditional notation forthej-thinterval,[tj¡1;tj):
²Midpoint(tmj)-usefulforplotting thedensityand
thehazardfunction
²Width(bj=tj¡tj¡1)neededforcalculating thehazard
inthej-thinterval
Quantitiesestimated:
²Conditional probabilit yofdying
^qj=dj=r0
j
²Conditional probabilit yofsurviving
^pj=1¡^qj
²Cumulativeprobabilit yofsurviving attj:
^S(tj)=Y
`·j^p`
=Y
`·j0
@1¡d`
r`01
A
82
Someimportantpointstonote:
²Because theintervalsarede¯nedas[tj¡1;tj),the¯rst
intervaltypicallystartswitht0=0.
²Stataestimates thesurvivalfunction attheright-hand
endpointofeachinterval,i.e.,S(tj)
²However,SASestimates thesurvivalfunction attheleft-
handendpoint,S(tj¡1).
²Theimplication inSASisthat^S(t0)=1and^S(t1)=p1
83
Otherquantitiesestimatedatthe
midpointofthej-thinterval:
²Hazard inthej-thinterval:
^¸(tmj)=dj
bj(r0j¡dj=2)
=^qj
bj(1¡^qj=2)
thenumberofdeathsintheintervaldividedbytheav-
eragenumberofsurvivorsatthemidpoint
²densityatthemidpointofthej-thinterval:
^f(tmj)=^S(tj¡1)¡^S(tj)
bj
=^S(tj¡1)^qj
bj
Note:Another waytogetthisis:
^f(tmj)=^¸(tmj)^S(tmj)
=^¸(tmj)[^S(tj)+^S(tj¡1)]=2
84
ConstructingtheLifetableusingStata
Usestheltablecommand.
Iftherawdataarealreadygrouped,thenthefreqstatemen t
mustbeusedwhenreadingthedata.
.infileyearsstatuscountusingangina.dat
(32observations read)
.ltableyearsstatus[freq=count]
Beg. Std.
Interval TotalDeaths LostSurvival Error [95%Conf.Int.]
------------------------------------------------------------------- ------
012418 456 00.8114 0.0080 0.7952 0.8264
121962 226390.7170 0.0092 0.6986 0.7346
231697 152220.6524 0.0097 0.6329 0.6711
341523 171230.5786 0.0101 0.5584 0.5981
451329 135240.5193 0.0103 0.4989 0.5392
561170 1251070.4611 0.0104 0.4407 0.4813
67938 831330.4172 0.0105 0.3967 0.4376
78722 741020.3712 0.0106 0.3505 0.3919
89546 51680.3342 0.0107 0.3133 0.3553
910427 42640.2987 0.0109 0.2775 0.3201
1011321 43450.2557 0.0111 0.2341 0.2777
1112233 34530.2136 0.0114 0.1917 0.2363
1213146 18330.1839 0.0118 0.1614 0.2075
1314 95 9270.1636 0.0123 0.1404 0.1884
1415 59 6230.1429 0.0133 0.1180 0.1701
1516 30 0300.1429 0.0133 0.1180 0.1701
------------------------------------------------------------------- -------- ----
85
Itisalsopossibletogetestimates ofthehazardfunction, ^¸j,
anditsstandard errorusingthe\hazard"option:
.ltableyearsstatus[freq=count], hazard
Beg. Cum. Std. Std.
Interval Total Failure Error Hazard Error [95%ConfInt]
------------------------------------------------------------------- -------
012418 0.1886 0.0080 0.2082 0.0097 0.1892 0.2272
121962 0.2830 0.0092 0.1235 0.0082 0.1075 0.1396
231697 0.3476 0.0097 0.0944 0.0076 0.0794 0.1094
341523 0.4214 0.0101 0.1199 0.0092 0.1020 0.1379
451329 0.4807 0.0103 0.1080 0.0093 0.0898 0.1262
561170 0.5389 0.0104 0.1186 0.0106 0.0978 0.1393
679380.5828 0.0105 0.1000 0.0110 0.0785 0.1215
787220.6288 0.0106 0.1167 0.0135 0.0902 0.1433
895460.6658 0.0107 0.1048 0.0147 0.0761 0.1336
9104270.7013 0.0109 0.1123 0.0173 0.0784 0.1462
10113210.7443 0.0111 0.1552 0.0236 0.1090 0.2015
11122330.7864 0.0114 0.1794 0.0306 0.1194 0.2395
12131460.8161 0.0118 0.1494 0.0351 0.0806 0.2182
1314 950.8364 0.0123 0.1169 0.0389 0.0407 0.1931
1415 590.8571 0.0133 0.1348 0.0549 0.0272 0.2425
1516 300.8571 0.0133 0.0000 . . .
------------------------------------------------------------------- ------
Thereisalsoa\failure "optionwhichgivesthenumberof
failures(likethedefault), andalsoprovidesa95%con¯dence
intervalonthecumulativefailureprobabilit y.
86
ConstructingthelifetableusingSAS
Iftherawdataarealreadygrouped,thentheFREQstate-
mentmustbeusedwhenreadingthedata.
SASrequires thattheintervalendpointsbespeci¯ed,using
oneofthefollowing(seeSASmanualoronlinehelpformore
detail):
²intervals-specifythetheintervalendpoints
²width-specifythewidthofeachinterval
²ninterval-specifythenumberofintervals
Title'Actuarial Estimator forAnginaPectoris Example';
dataangina;
inputyearsstatuscount;
cards;
0.51456
1.51226
2.51152 /*anginacases*/
3.51171
4.51135
5.51125
.
.
0.500
1.5039
2.5022 /*censored */
3.5023
4.5024
5.50107
.
.
proclifetest data=angina outsurv=survres intervals=0 to15by1method=act;
timeyears*status(0);
freqcount;
87
SASoutput:
Actuarial Estimator forAnginaPectoris Example
TheLIFETEST Procedure
LifeTableSurvival Estimates
Conditional
Effective Conditional Probability
Interval Number Number Sample Probability Standard
[Lower, Upper) Failed Censored Size ofFailure Error
0 1456 02418.0 0.1886 0.00796
1 2226 39 1942.5 0.1163 0.00728
2 3152 22 1686.0 0.0902 0.00698
3 4171 23 1511.5 0.1131 0.00815
4 5135 24 1317.0 0.1025 0.00836
5 6125 107 1116.5 0.1120 0.00944
6 783 133 871.5 0.0952 0.00994
7 874 102 671.0 0.1103 0.0121
8 951 68 512.0 0.0996 0.0132
9 10 42 64 395.0 0.1063 0.0155
10 11 43 45 298.5 0.1441 0.0203
11 12 34 53 206.5 0.1646 0.0258
12 13 18 33 129.5 0.1390 0.0304
13 14 9 27 81.5 0.1104 0.0347
14 15 6 23 47.5 0.1263 0.0482
15 . 0 30 15.0 0 0
Survival Median Median
Interval Standard Residual Standard
[Lower, Upper) Survival Failure Error Lifetime Error
0 11.0000 0 05.3313 0.1749
1 20.8114 0.1886 0.00796 6.2499 0.2001
2 30.7170 0.2830 0.00918 6.3432 0.2361
3 40.6524 0.3476 0.00973 6.2262 0.2361
4 50.5786 0.4214 0.0101 6.2185 0.1853
5 60.5193 0.4807 0.0103 5.9077 0.1806
6 70.4611 0.5389 0.0104 5.5962 0.1855
7 80.4172 0.5828 0.0105 5.1671 0.2713
8 90.3712 0.6288 0.0106 4.9421 0.2763
9 100.3342 0.6658 0.0107 4.8258 0.4141
10 110.2987 0.7013 0.0109 4.6888 0.4183
11 120.2557 0.7443 0.0111 . .
12 130.2136 0.7864 0.0114 . .
13 140.1839 0.8161 0.0118 . .
14 150.1636 0.8364 0.0123 . .
15 .0.1429 0.8571 0.0133 . .
88
moreSASoutput: (estimated density^fjandhazard^¸j)
Evaluated attheMidpoint oftheInterval
PDF Hazard
Interval Standard Standard
[Lower, Upper) PDF Error Hazard Error
0 10.1886 0.00796 0.208219 0.009698
1 20.0944 0.00598 0.123531 0.008201
2 30.0646 0.00507 0.09441 0.007649
3 40.0738 0.00543 0.119916 0.009154
4 50.0593 0.00495 0.108043 0.009285
5 60.0581 0.00503 0.118596 0.010589
6 70.0439 0.00469 0.10.010963
7 80.0460 0.00518 0.116719 0.013545
8 90.0370 0.00502 0.10483 0.014659
9 100.0355 0.00531 0.112299 0.017301
10 110.0430 0.00627 0.155235 0.023602
11 120.0421 0.00685 0.17942 0.030646
12 130.0297 0.00668 0.149378 0.03511
13 140.0203 0.00651 0.116883 0.038894
14 150.0207 0.00804 0.134831 0.054919
15 . . . . .
Summary oftheNumberofCensored andUncensored Values
Total Failed Censored %Censored
2418 1625 79332.7957
89
Supposewewishtousetheactuarial method,butthedata
donotcomegrouped.
Consider thetreatednursinghomepatients,withlengthof
stay(los)groupedinto100dayintervals:
.usenurshome
.dropifrx==0 (keeponlythetreated patients)
(881observations deleted)
.stsetlosfail
.ltable losfail,intervals(100)
Beg. Std.
Interval TotalDeaths LostSurvival Error [95%Conf.Int.]
------------------------------------------------------------------- -----
0100 710328 00.5380 0.0187 0.5006 0.5739
100200 382 86 00.4169 0.0185 0.3805 0.4529
200300 296 65 00.3254 0.0176 0.2911 0.3600
300400 231 38 00.2718 0.0167 0.2396 0.3050
400500 193 32 10.2266 0.0157 0.1966 0.2581
500600 160 13 00.2082 0.0152 0.1792 0.2388
600700 147 13 00.1898 0.0147 0.1619 0.2195
700800 134 10300.1739 0.0143 0.1468 0.2029
800900 94 4290.1651 0.0143 0.1383 0.1941
9001000 61 4300.1508 0.0147 0.1233 0.1808
10001100 27 0270.1508 0.0147 0.1233 0.1808
------------------------------------------------------------------- ------
90
SASCommandsforlifetableanalysis-grouping
data
Title'Actuarial Estimator fornursing homedata';
datamorris;
infile'ch12.dat' ;
inputlosagetrtgendermarstat hltstat cens;
datamorristr;
setmorris;
iftrt=1;
proclifetest data=morristr outsurv=survres
intervals=0 to1100by100method=act;
timelos*cens(1);
run;
procprintdata=survres;
run;
91
Actuarial estimator fortreatednursinghomepatients
Actuarial Estimator forNursing HomePatients
TheLIFETEST Procedure
LifeTableSurvival Estimates
Effective Conditional
Interval Number Number Sample Probability
[Lower, Upper) Failed Censored Size ofFailure
0100 330 0 712.0 0.4635
100 200 86 0 382.0 0.2251
200 300 65 0 296.0 0.2196
300 400 38 0 231.0 0.1645
400 500 32 1 192.5 0.1662
500 600 13 0 160.0 0.0813
600 700 13 0 147.0 0.0884
700 800 10 30 119.0 0.0840
800 900 4 29 79.5 0.0503
900 1000 4 30 46.0 0.0870
1000 1100 0 27 13.5 0
Conditional
Probability Survival Median
Interval Standard Standard Residual
[Lower, Upper) Error Survival Failure Error Lifetime
0100 0.0187 1.0000 0 0130.2
100 200 0.0214 0.5365 0.4635 0.0187 306.2
200 300 0.0241 0.4157 0.5843 0.0185 398.8
300 400 0.0244 0.3244 0.6756 0.0175 617.0
400 500 0.0268 0.2711 0.7289 0.0167 .
500 600 0.0216 0.2260 0.7740 0.0157 .
600 700 0.0234 0.2076 0.7924 0.0152 .
700 800 0.0254 0.1893 0.8107 0.0147 .
800 900 0.0245 0.1734 0.8266 0.0143 .
900 1000 0.0415 0.1647 0.8353 0.0142 .
1000 1100 00.1503 0.8497 0.0147 .
92
Actuarial estimator fortreatednursinghomepatients,cont'd
Evaluated attheMidpoint
oftheInterval
Median PDF Hazard
Interval Standard Standard Standard
[Lower, Upper) Error PDF Error Hazard Error
010015.5136 0.00463 0.000187 0.006033 0.000317
100 20030.4597 0.00121 0.000122 0.002537 0.000271
200 30065.7947 0.000913 0.000108 0.002467 0.000304
300 40074.5466 0.000534 0.000084 0.001792 0.00029
400 500 .0.000451 0.000078 0.001813 0.000319
500 600 .0.000184 0.00005 0.000847 0.000235
600 700 .0.000184 0.00005 0.000925 0.000256
700 800 .0.000159 0.00005 0.000877 0.000277
800 900 .0.000087 0.000043 0.000516 0.000258
900 1000 .0.000143 0.00007 0.000909 0.000454
1000 1100 . 0 . 0 .
Summary oftheNumberofCensored andUncensored Values
Total Failed Censored %Censored
712 595 11716.4326
93
Actuarial estimator fortreatednursinghomepatients,cont'd
OutputfromSURVRESdataset
Actuarial Estimator forNursing HomePatients
OBS LOSSURVIVAL SDF_LCL SDF_UCL MIDPOINT PDF
1 01.00000 1.00000 1.00000 50 .0046348
2100 0.53652 0.49989 0.57315 150 .0012079
3200 0.41573 0.37953 0.45193 250 .0009129
4300 0.32444 0.29005 0.35883 350 .0005337
5400 0.27107 0.23842 0.30372 450 .0004506
6500 0.22601 0.19528 0.25674 550 .0001836
7600 0.20764 0.17783 0.23745 650 .0001836
8700 0.18928 0.16048 0.21808 750 .0001591
9800 0.17337 0.14536 0.20139 850 .0000872
10900 0.16465 0.13677 0.19253 950 .0001432
111000 0.15033 0.12157 0.17910 1050 .0000000
OBS PDF_LCL PDF_UCL HAZARD HAZ_LCL HAZ_UCL
1.0042685 .0050011 .0060329 .0054123 .0066535
2.0009685 .0014472 .0025369 .0020050 .0030687
3.0007014 .0011245 .0024668 .0018717 .0030619
4.0003686 .0006988 .0017925 .0012248 .0023601
5.0002981 .0006031 .0018130 .0011874 .0024386
6.0000847 .0002825 .0008469 .0003869 .0013069
7.0000847 .0002825 .0009253 .0004228 .0014277
8.0000617 .0002565 .0008772 .0003340 .0014203
9.0000027 .0001717 .0005161 .0000105 .0010218
10.0000069 .0002794 .0009091 .0000191 .0017991
11. . .0000000 . .
94
ExamplesforNursinghomedata:
EstimatedSurvival:
Estimated Survival
0.00.10.20.30.40.50.60.70.80.91.0
Lower Limit of Time Interval01002003004005006007008009001000
95
Estimatedhazard:
Estimated hazard
0.0000.0020.0040.0060.0080.010
Lower Limit of Time Interval01002003004005006007008009001000
96
(3)Estimatingthecumulativehazard
(Nelson-Aalen estimator)
Supposewewanttoestimate ¤(t)=Rt
0¸(u)du,thecumula-
tivehazardattimet.
JustaswedidfortheKM,thinkofdividing theobserved
timespan ofthestudyintoaseriesof¯neintervalssothat
thereisonlyoneeventperinterval:
D C CDDD
¤(t)canthenbeapproximated byasum:
^¤(t)=X
j¸j¢
wherethesumisoverintervals,¸jisthevalueofthehazard
inthej-thintervaland¢isthewidthofeachinterval.Since
^¸¢isapproximately theprobabilit yofdyingintheinterval,
wecanfurtherapproximateby
^¤(t)=X
jdj=rj
Itfollowsthat¤(t)willchangeonlyatdeathtimes,and
hencewewritetheNelson-Aalen estimator as:
^¤NA(t)=X
j:¿j<tdj=rj
97
D C CDDD
rjnnnn-
1n-
1n-2n-2n-3n-4
dj001000011
cj000010100
^¸(tj)001/n00001
n¡31
n¡4
^¤(tj)001/n1/n1/n1/n1/n
Oncewehave^¤NA(t),wecanalso¯ndanotherestimator of
S(t)(Fleming-Harrington):
^SFH(t)=exp(¡^¤NA(t))
Ingeneral, thisestimator ofthesurvivalfunction willbe
closetotheKaplan-Meier estimator, ^SKM(t)
Wecanalsogotheotherway...wecantaketheKaplan-
Meierestimate ofS(t),anduseittocalculate analternativ e
estimate ofthecumulativehazardfunction:
^¤KM(t)=¡log^SKM(t)
98
StatacommandsforFHSurvivalEstimate
SaywewanttoobtaintheFleming-Harrington estimate of
thesurvivalfunction formarried females, inthehealthiest
initialsubgroup, whoarerandomized totheuntreatedgroup
ofthenursinghomestudy.
First,weusethefollowingcommands tocalculate theNelson-
Aalencumulativehazardestimator:
.usenurshome
.keepifrx==0&gender==0 &health==2 &married==1
(1579observations deleted)
.stslist,na
failure _d:fail
analysis time_t:los
Beg. NetNelson-Aalen Std.
Time Total FailLost Cum.Haz. Error [95%Conf.Int.]
------------------------------------------------------------------- ---
14 12 100.0833 0.0833 0.0117 0.5916
24 11 100.1742 0.1233 0.0435 0.6976
25 10 100.2742 0.1588 0.0882 0.8530
38 9100.3854 0.1938 0.1438 1.0326
64 8100.5104 0.2306 0.2105 1.2374
89 7100.6532 0.2713 0.2894 1.4742
113 6100.8199 0.3184 0.3830 1.7551
123 5101.0199 0.3760 0.4952 2.1006
149 4101.2699 0.4515 0.6326 2.5493
168 3101.6032 0.5612 0.8073 3.1840
185 2102.1032 0.7516 1.0439 4.2373
234 1103.1032 1.2510 1.4082 6.8384
------------------------------------------------------------------- ---
99
Aftergenerating theNelson-Aalen estimator, wemanually
havetocreateavariableforthesurvivalestimate:
.stsgennelson=na
.gensfh=exp(-nelson)
.listsfh
sfh
1..9200444
2..8400932
3..7601478
4..6802101
5..6002833
6..5203723
7..4404857
8..3606392
9..2808661
10..2012493
11..1220639
12..0449048
Additional built-infunctions canbeusedtogenerate 95%
con¯dence intervalsontheFHsurvivalestimate.
100
Wecancompare theFleming-Harrington survivalestimate
totheKMestimate byrerunning thestslistcommand:
.stslist
.stsgenskm=s
.listskmsfh
skm sfh
1..91666667 .9200444
2..83333333 .8400932
3. .75.7601478
4..66666667 .6802101
5..58333333 .6002833
6. .5.5203723
7..41666667 .4404857
8..33333333 .3606392
9. .25.2808661
10..16666667 .2012493
11..08333333 .1220639
12. 0.0449048
Inthisexample, itlooksliketheFleming-Harrington estima-
torisslightlyhigherthantheKMateverytimepoint,but
withlargerdatasets thetwowilltypicallybemuchcloser.
101
SplusCommandsforFleming-Harrington Esti-
mator:
(Nursing homedata:females, untreated,married, healthy)
Fleming-Harrington:
>fh<-surv.fit(los,cens,type="f",conf.type="log-log")
>fh
95percent confidence interval isoftype"log-log"
timen.riskn.event survival std.dev lower95%CIupper95%CI
1412 10.9200444 0.08007959 0.5244209125 0.9892988
2411 10.8400932 0.10845557 0.4750041174 0.9600371
2510 10.7601478 0.12669130 0.4055610500 0.9200425
38 9 10.6802101 0.13884731 0.3367907188 0.8724502
64 8 10.6002833 0.14645413 0.2718422278 0.8187596
89 7 10.5203723 0.15021856 0.2115701242 0.7597900
113 6 10.4404857 0.15045450 0.1564397006 0.6960354
123 5 10.3606392 0.14723033 0.1069925657 0.6278888
149 4 10.2808661 0.14043303 0.0640979523 0.5560134
168 3 10.2012493 0.12990589 0.0293208029 0.4827590
185 2 10.1220639 0.11686728 0.0058990525 0.4224087
234 1 10.0449048 0.06216787 0.0005874321 0.2740658
Kaplan-Meier:
>km<-surv.fit(los,cens,conf.type="log-log")
>km
95percent confidence interval isoftype"log-log"
timen.riskn.event survival std.dev lower95%CIupper95%CI
1412 10.91666667 0.07978559 0.538977181 0.9878256
2411 10.83333333 0.10758287 0.481714942 0.9555094
2510 10.75000000 0.12500000 0.408415913 0.9117204
38 9 10.66666667 0.13608276 0.337018933 0.8597118
64 8 10.58333333 0.14231876 0.270138924 0.8009402
89 7 10.50000000 0.14433757 0.208477143 0.7360731
113 6 10.41666667 0.14231876 0.152471264 0.6653015
123 5 10.33333333 0.13608276 0.102703980 0.5884189
149 4 10.25000000 0.12500000 0.060144556 0.5047588
168 3 10.16666667 0.10758287 0.026510427 0.4129803
185 2 10.08333333 0.07978559 0.005052835 0.3110704
234 1 10.00000000 NA NA NA
102
Comparison ofSurvivalCurves
Wespentthelastclasslookingatsomenonparametric ap-
proachesforestimating thesurvivalfunction, ^S(t),overtime
forasinglesampleofindividuals.
Nowwewanttocompare thesurvivalestimates betweentwo
groups.
Example:Timetoremissionofleukemiapatients
103
Howcanweformabasisforcomparison?
Ataspeci¯cpointintime,wecouldseewhether thecon¯-
denceintervalsforthesurvivalcurvesoverlap.
However,thecon¯dence intervalswehavebeencalculating
are\pointwise")theycorrespondtoacon¯dence inter-
valfor^S(t¤)atasinglepointintime,t¤.
Inotherwords,wecan'tsaythatthetruesurvivalfunction
S(t)iscontainedbetweenthepointwisecon¯dence intervals
with95%probabilit y.
(Aside: ifyou'reinterested, theissueofcon¯dencebands
fortheestimated survivalfunction arediscussed inSection
4.4ofKleinandMoeschberger)
104
Lookingatwhether thecon¯dence intervalsfor^S(t¤)overlap
betweenthe6MPandplacebogroupswouldonlyfocuson
comparing thetwotreatmen tgroupsatasinglepointin
time,t¤.Wewantanoverallcomparison.
Shouldwebaseouroverallcomparisonof^S(t)on:
²thefurthestdistance betweenthetwocurves?
²themediansurvivalforeachgroup?
²theaveragehazard? (forexponentialdistributions, this
wouldbelikecomparing themeaneventtimes)
²addingupthedi®erence betweenthetwosurvivalesti-
matesovertime?
X
j·^S(tjA)¡^S(tjB)¸
²aweightedsumofdi®erences, wheretheweightsre°ect
thenumberatriskateachtime?
²arank-based test?i.e.,wecouldrankalloftheevent
times,andthenseewhether thesumofranksforone
groupwaslessthantheother.
105
Nonparametric comparisonsofgroups
Alloftheseareprettyreasonable options,andwe'llseethat
therehavebeenseveralproposalsforhowtocompare the
survivaloftwogroups.Forthemoment,wearestickingto
nonparametric comparisons.
Whynonparametric?
²fairlyrobust
²e±cientrelativetoparametrictests
²oftensimpleandintuitive
Beforecontinuingthedescription ofthetwo-sample compar-
ison,I'mgoingtotrytoputthisinageneralframeworkto
giveaperspectiveofwherewe'reheadinginthisclass.
106
GeneralFrameworkforSurvivalAnalysis
Weobserve(Xi;±i;Zi)forindividuali,where
²Xiisacensored failuretimerandomvariable
²±iisthefailure/censoring indicator
²Zirepresentsasetofcovariates
NotethatZimightbeascalar(asinglecovariate,saytreat-
mentorgender)ormaybea(p£1)vector(represen ting
severaldi®erentcovariates).
Thesecovariatesmightbe:
²continuous
²discrete
²time-varying(morelater)
IfZiisascalarandisbinary,thenwearecomparing the
survivaloftwogroups,likeintheleukemiaexample.
Moregenerally though, itisusefultobuildamodelthat
characterizes therelationship betweensurvivalandallofthe
covariatesofinterest.
107
We'llproceedasfollows:
²Twogroupcomparisons
²Multigroup andstrati¯ed comparisons -strati¯ed logrank
²Failuretimeregression models
{Coxproportionalhazardsmodel
{Accelerated failuretimemodel
108
Twosampletests
²Mantel-Haenszel logranktest
²Peto&Peto'sversionofthelogranktest
²Gehan's Generalized Wilcoxon
²Peto&Peto'sandPrentice'sgeneralized Wilcoxon
²Tarone-WareandFleming-Harrington classes
²Cox'sF-test(non-parametric version)
References:
Hosmer&Lemesho wSection2.4
Collett Section2.5
Klein&Moeschberger Section7.3
Kleinbaum Chapter 2
Lee Chapter 5
109
Mantel-HaenszelLogranktest
Thelogranktestisthemostwellknownandwidelyused.
Italsohasanintuitiveappeal,building onstandard meth-
odsforbinarydata.(Laterwewillseethatitcanalsobe
obtained asthescoretestfromapartiallikelihoodfromthe
CoxProportionalHazards model.)
Firstconsider thefollowing(2£2)tableclassifying those
withandwithouttheeventofinterestinatwogroupsetting:
Event
Group Yes No Total
0d0n0¡d0n0
1d1n1¡d1n1
Totaldn¡dn
110
Ifthemargins ofthistableareconsidered ¯xed,thend0
followsa ? distribution. Under
thenullhypothesisofnoassociationbetweentheeventand
group,itfollowsthat
E(d0)=n0d
n
Var(d0)=n0n1d(n¡d)
n2(n¡1)
Therefore, underH0:
Â2
MH=[d0¡n0d=n]2
n0n1d(n¡d)
n2(n¡1)»Â2
1
ThisistheMantel-Haenszel statistic andisapproximately
equivalenttothePearsonÂ2testforequalityofthetwo
groupsgivenby:
Â2
p=X(o¡e)2
e
Note:recallthatthePearsonÂ2testwasderivedforthe
casewhereonlytherowmargins were¯xed,andthusthe
varianceabovewasreplaced by:
Var(d0¡n0(d0+d1)
n)=n0n1d(n¡d)
n3
111
Example: Toxicityinaclinicaltrialwithtwotreatmen ts
Toxicity
Group YesNo Total
0 842 50
1 248 50
Total 10 90 100
Â2
p=4:00(p=0:046)
Â2
MH=3:96(p=0:047)
112
NowsupposewehaveK(2£2)tables,allindependent,and
wewanttotestforacommon groupe®ect.TheCochran-
Mantel-Haenszel testforacommon oddsrationotequalto
1canbewrittenas:
Â2
CMH=[PK
j=1(d0j¡n0j¤dj=nj)]2
PKj=1n1jn0jdj(nj¡dj)=[n2j(nj¡1)]
wherethesubscriptjreferstothej-thtable:
Event
Group Yes No Total
0d0jn0j¡d0jn0j
1d1jn1j¡d1jn1j
Totaldjnj¡djnj
Thisstatistic isdistributed approximately asÂ2
1.
113
Howdoesthisapplyinsurvivalanalysis?
Supposeweobserve
Group1:(X11;±11):::(X1n1;±1n1)
Group0:(X01;±01):::(X0n0;±0n0)
Wecouldjustcountthenumbersoffailures: eg.,d1=
PK
j=1±1j
Example:Leukemiadata,justcountingupthenumber
ofremissions ineachtreatmen tgroup.
Fail
Group YesNo Total
021 0 21
1 912 21
Total 30 12 42
Â2
p=16:8(p=0:001)
Â2
MH=16:4(p=0:001)
But,thisdoesn'taccountforthetimeatrisk.
Conceptually ,wewouldliketocompare theKMsurvival
curves.Let'sputthecomponentsside-by-sideandcompare.
114
Cox&OakesTable1.1Leukemiaexample
Ordered Group0Group1
DeathTimesdjcjrjdjcjrj
120210021
220190021
310170021
420160021
520140021
600123121
700121017
840120016
90080116
100081115
112080113
122060012
130041012
151040011
160031011
171030110
19002019
20002018
22102107
23101106
25000015
NotethatIwrotedownthenumberatriskforGroup1fortimes
1-5eventhoughtherewerenoeventsorcensoringsatthosetimes.
115
LogrankTest:FormalDe¯nition
Thelogranktestisobtained byconstructing a(2£2)ta-
bleateachdistinctdeathtime,andcomparing thedeath
ratesbetweenthetwogroups,conditional onthenumberat
riskinthegroups.Thetablesarethencombinedusingthe
Cochran-Man tel-Haenszel test.
Note:ThelogrankissometimescalledtheCox-Manteltest.
Lett1;:::;tKrepresenttheKordered, distinctdeathtimes.
Atthej-thdeathtime,wehavethefollowingtable:
Die/Fail
Group Yes No Total
0d0jr0j¡d0jr0j
1d1jr1j¡d1jr1j
Totaldjrj¡djrj
whered0jandd1jarethenumberofdeathsingroup0and
1,respectivelyatthej-thdeathtime,andr0jandr1jare
thenumberatriskatthattime,ingroups0and1.
116
Thelogranktestis:
Â2
logrank=[PK
j=1(d0j¡r0j¤dj=rj)]2
PKj=1r1jr0jdj(rj¡dj)
[r2j(rj¡1)]
Assuming thetablesareallindependent,thenthisstatistic
willhaveanapproximateÂ2distribution with1df.
Basedonthemotivationforthelogranktest,
whichofthesurvival-relatedquantitiesarewe
comparingateachtimepoint?
²PK
j=1wj·^S1(tj)¡^S2(tj)¸
?
²PK
j=1wj·^¸1(tj)¡^¸2(tj)¸
?
²PK
j=1wj·^¤1(tj)¡^¤2(tj)¸
?
117
Firstseveraltablesofleukemiadata
CMHanalysis ofleukemia data
TABLE1OFTRTMTBYREMISS TABLE3OFTRTMTBYREMISS
CONTROLLING FORFAILTIME=1 CONTROLLING FORFAILTIME=3
TRTMT REMISS TRTMT REMISS
Frequency| Frequency|
Expected | 0| 1|Total Expected | 0| 1|Total
---------+--------+--------+ ---------+--------+--------+
0|19|2|21 0|16|1|17
|20|1| |16.553|0.4474|
---------+--------+--------+ ---------+--------+--------+
1|21|0|21 1|21|0|21
|20|1| |20.447|0.5526|
---------+--------+--------+ ---------+--------+--------+
Total 40 2 42 Total 37 1 38
TABLE2OFTRTMTBYREMISS TABLE4OFTRTMTBYREMISS
CONTROLLING FORFAILTIME=2 CONTROLLING FORFAILTIME=4
TRTMT REMISS TRTMT REMISS
Frequency| Frequency|
Expected | 0| 1|Total Expected | 0| 1|Total
---------+--------+--------+ ---------+--------+--------+
0|17|2|19 0|14|2|16
|18.05|0.95| |15.135|0.8649|
---------+--------+--------+ ---------+--------+--------+
1|21|0|21 1|21|0|21
|19.95|1.05| |19.865|1.1351|
---------+--------+--------+ ---------+--------+--------+
Total 38 2 40 Total 35 2 37
118
CMHstatistic=logrankstatistic
SUMMARY STATISTICS FORTRTMTBYREMISS
CONTROLLING FORFAILTIME
Cochran-Mantel-Haenszel Statistics (BasedonTableScores)
Statistic Alternative Hypothesis DF Value Prob
-----------------------------------------------------------------
1 Nonzero Correlation 116.793 0.001
2 RowMeanScoresDiffer 116.793 0.001
3 General Association 116.793 0.001<===LOGRANK
TEST
Note:Although CMHworkstogetthecorrectlogranktest,
itwouldrequireinputting thedjandrjateachtimeofdeath
foreachtreatmen tgroup.There'saneasierwaytogetthe
teststatistic, whichI'llshowyoushortly.
119
Calculatinglogrankstatisticbyhand
LeukemiaExample:
OrderedGroup0Combined
DeathTimesd0jr0jdjrjejoj¡ejvj
12212421.001.000.488
22192400.951.05
31171380.450.55
42162370.861.14
5214235
6012333
7012129
8412428
1008123
1128221
1226218
1304116
1514115
1603114
1713113
221229
231127
Sum 10.2516.257
oj=d0j
ej=djr0j=rj
vj=r1jr0jdj(rj¡dj)=[r2
j(rj¡1)]
Â2
logrank=(10:251)2
6:257=16:793
120
Notesaboutlogranktest:
²Thelogrankstatistic dependsonranksofeventtimes
only
²Iftherearenotieddeaths,thenthelogrankhastheform:
[PK
j=1(d0j¡r0j
rj)]2
PKj=1r1jr0j=r2j
²Numerator canbeinterpreted asP(o¡e)where\o"is
theobservednumberofdeathsingroup0,and\e"is
theexpectednumber,giventheriskset.Theexpected
numberequals#deaths£proportioningroup0atrisk.
²The(o¡e)termsinthenumerator canbewrittenas
r0jr1j
rj(^¸1j¡^¸0j)
²Itdoesnotmatterwhichgroupyouchoosetosumover.
Toseethis,notethatifwesummedup(o-e)overthedeath
timesforthe6MPgroupwewouldget-10.251,andthesumof
thevariancesisthesame.Sowhenwesquarethenumerator,
theteststatisticisthesame.
121
Analogous totheCMHtestforaseriesoftablesatdi®erent
levelsofaconfounder, thelogranktestismostpowerfulwhen
\oddsratios"areconstantovertimeintervals.Thatis,itis
mostpowerfulforproportionalhazards .
Checkingtheassumptionofproportionalhazards:
²checktoseeiftheestimated survivalcurvescross-if
theydo,thenthisisevidence thatthehazardsarenot
proportional
²moreformaltest:anyideas?
Whatshouldbedoneifthehazardsarenot
proportional?
²Ifthedi®erence betweenhazardshasaconsisten tsign,
thelogranktestusuallydoeswell.
²Othertestsareavailablethataremorepowerfulagainst
di®erentalternativ es.
122
GettingthelogrankstatisticusingStata:
Afterdeclaringdataassurvivaltypedatausing
the\stset"command,issuethe\ststest"com-
mand
.stsetremissstatus
datasetname:leukem
id:-- (meaning eachrecordauniquesubject)
entrytime:-- (meaning allentered attime0)
exittime:remiss
failure/censor: status
.stslist,by(trt)
Beg. NetSurvivor Std.
Time Total FailLostFunction Error [95%Conf.Int.]
------------------------------------------------------------------- ---
trt=0
1 21 200.9048 0.0641 0.6700 0.9753
2 19 200.8095 0.0857 0.5689 0.9239
3 17 100.7619 0.0929 0.5194 0.8933
4 16 200.6667 0.1029 0.4254 0.8250
.
.(etc)
.ststesttrt
Log-rank testforequality ofsurvivor functions
------------------------------------------------
|Events
trt|observed expected
------+-------------------------
0| 21 10.75
1| 9 19.25
------+-------------------------
Total| 30 30.00
chi2(1) =16.79
Pr>chi2 =0.0000
123
GettingthelogrankstatisticusingSAS
²StillusePROCLIFETEST
²Add\STRATA"command, withtreatmen tvariable
²Givesthechi-square test(2-sided), butalsogivesyou
thetermsyouneedtocalculate the1-sidedtest;thisis
usefulifwewanttoknowwhichofthetwogroupshas
thehigherestimated hazardovertime.
²TheSTRATAcommand alsogivestheGehan-Wilco xon
test(whichwewilltalkaboutnext)
Title'CoxandOakesexample';
dataleukemia;
inputweeksremisstrtmt;
cards;
601
611
611
611 /*datafor6MPgroup*/
711
901
etc
110
110 /*dataforplacebo group*/
210
210
etc
;
proclifetest data=leukemia;
timeweeks*remiss(0);
stratatrtmt;
title'Logrank testforleukemia data';
run;
124
Outputfromleukemiaexample:
Logrank testforleukemia data
Summary oftheNumberofCensored andUncensored Values
TRTMT Total Failed Censored %Censored
6-MP 21 9 1257.1429
Control 21 21 00.0000
Total 42 30 1228.5714
Testing Homogeneity ofSurvival CurvesoverStrata
TimeVariable FAILTIME
RankStatistics
TRTMT Log-Rank Wilcoxon
6-MP -10.251 -271.00
Control 10.251 271.00
Covariance MatrixfortheLog-Rank Statistics
TRTMT 6-MP Control
6-MP 6.25696 -6.25696
Control -6.25696 6.25696
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 16.7929 10.0001 <==Here'stheonewewant!!
Wilcoxon 13.4579 10.0002
-2Log(LR) 16.4852 10.0001
125
GettingthelogrankstatisticusingSplus:
Insteadofthe\surv.¯t"command,usethe
\surv.di®"commandwitha\group"(treatment)
variable.
Mantel-Haenszel logrank:
>logrank<-surv.diff(weeks,remiss,trtmt)
>logrank
NObserved Expected (O-E)^2/E
021 2110.75 9.775
121 919.25 5.458
Chisq=16.8on1degrees offreedom, p=4.169e-05
126
Generalization oflogranktest
=)Linearranktests
Thelogrankandothertestscanbederivedbyassigning
scorestotheranksofthedeathtimes,andaremembersof
ageneralclassoflinearranktests(formoredetail,see
Lee,ch5)
First,de¯ne
^¤(t)=X
j:tj<tdj
rj
wheredjandrjarethenumberofdeathsandthenumber
atrisk,respectivelyatthej-thordereddeathtime.
Thenassignthesescores(suggested byPetoandPeto):
Event Score
Deathattjwj=1¡^¤(tj)
Censoring attjwj=¡^¤(tj)
Tocalculate thelogranktest,simplysumupthescoresfor
group0.
127
Example Group0:15,18,19,19,20
Group1:16+,18+,20+,23,24+
Calculationoflogrankasalinearrankstatistic
Ordered DataGroupdjrj^¤(tj)scorewj
15 01100.100 0.900
16+1090.100-0.100
18 0180.225 0.775
18+1070.225-0.225
19 0260.558 0.442
20 0140.808 0.192
20+1030.808-0.808
23 1121.308-0.308
24+1011.308-1.308
ThelogrankstatisticSissumofscoresforgroup0:
S=0:900+0:775+0:442+0:442+0:192=2:75
Thevarianceis:
Var(S)=n0n1Pn
j=1w2
j
n(n¡1)
Inthiscase,Var(S)=1:210,so
Z=2:75p
1:210=2:50=)Â2
logrank=(2:50)2=6:25
128
Whyisthisformofthelogrankequivalent?
Thelogrankstatistic SisequivalenttoP(o¡e)overthe
distinctdeathtimes,where\o"istheobservednumberof
deathsingroup0,and\e"istheexpectednumber,given
therisksets.
Atdeaths: weightsare1¡^¤
Atcensorings: weightsare¡^¤
Sowearesumming up\1's"fordeaths(togetd0j),andsub-
tracting¡^¤atbothdeathsandcensorings. Thisamountsto
subtracting dj=rjateachdeathorcensoring timeingroup
0,atorafterthej-thdeath.Sincethereareatotalofr0jof
these,wegete=r0j¤dj=rj.
Whyisitcalledthelogrank test?
SinceS(t)=exp(¡¤(t)),analternativ eestimator ofS(t)
is:
^S(t)=exp(¡^¤(t))=exp(¡X
j:tj<tdj
rj)
So,wecanthinkof^¤(t)=¡log(^S(t))asyieldingthe\log-
survival"scoresusedtocalculate thestatistic.
129
ComparingtheCMH-typeLogrankand
\LinearRank"logrank
A.CMH-typeLogrank:
Wemotivatedthelogranktestthrough theCMHstatistic
fortestingHo:OR=1overKtables,whereKisthe
numberofdistinctdeathtimes.Thisturnedouttobewhat
wegetwhenweusethelogrank(default) optioninStataor
the\strata"statemen tinSAS.
B.LinearRanklogrank:
Thelinearrankversionofthelogranktestisbasedonadding
up\scores" foroneofthetwotreatmen tgroups.Thepar-
ticularscoresthatgaveusthesamelogrankstatistic were
basedontheNelson-Aalen estimator, i.e.,^¤=P^¸(tj).This
iswhatyougetwhenyouusethe\test"statemen tinSAS.
Herearesomecomparisons, withanewexample toshow
whenthetwotypesoflogrankstatistics willbeequal.
130
First,let'sgobacktoourexample fromChapter 5ofLee:
Example Group0:15,18,19,19,20
Group1:16+,18+,20+,23,24+
A.TheCMH-typelogrankstatistic:
(usingthestratastatement)
RankStatistics
TRTMT Log-Rank Wilcoxon
Control 2.7500 18.000
Treated -2.7500 -18.000
Covariance MatrixfortheLog-Rank Statistics
TRTMT Control Treated
Control 1.08750 -1.08750
Treated -1.08750 1.08750
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 6.9540 10.0084
Wilcoxon 5.5479 10.0185
-2Log(LR) 3.3444 10.0674
131
Thisisexactlythesamechi-square testthatyouwouldget
ifyoucalculated thenumerator ofthelogrankasP(oj¡ej)
andthevarianceasvj=r1jr0jdj(rj¡dj)=[r2
j(rj¡1)]
OrderedGroup0Combined
DeathTimesd0jr0jdjrjejoj¡ejvj
15151100.500.500.2500
1814180.500.500.2500
1923261.001.000.4000
2011240.250.750.1870
2300120.000.000.0000
Sum 2.751.0875
Â2
logrank=(2:75)2
1:0875=6:954
132
B.The\linearrank"logrankstatistic:
(usingtheteststatement)
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
GROUP 2.7500 1.0897 6.3684 0.0116
Covariance MatrixfortheLOGRANKStatistics
Variable TRTMT
TRTMT 1.18750
Thisisactually veryclosetowhatwewouldgetifweuse
theNelson-Aalen based\scores":
Calculationoflogrankasalinearrankstatistic
Ordered DataGroupdjrj^¤(tj)scorewj
15 01100.100 0.900
16+1090.100-0.100
18 0180.225 0.775
18+1070.225-0.225
19 0260.558 0.442
20 0140.808 0.192
20+1030.808-0.808
23 1121.308-0.308
24+1111.308-1.308
Sum(grp 0) 2.750
133
Notethatthenumerator istheexactsamenumber(2.75)
inbothversionsofthelogranktest.Thedi®erence inthe
denominator isduetothewaythattiesarehandled.
CMH-typevariance:
var=Xr1jr0jdj(rj¡dj)
r2j(rj¡1)
=Xr1jr0j
rj(rj¡1)dj(rj¡dj)
rj
Linearranktypevariance:
var=n0n1Pn
j=1w2
j
n(n¡1)
134
Nowconsideranexamplewheretherearenotied
deathtimes
ExampleIGroup0:15,18,19,21,22
Group1:16+,17+,20+,23,24+
A.TheCMH-typelogrankstatistic:
(usingthestratastatement)
RankStatistics
TRTMT Log-Rank Wilcoxon
Control 2.5952 15.000
Treated -2.5952 -15.000
Covariance MatrixfortheLog-Rank Statistics
TRTMT Control Treated
Control 1.21712 -1.21712
Treated -1.21712 1.21712
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 5.5338 10.0187
Wilcoxon 4.3269 10.0375
-2Log(LR) 3.1202 10.0773
135
B.The\linearrank"logrankstatistic:
(usingtheteststatement)
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
TRTMT 2.5952 1.1032 5.5338 0.0187
Covariance MatrixfortheLOGRANKStatistics
Variable TRTMT
TRTMT 1.21712
Notethatthistime,thevariances ofthetwologrankstatis-
ticsareexactlythesame,equalto1.217.
Iftherearenotiedeventtimes,thenthe
twoversionsofthetestwillyieldidenti-
calresults.Themoretieswehave,the
moreitmatterswhichversionweuse.
136
Gehan'sGeneralizedWilcoxonTest
First,let'sreviewtheWilcoxontestforuncensored data:
Denoteobservationsfromtwosamplesby:
(X1;X2;:::;Xn)and(Y1;Y2;:::;Ym)
Orderthecombinedsampleandde¯ne:
Z(1)<Z(2)<¢¢¢<Z(m+n)
Ri1=rankofXi
R1=m+nX
i=1Ri1
RejectH0ifR1istoobigortoosmall,according to
R1¡E(R1)
r
Var(R1)»N(0;1)
where
E(R1)=m(m+n+1)
2
Var(R1)=mn(m+n+1)
12
137
TheMann-Whitney formoftheWilcoxonisde¯nedas:
U(Xi;Yj)=Uij=8
>>>>><
>>>>>:+1ifXi>Yj
0ifXi=Yj
¡1ifXi<Yj
and
U=nX
i=1mX
j=1Uij:
Thereisasimplecorrespondence betweenUandR1:
R1=m(m+n+1)=2+U=2
soU=2R1¡m(m+n+1)
Therefore,
E(U)=0
Var(U)=mn(m+n+1)=3
138
ExtendingWilcoxontocensoreddata
TheMann-Whitney formleadstoageneralization forcen-
soreddata.De¯ne
U(Xi;Yj)=Uij=8
>>>>><
>>>>>:+1ifxi>yjorx+
i¸yj
0ifxi=yiorlowervaluecensored
¡1ifxi<yjorxi·y+
j
Thende¯ne
W=nX
i=1mX
j=1Uij
Thus,thereisacontribution toWforeverycomparison
wherebothobservationsarefailures(exceptforties),or
whereacensored observationisgreaterthanorequaltoa
failure.
Lookingatallpossiblepairsofindividuals betweenthetwo
treatmen tgroupsmakesthisanightmaretocompute by
hand!
139
Gehanfoundaneasierwaytocompute theabove.First,
poolthesampleof(n+m)observationsintoasinglegroup,
thencompare eachindividual withtheremainingn+m¡1:
Forcomparing thei-thindividual withthej-th,de¯ne
Uij=8
>>>>><
>>>>>:+1ifti>tjort+
i¸tj
¡1ifti<tjorti·t+
j
0 otherwise
Then
Ui=m+nX
j=1Uij
Thus,forthei-thindividual, Uiisthenumberofobserva-
tionswhicharede¯nitely lessthantiminusthenumberof
observationsthatarede¯nitely greaterthanti.Weassume
censorings occurafterdeaths,sothatifti=18+andtj=18,
thenweadd1toUi.
TheGehanstatistic isde¯nedas
U=m+nX
i=1Ui1fiingroup0g
=W
Uhasmean0andvariance
var(U)=mn
(m+n)(m+n¡1)m+nX
i=1U2
i
140
Example fromLee:
Group0:15,18,19,19,20
Group1:16+,18+,20+,23,24+
TimeGroupUiU2
i
150-9 81
16+1 1 1
180-6 36
18+1 2 4
190-2 4
190-2 4
200 1 1
20+1 525
231 416
24+1 636
SUM -18 208
U=¡18
Var(U)=(5)(5)(208)
(10)(9)
=57:78
andÂ2=(¡18)2=57:78=5:61
141
SAScode:
dataleedata;
infile'lee.dat';
inputtimecensgroup;
proclifetest data=leedata;
timetime*cens(0);
stratagroup;
run;
SASOUTPUT: GehansWilcoxontest
RankStatistics
TRTMT Log-Rank Wilcoxon
Control 2.7500 18.000
Treated -2.7500 -18.000
Covariance MatrixfortheWilcoxon Statistics
TRTMT Control Treated
Control 58.4000 -58.4000
Treated -58.4000 58.4000
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 6.9540 10.0084
Wilcoxon 5.5479 10.0185**thisisGehan's test
-2Log(LR) 3.3444 10.0674
142
NotesaboutSASWilcoxonTest:
SAScalculates theWilcoxonas¡UinsteadofU,probably
sothatthesignoftheteststatistic isconsisten twiththe
logrank.
SASgetssomething slightlydi®erentforthevariance, and
thisdoesnotseemtodependonwhether thereareties.
Forexample, thehypothetical datasetonp.6without ties
yieldsU=¡15andPU2
i=182,so
Var(U)=(5)(5)(182)
(10)(9)=50:56andÂ2=(¡15)2
50:56=4:45
whileSASgivesthefollowing:
RankStatistics
TRTMT Log-Rank Wilcoxon
Control 2.5952 15.000
Treated -2.5952 -15.000
Covariance MatrixfortheWilcoxon Statistics
TRTMT Control Treated
Control 52.0000 -52.0000
Treated -52.0000 52.0000
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 5.5338 10.0187
Wilcoxon 4.3269 10.0375
-2Log(LR) 3.1202 10.0773
143
ObtainingtheWilcoxontestusingStata
Usetheststeststatement,withtheappropriate option
ststest varlist [ifexp][inrange]
[,[logrank|wilcoxon|cox] strata(varlist) detail
mat(matname1 matname2) notitle noshow]
logrank, wilcoxon, andcoxspecify whichtestofequality isdesired.
logrank isthedefault, andcoxyieldsalikelihood ratiotest
underacoxmodel.
Example:(leukemiadata)
.stsetremissstatus
.ststesttrt,wilcoxon
Wilcoxon (Breslow) testforequality ofsurvivor functions
----------------------------------------------------------
|Events Sumof
trt|observed expected ranks
------+--------------------------------------
0| 21 10.75 271
1| 9 19.25 -271
------+--------------------------------------
Total| 30 30.00 0
chi2(1) =13.46
Pr>chi2 =0.0002
144
GeneralizedWilcoxon
(Peto&Peto,Prentice)
Assignthefollowingscores:
Foradeathatt: ^S(t+)+^S(t¡)¡1
Foracensoring att:^S(t+)¡1
Theteststatistic isP(scores)forgroup0.
TimeGroupdjrj^S(t+)scorewj
1501100.900 0.900
16+1090.900 -0.100
180180.788 0.688
18+1070.788 -0.212
190260.525 0.313
200140.394 -0.081
20+1030.394 -0.606
231120.197 -0.409
24+1010.197 -0.803
Xwj1fjingroup0g=0:900+0:688+2¤(0:313)+(¡0:081)
=2:13
Var(S)=n0n1Pn
j=1w2
j
n(n¡1)=0:765
soZ=2:13=0:765=2:433
145
TheTarone-Wareclassoftests:
Thisgeneralclassoftestsislikethelogranktest,butadds
weightswj.Thelogrank test,Wilcoxontest,andPeto-
PrenticeWilcoxonareincluded asspecialcases.
Â2
tw=[PK
j=1wj(d1j¡r1j¤dj=rj)]2
PK
l=1w2jr1jr0jdj(rj¡dj)
r2j(rj¡1)
Test Weightwj
Logrank wj=1
Gehan's Wilcoxonwj=rj
Peto/Pren tice wj=ncS(tj)
Fleming-Harrington wj=[^S(tj)]®
Tarone-Ware wj=prj
Note:theseweightswjarenotthesameasthescoreswjwe'vebeen
talkingaboutearlier,andtheyapplytotheCMH-typeformofthe
teststatisticratherthanP(scores)overasingletreatmentgroup.
146
Whichtestshouldweused?
CMH-typeorLinearRank?
Iftherearenotahighproportionofties,thenitdoesn't
reallymattersince:
²ThetwoWilcoxonsaresimilartoeachother
²Thetwologranktestsaresimilartoeachother
Note:personally,ItendtousetheCMH-typetest,whichyougetwiththestrata
statementinSASandtheteststatementinSTATA.
LogrankorWilcoxon?
²BothtestshavetherightTypeIpowerfortestingthe
nullhypothesisofequalsurvival,Ho:S1(t)=S2(t)
²Thechoiceofwhichtestmaytherefore dependonthe
alternativ ehypothesis,whichwilldrivethepowerofthe
test.
147
²TheWilcoxonissensitivetoearlydi®erences between
survival,whilethelogrankissensitivetolaterones.This
canbeseenbytherelativeweightstheyassigntothetest
statistic:
LOGRANK numerator=X
j(oj¡ej)
WILCOXONnumerator=X
jrj(oj¡ej)
²Thelogrankismostpowerfulundertheassumption of
proportional hazards, whichimpliesanalternativ ein
termsofthesurvivalfunctions ofHa:S1(t)=[S2(t)]®
²TheWilcoxonhashighpowerwhenthefailuretimes
arelognormally distributed, withequalvarianceinboth
groupsbutadi®erentmean.Itwillturnoutthatthisis
theassumption ofanaccelerated failuretimemodel.
²Bothtestswilllackpowerifthesurvivalcurves(orhaz-
ards)\cross". However,thatdoesnotnecessarily make
theminvalid!
148
ComparisonbetweenTESTandSTRATAinSAS
for2examples:
DatafromLee(n=10) :
fromSTRATA:
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 6.9540 10.0084
Wilcoxon 5.5479 10.0185**thisisGehan's test
-2Log(LR) 3.3444 10.0674
fromTEST:
Univariate Chi-Squares fortheWILCOXON Test
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
GROUP 1.8975 0.7508 6.3882 0.0115
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
GROUP 2.7500 1.0897 6.3684 0.0116
149
Previousexamplewithleukemiadata:
fromSTRATA:
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 16.7929 10.0001
Wilcoxon 13.4579 10.0002
-2Log(LR) 16.4852 10.0001
fromTEST:
Univariate Chi-Squares fortheWILCOXON Test
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
GROUP 6.6928 1.7874 14.0216 0.0002
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
GROUP 10.2505 2.5682 15.9305 0.0001
150
P-sampleandstrati¯edlogranktests
Wehavebeendiscussing twosampleproblems. Inpractice,
morecomplex settingsoftenarise:
²Therearemorethantwotreatmen tsorgroups,andthe
question ofinterestiswhetherthegroupsdi®erfromeach
other.
²Weareinterested inacomparison betweentwogroups,
butwewishtoadjustforanotherfactorthatmaycon-
foundtheanalysis
²Wewanttoadjustforlotsofcovariates.
Wewill¯rsttalkaboutcomparing thesurvivaldistributions
betweenmorethan2groups,andthenaboutadjusting for
othercovariates.
151
P-samplelogrank
SupposeweobservedatafromPdi®erentgroups,andthe
datafromgroupp(p=1;:::;P)are:
(Xp1;±p1):::(Xpnp;±pnp)
Wenowconstruct a(P£2)tableateachoftheKdistinct
deathtimes,andcompare thedeathratesbetweentheP
groups,conditional onthenumberatrisk.
Lett1;::::tKrepresenttheKordered, distinctdeathtimes.
Atthej-thdeathtime,wehavethefollowingtable:
Die/Fail
Group Yes No Total
1d1jr1l¡d1jr1j
. . . .
PdPjrPj¡dPjrPj
Totaldjrj¡djrj
wheredpjisthenumberofdeathsingrouppatthej-th
deathtime,andrpjisthenumberatriskatthattime.
ThetablesarethencombinedusingtheCMHapproach.
152
Ifwewerejustfocusingonthisonetable,thenaÂ2
(P¡1)test
statistic couldbeconstructed through acomparison of\o"s
and\e"s,likebefore.
Example: Toxicityinaclinicaltrialwith3treatmen ts
TABLEOFGROUPBYTOXICITY
GROUP TOXICITY
Frequency|
RowPct|no |yes |Total
---------+--------+--------+
1|42|8|50
|84.00|16.00|
---------+--------+--------+
2|48|2|50
|96.00|4.00|
---------+--------+--------+
3|38|12|50
|76.00|24.00|
---------+--------+--------+
Total 128 22 150
STATISTICS FORTABLEOFGROUPBYTOXICITY
Statistic DFValue Prob
------------------------------------------------------
Chi-Square 28.097 0.017
Likelihood RatioChi-Square 29.196 0.010
Mantel-Haenszel Chi-Square 11.270 0.260
Cochran-Mantel-Haenszel Statistics (BasedonTableScores)
Statistic Alternative Hypothesis DFValue Prob
----------------------------------------------------------
1 Nonzero Correlation 11.270 0.260
2 RowMeanScoresDiffer 28.043 0.018
3 General Association 28.043 0.018
153
FormalCalculations:
LetOj=(d1j;:::d(P¡1)j)Tbeavectoroftheobservednum-
beroffailuresingroups1to(P¡1),respectively,atthe
j-thdeathtime.Giventherisksetsr1j,...rPj,andthefact
thattherearedjdeaths,thenOjhasadistribution likea
multivariateversionoftheHypergeometric.Ojhasmean:
Ej=(djr1j
rj;:::;djr(P¡1)j
rj)T
andvariancecovariancematrix:
Vj=0
BBBBBBBB@v11jv12j:::v1(P¡1)j
v22j:::v2(P¡1)j
:::::::::
v(P¡1)(P¡1)j1
CCCCCCCCA
wherethe`-thdiagonal elementis:
v``j=r`j(rj¡r`j)dj(rj¡dj)=[r2
j(rj¡1)]
andthe`m-tho®-diagonal elementis:
v`mj=r`jrmjdj(rj¡dj)=[r2
j(rj¡1)]
154
TheresultingÂ2testforasingle(P£1)tablewouldhave
(P-1)degreesandisconstructed asfollows:
(Oj¡Ej)TV¡1
j(Oj¡Ej)
GeneralizingtoKtables
Analogous towhatwedidforthetwosamplelogrank, we
replacetheOj,EjandVjwiththesumsovertheKdistinct
deathtimes.Thatis,letO=Pk
j=1Oj,E=Pk
j=1Ej,and
V=Pk
j=1Vj.Then,theteststatistic is:
(O¡E)TV¡1(O¡E)
155
Example:
Timetakento¯nishatestwith3di®erentnoisedistractions.
Alltestswerestoppedafter12minutes.
NoiseLevel
Group Group Group
1 2 3
9.0 10.0 12.0
9.5 12.0 12+
9.0 12+12+
8.5 11.0 12+
10.0 12.0 12+
10.5 10.5 12+
156
Letsstartthecalculations...
Observeddatatable
OrderedGroup1Group2Group3Combined
Timesd1jr1jd2jr2jd3jr3jdjrj
8.5160606
9.0250606
9.5130606
10.0121606
10.5111506
11.0001406
12.0002316
Expectedtable
OrderedGroup1Group2Group3Combined
Timeso1je1jo2je2jo3je3jojej
8.5
9.0
9.5
10.0
10.5
11.0
12.0
DoingtheP-sampletestbyhandiscumbersome...
Luckily,moststatistical packageswilldoitforyou!
157
P-samplelogrankinStata
.stsgraph,by(group)
.ststestgroup,logrank
Log-rank testforequality ofsurvivor functions
------------------------------------------------
|Events
group|observed expected
------+-------------------------
1| 6 1.57
2| 5 4.53
3| 1 5.90
------+-------------------------
Total| 12 12.00
chi2(2) =20.38
Pr>chi2 =0.0000
.ststestgroup,wilcoxon
Wilcoxon (Breslow) testforequality ofsurvivor functions
----------------------------------------------------------
|Events Sumof
group|observed expected ranks
------+--------------------------------------
1| 6 1.57 68
2| 5 4.53 -5
3| 1 5.90 -63
------+--------------------------------------
Total| 12 12.00 0
chi2(2) =18.33
Pr>chi2 =0.0001
158
SASprogramforP-samplelogrank
Title'Testing withnoiseexample';
datanoise;
inputtesttime finishgroup;
cards;
9 11
9.5 11
9.0 11
8.5 11
10 11
10.5 11
10.0 12
12 12
12 02
11 12
12 12
10.5 12
12 13
12 03
12 03
12 03
12 03
12 03
;
proclifetest data=noise;
timetesttime*finish(0);
stratagroup;
run;
159
Testing Homogeneity ofSurvival CurvesoverStrata
TimeVariable TESTTIME
RankStatistics
GROUP Log-Rank Wilcoxon
1 4.4261 68.000
2 0.4703 -5.000
3 -4.8964 -63.000
Covariance MatrixfortheLog-Rank Statistics
GROUP 1 2 3
1 1.13644 -0.56191 -0.57454
2 -0.56191 2.52446 -1.96255
3 -0.57454 -1.96255 2.53709
Covariance MatrixfortheWilcoxon Statistics
GROUP 1 2 3
1 284.808 -141.495 -143.313
2 -141.495 466.502 -325.007
3 -143.313 -325.007 468.320
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 20.3844 20.0001
Wilcoxon 18.3265 20.0001
-2Log(LR) 5.5470 20.0624
160
Note:donotuseTestinSASPROCLIFETEST ifyou
wantaP-sample logrank. Testwillinterpretthegroup
variableasameasured covariate(i.e.,eitherordinalorcon-
tinuous).
Inotherwords,youwillgetatrendtestwithonly1degree
offreedom, ratherthanaP-sample testwith(p-1)df.
Forexample, here'swhatwegetifweusetheTESTstate-
mentonthenoiseexample:
proclifetest data=noise;
timetesttime*finish(0);
testgroup;
run;
SASOUTPUT:
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
GROUP 9.3224 2.2846 16.6503 0.0001
Covariance MatrixfortheLOGRANKStatistics
Variable GROUP
GROUP 5.21957
Forward Stepwise Sequence ofChi-Squares fortheLOGRANKTest
Pr>Chi-Square Pr>
Variable DFChi-Square Chi-Square Increment Increment
GROUP 116.6503 0.0001 16.6503 0.0001
161
TheStrati¯edLogrank
Sometimes, eventhoughweareinterestedincomparing two
groups(ormaybeP)groups,weknowthereareotherfactors
thatalsoa®ecttheoutcome. Itwouldbeusefultoadjustfor
theseotherfactorsinsomeway.
Example: Forthenursinghomedata,alogranktestcom-
paringlengthofstayforthoseunderandover85yearsof
agesuggests asigni¯can tdi®erence (p=0.03).
However,weknowthatgenderhasastrongassociationwith
lengthofstay,andalsoage.Hence,itwouldbeagoodidea
toSTRATIFYtheanalysisbygenderwhentryingtoassess
theagee®ect.
Astrati¯edlogrank allowsonetocompare groups,but
allowstheshapesofthehazardsofthedi®erentgroupsto
di®eracrossstrata.Itmakestheassumption thatthegroup
1vsgroup2hazardratioisconstantacrossstrata.
Inotherwords:¸1s(t)
¸2s(t)=µwhereµisconstantoverthestrata
(s=1;:::;S).
Thismethodofadjusting forothervariablesisnotas°exible
asthatbasedonamodellingapproach.
162
Generalsetupforthestrati¯edlogrank:
Supposewewanttoassesstheassociationbetweensurvival
andafactor(callthisX)thathastwodi®erentlevels.Sup-
posehowever,thatwewanttostratifybyasecondfactor,
thathasSdi®erentlevels.
First,dividethedataintoSseparate groups.Withingroup
s(s=1;:::;S),proceedasthoughyouwereconstructing
thelogranktoassesstheassociationbetweensurvivaland
thevariableX.Thatis,lett1s;:::;tKssrepresenttheKs
ordered, distinctdeathtimesinthes-thgroup.
Atthej-thdeathtimeingroups,wehavethefollowing
table:
Die/Fail
XYes No Total
1ds1jrs1j¡ds1jrs1j
2ds2jrs2j¡ds2jrs2j
Totaldsjrsj¡dsjrsj
163
LetOsbethesumofthe\o"sobtained byapplying the
logrankcalculations intheusualwaytothedatafromgroup
s.Similarly ,letEsbethesumofthe\e"s,andVsbethe
sumofthe\v"s.
Thestrati¯edlogrank is
Z=PS
s=1(Os¡Es)
rPSs=1(Vs)
164
Strati¯edlogrankusingStata:
.usenurshome
.genage1=0
.replace age1=1ifage>85
.ststestage1,strata(gender)
failure _d:cens
analysis time_t:los
Stratified log-rank testforequality ofsurvivor functions
-----------------------------------------------------------
|Events
age1|observed expected(*)
------+-------------------------
0| 795 764.36
1| 474 504.64
------+-------------------------
Total|1269 1269.00
(*)sumovercalculations withingender
chi2(1) = 3.22
Pr>chi2 =0.0728
165
Strati¯edlogrankusingSAS:
datapop1;
setpop;
age1=0;
ifage>85thenage1=1;
proclifetest data=pop1 outsurv=survres;
timestay*censor(1);
testage1;
stratagender;
RESULTS(justthelogrankpart....youcanalsodoastrati¯ed
Wilcoxon)
TheLIFETEST Procedure
RankTestsfortheAssociation ofLSTAYwithCovariates
PooledoverStrata
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
AGE1 29.1508 17.1941 2.8744 0.0900
Covariance MatrixfortheLOGRANKStatistics
Variable AGE1
AGE1 295.636
Forward Stepwise Sequence ofChi-Squares fortheLOGRANKTest
Pr> Chi-Square Pr>
Variable DFChi-Square Chi-Square Increment Increment
AGE1 1 2.8744 0.0900 2.8744 0.0900
166
ModelingofSurvivalData
Nowwewillexploretherelationship betweensurvivaland
explanatory variablesbymodeling.Inthisclass,weconsider
twobroadclassesofregression models:
²ProportionalHazards(PH)models
¸(t;Z)=¸0(t)ª(Z)
Mostcommonly ,wewritethesecondtermas:
ª(Z)=e¯Z
SupposeZ=1fortreatedsubjectsandZ=0forun-
treatedsubjects.Thenthismodelsaysthatthehazard
isincreased byafactorofe¯fortreatedsubjectsversus
untreatedsubjects(c¯mightbe<1).
Thisisanexample ofasemi-parametric model.
²AcceleratedFailureTime(AFT)models
log(T)=¹+¯Z+¾w
wherewisan\errordistribution". Typically,weplace
aparametric assumption onw:
{exponential,Weibull,Gamma
{lognormal
167
Covariates:
Ingeneral,Zisavectorofcovariatesofinterest.
Zmayinclude:
²continuousfactors(eg,age,bloodpressure),
²discretefactors(gender, maritalstatus),
²possibleinteractions (agebysexinteraction)
DiscreteCovariates:
Justasinstandard linearregression, ifwehaveadiscrete
covariateAwithalevels,thenwewillneedtoinclude(a¡1)
dummyvariables(U1;U2;:::;Ua)suchthatUj=1ifA=
j.Then
¸i(t)=¸0(t)exp(¯2U2+¯3U3+¢¢¢+¯aUa)
(Intheabovemodel,thesubgroup withA=1orU1=1is
thereference group.)
Interactions:
Twofactors,AandB,interactifthehazardofdeathde-
pendsonthecombination oflevelsofAandB.
Weusuallyfollowtheprinciple ofhierarchicalmodels,and
onlyincludeinteractions ifallofthecorrespondingmain
e®ectsarealsoincluded.
168
Theexample Ijustgavewasbasedonaproportionalhazards
model,butthedescription ofthetypesofcovariateswemight
wanttoincludeinourmodelappliestoboththeAFTand
PHmodel.
We'llstartoutbyfocusingontheCoxPHmodel,andad-
dresssomeofthefollowingquestions:
²Whatdoestheterm¸0(t)mean?
²What's\proportional" aboutthePHmodel?
²Howdoweestimate theparameters inthemodel?
²Howdoweinterprettheestimated values?
²Howcanweconstruct testsofwhether thecovariates
haveasigni¯can te®ectonthedistribution ofsurvival
times?
²Howdothesetestscompare tothelogranktestorthe
Wilcoxontest?
169
TheCoxProportionalHazardsmodel
¸(t;Z)=¸0(t)exp(¯Z)
Thisisthemostcommon modelusedforsurvivaldata.
Why?
²°exiblechoiceofcovariates
²fairlyeasyto¯t
²standard softwareexists
References: Collett,Chapter 3*
Lee,Chapter 10*
Hosmer&Lemesho w,Chapters 3-7
Allison,Chapter 5
CoxandOakes,Chapter 7
Kleinbaum,Chapter 3
KleinandMoeschberger,Chapters 8&9
Kalb°eisc handPrentice
Note:somebooks(likeCollettandH&L)useh(t;X)as
theirstandard notation forthehazardinsteadof¸(t;Z),and
H(t)forthecumulativehazardinsteadof¤(t).
170
Whydowecallitproportionalhazards?
Thinkofthe¯rstexample, whereZ=1fortreatedandZ=
0forcontrol.Thenifwethinkof¸1(t)asthehazardrate
forthetreatedgroup,and¸0(t)asthehazardforcontrol,
thenwecanwrite:
¸1(t)=¸(t;Z=1)=¸0(t)exp(¯Z)
=¸0(t)exp(¯)
Thisimpliesthattheratioofthetwohazardsisaconstant,
Á,whichdoesNOTdependontime,t.Inotherwords,the
hazardsofthetwogroupsremainproportionalovertime.
Á=¸1(t)
¸0(t)=e¯
Áisreferredtoasthehazardratio.
Whatistheinterpretationof¯here?
171
TheBaselineHazardFunction
Intheexample ofcomparing twotreatmen tgroups,¸0(t)is
thehazardrateforthecontrolgroup.
Ingeneral,¸0(t)iscalledthebaselinehazardfunction ,
andre°ectstheunderlying hazardforsubjectswithallco-
variatesZ1;:::;Zpequalto0(i.e.,the\reference group").
Thegeneralformis:
¸(t;Z)=¸0(t)exp(¯1Z1+¯2Z2+¢¢¢+¯pZp)
Sowhenwesubstitute alloftheZj'sequalto0,weget:
¸(t;Z=0)=¸0(t)exp(¯1¤0+¯2¤0+¢¢¢+¯p¤0)
=¸0(t)
Inthegeneralcase,wethinkofthei-thindividual havinga
setofcovariatesZi=(Z1i;Z2i;:::;Zpi),andwemodeltheir
hazardrateassomemultipleofthebaseline hazardrate:
¸i(t;Zi)=¸0(t)exp(¯1Z1i+¢¢¢+¯pZpi)
172
Thismeanswecanwritethelogofthehazardratioforthe
i-thindividual tothereference groupas:
log0
B@¸i(t)
¸0(t)1
CA=¯1Z1i+¯2Z2i+¢¢¢+¯pZpi
TheCoxProportionalHazardsmodelisa
linearmodelforthelogofthehazardratio
OneofthebiggestadvantagesoftheframeworkoftheCox
PHmodelisthatwecanestimate theparameters¯which
re°ectthee®ectsoftreatmen tandothercovariateswithout
havingtomakeanyassumptions abouttheformof¸0(t).
Inotherwords,wedon'thavetoassumethat¸0(t)follows
anexponentialmodel,oraWeibullmodel,oranyotherpar-
ticularparametric model.
That'swhatmakesthemodelsemi-parametric.
Questions:
1.Whydon'twejustmodelthehazardratio,
Á=¸i(t)=¸0(t),directlyasalinearfunctionofthe
covariatesZ?
2.Whydoesn'tthemodelhaveanintercept?
173
Howdoweestimatethemodelparameters?
ThebasicideaisthatunderPH,information about¯can
beobtained fromtherelativeorderings (i.e.,ranks)ofthe
survivaltimes,ratherthantheactualvalues.Why?
SupposeTfollowsaPHmodel:
¸(t;Z)=¸0(t)e¯Z
NowconsiderT¤=g(T),wheregisamonotonic increasing
function. WecanshowthatT¤alsofollowsthePHmodel,
withthesamemultiplier,e¯Z.
Therefore, whenweconsider likelihoodmethodsforestimat-
ingthemodelparameters, weonlyhavetoworryaboutthe
ranksofthesurvivaltimes.
174
LikelihoodEstimationforthePHModel
Kalb°eisc handPrenticederivealikelihoodinvolvingonly
¯andZ(not¸0(t))basedonthemarginal distribution of
theranksoftheobservedfailuretimes(intheabsence of
censoring).
Cox(1972)derivedthesamelikelihood,andgeneralized it
forcensoring, usingtheideaofapartiallikelihood
Supposeweobserve(Xi;±i;Zi)forindividuali,where
²Xiisacensored failuretimerandomvariable
²±iisthefailure/censoring indicator (1=fail,0=censor)
²Zirepresentsasetofcovariates
Thecovariatesmaybecontinuous,discrete, ortime-varying.
175
SupposethereareKdistinctfailure(ordeath)times,and
let¿1;::::¿KrepresenttheKordered, distinctdeathtimes.
Fornow,assumetherearenotieddeathtimes.
LetR(t)=fi:xi¸tgdenotethesetofindividuals who
are\atrisk"forfailureattimet.
Moreaboutrisksets:
²IwillrefertoR(¿j)astherisksetatthejthfailuretime
²IwillrefertoR(Xi)astherisksetatthefailuretimeof
individuali
²Therewillstillberjindividuals inR(¿j).
²rjisanumber,whileR(¿j)identi¯estheactualsubjects
atrisk
176
Whatisthepartiallikelihood?
Intuitively,itisaproductoverthesetofobserveddeath
timesoftheconditional probabilities ofseeingtheobserved
deaths,giventhesetofindividuals atriskatthosetimes.
Ateachdeathtime¿j,thecontribution tothelikelihoodis:
Lj(¯)=Pr(individual jfailsj1failurefromR(¿j))
=Pr(individual jfailsjatriskat¿j)
P
`2R(¿j)Pr(individual `failsjatriskat¿j)
=¸(¿j;Zj)
P
`2R(¿j)¸(¿j;Z`)
UnderthePHassumption, ¸(t;Z)=¸0(t)e¯Z,soweget:
Lpartial(¯)=KY
j=1¸0(¿j)e¯Zj
P
`2R(¿j)¸0(¿j)e¯Z`
=KY
j=1e¯Zj
P
`2R(¿j)e¯Z`
177
Anotherderivation:
Ingeneral,thelikelihoodcontributions forcensored datafall
intotwocategories:
²IndividualiscensoredatXi:
Li(¯)=S(Xi)=exp[¡ZXi
0¸i(u)du]
²IndividualfailsatXi:
Li(¯)=S(Xi)¸i(Xi)=¸i(Xi)exp[¡ZXi
0¸i(u)du]
Thus,everyonecontributesS(Xi)tothelikelihood,andonly
thosewhofailcontribute¸i(Xi).
Thismeanswegetatotallikelihoodof:
L(¯)=nY
i=1¸i(Xi)±iexp[¡ZXi
0¸i(u)du]
Theabovelikelihoodholdsforallcensored survivaldata,
withgeneralhazardfunction¸(t).Inotherwords,wehaven't
usedtheCoxPHassumption atallyet.
178
Now,let'smultiplyanddividebytheterm·P
j2R(Xi)¸i(Xi)¸±i:
L(¯)=nY
i=12
4¸i(Xi)
P
j2R(Xi)¸i(Xi)3
5±i2
64X
j2R(Xi)¸i(Xi)3
75±i
exp[¡ZXi
0¸i(u)du]
Cox(1972)arguedthatthe¯rstterminthisproductcon-
tainedalmostalloftheinformation about¯,whilethesec-
ondtwotermscontainedtheinformation about¸0(t),i.e.,
thebaseline hazard.
Ifwejustfocusonthe¯rstterm,thenundertheCoxPH
assumption:
L(¯)=nY
i=12
64¸i(Xi)
P
j2R(Xi)¸i(Xi)3
75±i
=nY
i=12
64¸0(Xi)exp(¯Zi)
P
j2R(Xi)¸0(Xi)exp(¯Zj)3
75±i
=nY
i=12
64exp(¯Zi)
P
j2R(Xi)exp(¯Zj)3
75±i
Thisisthepartiallikelihoodde¯nedbyCox.Notethatit
doesnotdependontheunderlying hazardfunction¸0(¢).
Coxrecommends treating thisasanordinary likelihoodfor
makinginferences about¯inthepresence ofthenuisance
parameter¸0(¢).
179
Asimpleexample:
individual Xi±iZi
1 914
2 805
3 617
4 1013
Nowlet'scompilethepiecesthatgointothepartiallikeli-
hoodcontributions ateachfailuretime:
ordered
failure Likelihoodcontribution
jtimeXiR(Xi)ij·
e¯Zi=P
j2R(Xi)e¯Zj¸±i
16f1,2,3,4g3e7¯=[e4¯+e5¯+e7¯+e3¯]
28f1,2,4g2 1
39f1,4g1e4¯=[e4¯+e3¯]
410f4g4 e3¯=e3¯=1
Thepartiallikelihoodwouldbetheproductofthesefour
terms.
180
Notesonthepartiallikelihood:
L(¯)=nY
j=12
664e¯Zj
P
`2R(Xj)e¯Z`3
775±j
=KY
j=1e¯Zj
P
`2R(¿j)e¯Z`
wheretheproductisovertheKdeath(orfailure)times.
²contributions onlyatthedeathtimes
²thepartiallikelihoodisNOTaproductofindependent
terms,butofconditional probabilities
²Thereareotherchoicesbesidesª(Z)=e¯Z,butthis
isthemostcommon andtheoneforwhichsoftwareis
generally available.
181
PartialLikelihoodinference
Inference canbeconducted bytreatingthepartiallikelihood
asthoughitsatis¯ed alltheregularlikelihoodproperties.
Thelog-partiallikelihoodis:
`(¯)=log2
664nY
j=1e¯Zj
P
`2R(¿j)e¯Z`3
775±j
=log2
664KY
j=1e¯Zj
P
`2R(¿j)e¯Z`3
775
=KX
j=12
664¯Zj¡log[X
`2R(¿j)e¯Z`]3
775
=KX
j=1lj(¯)
whereljisthelog-partial likelihoodcontribution atthej-th
ordereddeathtime.
182
Supposethereisonlyonecovariate(¯isone-dimensional):
Thepartiallikelihoodscoreequations are:
U(¯)=@
@¯`(¯)=nX
j=1±j2
664Zj¡P
`2R(¿j)Z`e¯Z`
P
`2R(¿j)e¯Z`3
775
WecanexpressU(¯)intuitivelyasasumof\observed"mi-
nus\expected"values:
U(¯)=@
@¯`(¯)=nX
j=1±j(Zj¡¹Zj)
where¹Zjisthe\weightedaverage"ofthecovariateZover
alltheindividuals intherisksetattime¿j.Notethat¯is
involvedthrough theterm¹Zj.
Themaximumpartiallikelihoodestimators canbefoundby
solvingU(¯)=0.
183
Analogous tostandard likelihoodtheory,itcanbeshown
(though noteasily)that
(c¯¡¯)
se(^¯)»N(0;1)
Thevarianceof^¯canbeobtained byinvertingthesecond
derivativeofthepartiallikelihood,
var(^¯)»2
64¡@2
@¯2`(¯)3
75¡1
Fromtheaboveexpression forU(¯),wehave:
@2
@¯2`(¯)=nX
j=1±j2
664¡P
`2R(¿j)(Zj¡¹Zj)2e¯Z`
P
`2R(¿j)e¯Z`3
775
Note:
Thetruevarianceof^¯endsupbeingafunctionof¯,which
isunknown.Wecalculate the\observed"information by
substituting inourpartiallikelihoodestimateof¯intothe
aboveformulaforthevariance
184
SimpleExamplefor2-groupcomparison:(noties)
Group0:4+;7;8+;9;10+=)Zi=0
Group1:3;5;5+;6;8+=)Zi=1
orderedfailureNumberatriskLikelihoodcontribution
jtimeXiGroup0Group1h
e¯Zi=P
j2R(Xi)e¯Zji±i
13 55 e¯=[5+5e¯]
25 44 e¯=[4+4e¯]
36 42 e¯=[4+2e¯]
47 41 e¯=[4+1e¯]
59 20 e0=[2+0]=1=2
Again,wetaketheproductoverthelikelihoodcontributions,
thenmaximize togetthepartialMLEfor¯.
Whatdoes¯representinthiscase?
185
Notes
²The\observed"information matrixisgenerally usedbe-
causeinpractice, people¯ndithasbetterproperties.
Also,the\expected"isveryhardtocalculate.
²Thereisaniceanalogy withthescoreandinforma-
tionmatrices frommorestandard regression problems,
exceptthatherewearesumming overobserveddeath
times,ratherthanindividuals.
²NewtonRaphson isusedbymanyofthecomputer pack-
agestosolvethepartiallikelihoodequations.
186
FittingCoxPHmodelwithStata
Usesthe\stcox"command.
First,trytyping\helpstcox"
------------------------------------------------------------------- ---
helpforstcox
------------------------------------------------------------------- ---
Estimate Coxproportional hazards model
---------------------------------------
stcox[varlist] [ifexp][inrange]
[,nohrstrata(varnames) robustcluster(varname) noadjust
mgale(newvar) esr(newvars)
schoenfeld(newvar) scaledsch(newvar)
basehazard(newvar) basechazard(newvar) basesurv(newvar)
{breslow |efron|exactm|exactp} cmdestimate noshow
offsetlevel(#) maximize-options ]
stphtest [,kmlogranktime(varname) plot(varname) detail
graph-options ksm-options]
stcoxisforusewithsurvival-time data;seehelpst.Youmust
havestsetyourdatabeforeusingthiscommand; seehelpstset.
Description
-----------
stcoxestimates maximum-likelihood proportional hazards modelsonstdata.
Options (manymore!)
-------
nohrreports theestimated coefficients ratherthanhazardratios; i.e.,
bratherthanexp(b). Standard errorsandconfidence intervals are
similarly transformed. Thisoptionaffects howresults aredisplayed,
nothowtheyareestimated.
187
Ex.LeukemiaData
.stcoxtrt
Iteration 0:loglikelihood =-93.98505
Iteration 1:loglikelihood =-86.385606
Iteration 2:loglikelihood =-86.379623
Iteration 3:loglikelihood =-86.379622
Refining estimates:
Iteration 0:loglikelihood =-86.379622
Coxregression --Breslow methodforties
No.ofsubjects = 42 Numberofobs= 42
No.offailures = 30
Timeatrisk = 541
LRchi2(1) =15.21
Loglikelihood =-86.379622 Prob>chi2 =0.0001
------------------------------------------------------------------- -----------
_t|
_d|Haz.Ratio Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
trt|.2210887 .0905501 -3.685 0.000 .0990706 .4933877
------------------------------------------------------------------- -----------
.stcoxtrt,nohr
(sameiterations forlog-likelihood)
Coxregression --Breslow methodforties
No.ofsubjects = 42 Numberofobs= 42
No.offailures = 30
Timeatrisk = 541
LRchi2(1) =15.21
Loglikelihood =-86.379622 Prob>chi2 =0.0001
------------------------------------------------------------------- -----------
_t|
_d|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
trt|-1.509191 .4095644 -3.685 0.000 -2.311923 -.7064599
------------------------------------------------------------------- -----------
188
FittingPHmodelsinSAS-PROCPHREG
Ex.Leukemiadata
Title'CoxandOakesexample';
dataleukemia;
inputweeksremisstrtmt;
cards;
601
611
611
611 /*datafor6MPgroup*/
711
901
etc
110
110 /*dataforplacebo group*/
210
210
etc
;
procphregdata=leukemia;
modelweeks*remiss(0)=trtmt;
title'CoxPHModelforleukemia data';
run;
189
PROCPHREGOutput:
ThePHREGProcedure
DataSet:WORK.LEUKEM
Dependent Variable: FAILTIME TimetoRelapse
Censoring Variable: FAIL
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
42 30 12 28.57
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 187.970 172.759 15.211with1DF(p=0.0001)
Score . . 15.931with1DF(p=0.0001)
Wald . . 13.578with1DF(p=0.0002)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DFEstimate ErrorChi-Square Chi-Square Ratio
TRTMT 1-1.509191 0.40956 13.57826 0.0002 0.221
190
FittingPHmodelsinS-plus:coxphfunction
Herearesomeofthedatainleuk.dat:
tfx
110
110
210
210
310
...
1901
2001
2211
2311
2501
3201
3201
3401
3501
leuk_read.table("leuk.dat",header=T)
#specify Breslow handling ofties
print(coxph(Surv(t,f) ~x,leuk,method="breslow"))
#specify Efronhandling ofties(default)
print(coxph(Surv(t,f) ~x,leuk))
191
coxphOutput:
Call:
coxph(formula =Surv(t, f)~x,data=leuk,method="breslow")
coefexp(coef) se(coef) z p
x-1.51 0.221 0.41-3.680.00023
Likelihood ratiotest=15.2 on1df,p=0.0000961 n=42
Call:
coxph(formula =Surv(t, f)~x,data=leuk)
coefexp(coef) se(coef) z p
x-1.57 0.208 0.412-3.810.00014
Likelihood ratiotest=16.4 on1df,p=0.0000526 n=42
192
Comparethiswiththelogranktest
fromProcLifetest
(Usingthe\Test"statement)
TheLIFETEST Procedure
RankTestsfortheAssociation ofFAILTIME withCovariates
PooledoverStrata
Univariate Chi-Squares fortheLOGRANKTest
Test Standard Pr>
Variable Statistic Deviation Chi-Square Chi-Square
TRTMT 10.2505 2.5682 15.9305 0.0001
Notes:
²Thelogranktest=scoretestfromProcphreg!
Ingeneral,thescoretestwouldbeforallofthevariables
inthemodel,butinthiscase,wehaveonly\trtmt".
²Statadoesnotprovideascoretestinitsoutputfrom
theCoxmodel.However,thestcoxcommand with
thebreslow optionfortiesyieldsthesameLRtestas
theCMH-versionlogranktestfromtheststest,cox
command.
193
MoreNotes:
²TheCoxProportionalhazardsmodelhastheadvantage
overasimplelogranktestofgivingusanestimate of
the\riskratio"(i.e.,Á=¸1(t)=¸0(t)).Thisismore
informativ ethanjustateststatistic, andwecanalso
formcon¯dence intervalsfortheriskratio.
²Inthiscase,^Á=0:221,whichcanbeinterpreted tomean
thatthehazardforrelapseamongpatientstreatedwith
6-MPislessthan25%ofthatforplacebopatients.
²Fromthestslistcommand inStataorProclifetest
inSAS,wewereabletogetestimates oftheentiresur-
vivaldistribution ^S(t)foreachtreatmen tgroup;wecan't
immediately getthisfromourCoxmodelwithout fur-
therassumptions.Whynot?
194
Adjustmentsforties
Theproportionalhazardsmodelassumes acontinuoushaz-
ard{tiesarenotpossible.Therearefourproposedmodi¯-
cationstothelikelihoodtoadjustforties.
(1)Cox's(1972)modi¯cation: \discrete" method
(2)Peto-Breslowmethod
(3)Efron's(1977)method
(4)Exactmethod(Kalb°eischandPrentice)
(5)Exactmarginalmethod(stata)
Somenotation:
¿1;::::¿K theKordered, distinctdeathtimes
dj thenumberoffailuresat¿j
Hj the\history" oftheentiredataset,uptothe
j-thdeathorfailuretime,including thetime
ofthefailure,butnottheidentitiesofthedj
whofailthere.
ij1;:::ijdjtheidentitiesofthedjindividuals whofailat¿j
195
(1)Cox's(1972)modi¯cation: \discrete" method
Cox'smethodassumes thatiftherearetiedfailuretimes,
theytrulyhappenedatthesametime.Itisbasedona
discretelikelihood.
Thepartiallikelihoodis:
L(¯)=KY
j=1Pr(ij1;:::ijdjfailjdjfailat¿j;fromR)
=KY
j=1Pr(ij1;:::ijdjfailjinR(¿j))
P
`2s(j;dj)Pr(`1;::::`djfailjinR(¿j))
=KY
j=1exp(¯Zij1)¢¢¢exp(¯Zijdj)
P
`2s(j;dj)exp(¯Z`1)¢¢¢exp(¯Z`dj)
=KY
j=1exp(¯Sj)
P
`2s(j;dj)exp(¯Sj`)
where
²s(j;dj)isthesetofallpossiblesetsofdjindividuals that
canpossiblybedrawnfromtherisksetattime¿j
²SjisthesumoftheZ'sforallthedjindividuals who
failat¿j
²Sj`isthesumoftheZ'sforallthedjindividuals inthe
`-thsetdrawnoutofs(j;dj)
196
Whatdoesthisallmean??!!
Let'smodifyourprevious simpleexample toincludeties.
SimpleExample(withties)
Group0:4+;6;8+;9;10+=)Zi=0
Group1:3;5;5+;6;8+=)Zi=1
Ordered
failureNumberatriskLikelihoodContribution
jtimeXiGroup0Group1e¯Sj=P
`2s(j;dj)e¯Sj`
1355 e¯=[5+5e¯]
2544 e¯=[4+4e¯]
3642 e¯=[6+8e¯+e2¯]
4920 e0=2=1=2
Thetieoccursatt=6,whenR(¿j)=fZ=0:(6;8+;9;10+);
Z=1:(6;8+)g.Ofthe³6
2´=15possiblepairsofsubjects
atriskatt=6,thereare6pairsformedwherebotharefrom
group0(Sj=0),8pairsformedwithoneineachgroup
(Sj=1),and1pairsformedwithbothingroup1(Sj=2).
Problem: Withlargenumbersofties,thedenominator can
havemanymanytermsandbedi±culttocalculate.
197
(2)Breslowmethod:(default)
BreslowandPetosuggested replacing thetermP
`2s(j;dj)e¯Sj`
inthedenominator bythetermµP
`2R(¿j)e¯Z`¶dj,sothatthe
followingmodi¯edpartiallikelihoodwouldbeused:
L(¯)=KY
j=1e¯Sj
P
`2s(j;dj)e¯Sj`¼KY
j=1e¯Sj
µP
`2R(¿j)e¯Z`¶dj
Justi¯cation:
Supposeindividuals 1and2failfromf1;2;3;4gattime¿j.
LetÁ(i)bethehazardratioforindividual i(compared to
baseline).
e¯Sj
P
`2s(j;dj)e¯Sj`=Á(1)
Á(1)+Á(2)+Á(3)+Á(4)£Á(2)
Á(2)+Á(3)+Á(4)
+Á(2)
Á(1)+Á(2)+Á(3)+Á(4)£Á(1)
Á(1)+Á(3)+Á(4)
¼2Á(1)Á(2)
[Á(1)+Á(2)+Á(3)+Á(4)]2
ThePeto(Breslow)approximation willbreakdownwhen
thenumberoftiesarelargerelativetothesizeoftherisk
sets,andthentendstoyieldestimates of¯whicharebiased
toward0.
198
(3)Efron's(1977)method:
Efronsuggested anevencloserapproximation tothediscrete
likelihood:
L(¯)=KY
j=1e¯Sj
Ã
P
`2R(¿j)e¯Z`+j¡1
djP
`2D(¿j)e¯Z`!dj
LiketheBreslowapproximation, Efron'smethodwillyield
estimates of¯whicharebiasedtoward0whenthereare
manyties.
However,Allison(1995)recommends theEfronapproxima-
tionsinceitismuchfasterthantheexactmethodsandtends
toyieldmuchcloserestimates thanthedefaultBreslowap-
proach.
199
(4)Exactmethod(Kalb°eischandPrentice):
The\discrete" optionthatwediscussed in(1)isanexact
methodbasedonadiscretelikelihood(assuming thattied
eventstrulyAREtied).
Thissecondexactmethodisbasedonthecontinuouslike-
lihood,undertheassumption thatiftherearetiedevents,
thatisduetotheimprecise natureofourmeasuremen t,and
thattheremustbesometrueordering.
Allpossibleorderings ofthetiedeventsarecalculated, and
theprobabilities ofeacharesummed.
Example with2tiedevents(1,2)fromriskset(1,2,3,4):
e¯Sj
P
`2s(j;dj)e¯Sj`=e¯S1
e¯S1+e¯S2+e¯S3+e¯S4£e¯S2
e¯S2+e¯S3+e¯S4
+e¯S2
e¯S1+e¯S2+e¯S3+e¯S4£e¯S1
e¯S1+e¯S3+e¯S4
200
BottomLine:ImplicationsofTies
(SeeAllison(1995),p.127-137)
(1)Whentherearenoties,alloptionsgiveexactlythe
sameresults.
(2)Whenthereareonlyafewties,itwon'tmake
muchdi®erence whichmethodisused.However,since
theexactmethodswon'ttakemuchextracomputing
time,youmightaswelluseoneofthem.
(3)Whentherearemanyties(relativetothenumber
atrisk),theBreslowoption(default) performspoorly
(Farewell&Prentice,1980;Hsieh,1995).Bothofthe
approximatemethods,BreslowandEfron,yieldcoe±-
cientsthatareattenuated(biasedtoward0).
(4)Thechoiceofwhichexactmethodtouseshould
bebasedonsubstantivegrounds -arethetiedevent
timestrulytied?...oraretheytheresultofimprecise
measuremen t?
(5)Computingtimeofexactmethodsismuchlonger
thanthatoftheapproximatemethods.However,inmost
casesitwillstillbelessthan30secondsevenfortheexact
methods.
(6)Bestapproximatemethod-theEfronapproxi-
mationnearlyalwaysworksbetterthantheBreslow
method,withnoincreaseincomputing time,sousethis
optionifexactmethodsaretoocomputer-in tensive.
201
Example:Thefecundabilitystudy
Womenwhohadrecentlygivenbirth(orhadtriedtoget
pregnantforatleastayear)wereaskedtorecallhowlong
ittookthemtobecomepregnant,andwhether ornotthey
smokedduringthattime.Theoutcome ofinterestistimeto
pregnancy (measured inmenstrual cycles).
datafecund;
input smoke cycle status count;
cards;
0 1 1 198
0 2 1 107
0 3 1 55
0 4 1 38
0 5 1 18
0 6 1 22
..........................................
1 10 1 1
1 11 1 1
1 12 1 3
1 12 0 7
;
procphreg;
modelcycle*status(0) =smoke/ties=breslow; /*default */
freqcount;
procphreg;
modelcycle*status(0) =smoke/ties=discrete;
freqcount;
procphreg;
modelcycle*status(0) =smoke/ties=exact;
freqcount;
procphreg;
modelcycle*status(0) =smoke/ties=efron;
freqcount;
202
SASOutputforFecundabilitystudy:
AccountingforTies
******************************************************************* ********
TiesHandling: BRESLOW
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
SMOKE 1-0.329054 0.11412 8.31390 0.0039 0.720
******************************************************************* ********
TiesHandling: DISCRETE
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
SMOKE 1-0.461246 0.13248 12.12116 0.0005 0.630
******************************************************************* ********
TiesHandling: EXACT
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
SMOKE 1-0.391548 0.11450 11.69359 0.0006 0.676
******************************************************************* ********
TiesHandling: EFRON
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
SMOKE 1-0.387793 0.11402 11.56743 0.0007 0.679
******************************************************************* ********
Forthisparticulardataset,doesitseemlikeit
wouldbeimportanttoconsiderthee®ectoftied
failuretimes?Whichmethodwouldbebest?
203
StataCommandsforPHModelwithTies:
Stataalsoo®ersfouroptionsforadjustmen tswithtieddata:
²breslow (default)
²efron
²exactp(sameasthe\discrete" optioninSAS)
²exactm-anexactmarginal likelihoodcalculation
(di®erentthanthe\exact"optioninSAS)
FecundabilityDataExample:
.stcoxsmoker, efronnohr
failure _d:status
analysis time_t:cycle
Iteration 0:loglikelihood =-3113.5313
Iteration 1:loglikelihood =-3107.3102
Iteration 2:loglikelihood =-3107.2464
Iteration 3:loglikelihood =-3107.2464
Refining estimates:
Iteration 0:loglikelihood =-3107.2464
Coxregression --Efronmethodforties
No.ofsubjects = 586 Numberofobs= 586
No.offailures = 567
Timeatrisk = 1844
LRchi2(1) =12.57
Loglikelihood =-3107.2464 Prob>chi2 =0.0004
------------------------------------------------------------------- -----------
_t|
_d|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
smoker|-.3877931 .1140202 -3.401 0.001 -.6112685 -.1643177
------------------------------------------------------------------- -----------
204
Aspecialcase:thetwo-sampleproblem
Previously ,wederivedthelogranktestfromanintuitiveper-
spective,assuming thatwehave(X01;±01):::(X0n0;±0n0)from
group0and(X11;±11);:::;(X1n1;±1n1)fromgroup1.
JustasaÂ2testforbinarydatacanbederivedfromalogistic
model,wewillseeherethatthelogranktestcanbederived
asaspecialcaseoftheCoxProportionalHazards model.
First,let'sre-de¯ne ournotation intermsof(Xi;±i;Zi):
(X01;±01);:::;(X0n0;±0n0)=)(X1;±1;0);:::;(Xn0;±n0;0)
(X11;±11);:::;(X1n1;±1n1)=)(Xn0+1;±n0+1;1);:::;(Xn0+n1;±n0+n1;1)
Inotherwords,wehaven0rowsofdata(Xi;±i;0)forthe
group0subjects,thenn1rowsofdata(Xi;±i;1)forthe
group1subjects.
Usingtheproportionalhazardsformulation,wehave
¸(t;Z)=¸0(t)e¯Z
Group0hazard: ¸0(t)
Group1hazard: ¸0(t)e¯
205
Thelog-partial likelihoodis:
logL(¯)=log2
664KY
j=1e¯Zj
P
`2R(¿j)e¯Z`3
775
=KX
j=12
664¯Zj¡log[X
`2R(¿j)e¯Z`]3
775
Takingthederivativewithrespectto¯,weget:
U(¯)=@
@¯`(¯)
=nX
j=1±j2
664Zj¡P
`2R(¿j)Z`e¯Z`
P
`2R(¿j)e¯Z`3
775
=nX
j=1±j(Zj¡¹Zj)
where¹Zj=P
`2R(¿j)Z`e¯Z`
P
`2R(¿j)e¯Z`
U(¯)iscalledthe\score" .
206
Aswediscussed earlierintheclass,oneusefulformofa
likelihood-basedtestisthescoretest.Thisisobtained by
usingthescoreU(¯)evaluatedatHoasateststatistic.
Let'slookmorecloselyattheformofthescore:
±jZjobservednumberofdeathsingroup1at¿j
±j¹Zjexpectednumberofdeathsingroup1at¿j
Why?UnderH0:¯=0,¹Zjissimplythenumberof
individuals fromgroup1intherisksetattime¿j(callthis
r1j),dividedbythetotalnumberintherisksetatthattime
(callthisrj).Thus,¹Zjapproximates theprobabilit ythat
giventhereisadeathat¿j,itisfromgroup1.
Thus,thescorestatisticisoftheform:
nX
j=1(Oj¡Ej)
Whenthereareties,thelikelihoodhastobereplaced byone
thatallowsforties.
InSASorStata:
discrete/exactp !Mantel-Haenszel logranktest
breslow!linearrankversionofthelogranktest
207
Ialreadyshowedyoutheequivalenceofthelinearranklo-
granktestandtheBreslow(default) CoxPHmodelinSAS
(p.24-25)
HereistheoutputfromSASfortheleukemiadatausingthe
method=discrete option:
Logrank testwithproclifetest -stratastatement
TestofEquality overStrata
Pr>
Test Chi-Square DFChi-Square
Log-Rank 16.7929 10.0001
Wilcoxon 13.4579 10.0002
-2Log(LR) 16.4852 10.0001
ThePHREGProcedure
DataSet:WORK.LEUKEM
Dependent Variable: FAILTIME TimetoRelapse
Censoring Variable: FAIL
Censoring Value(s): 0
TiesHandling: DISCRETE
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 165.339 149.086 16.252with1DF(p=0.0001)
Score . . 16.793with1DF(p=0.0001)
Wald . . 14.132with1DF(p=0.0002)
208
MoreontheCoxPHmodel
I.Con¯denceintervalsandhypothesistests
{Twomethodsforcon¯denceintervals
{Waldtestsandlikelihoodratiotests
{Interpretationofparameterestimates
{AnexamplewithrealdatafromanAIDS
clinicaltrial
II.Predictedsurvivalunderproportionalhazards
III.PredictedmediansandP-yearsurvival
209
I.ConstructingCon¯denceintervalsandtestsfor
theHazardRatio(seeH&L4.2,Collett3.4):
Manysoftwarepackagesprovideestimates of¯,butthehaz-
ardratioHR=exp(¯)isusuallytheparameter ofinterest.
Wecanusethedeltamethodtogetstandard errorsfor
exp(^¯):
Var(dHR)=Var(exp(^¯))=exp(2^¯)Var(^¯)
Constructingcon¯denceintervalsforexp(¯)
Twooptions:(assumingthat¯isascalar)
I.Usingse(exp^¯)obtained aboveviathedeltamethodas
se(exp^¯)=r
[Var(exp(^¯))],calculate theendpointsas:
[L;U]=[dOR¡1:96se(dOR);dOR+1:96se(dOR)]
II.Formacon¯dence intervalfor^¯,andthenexponentiate
theendpoints.
[L;U]=[e^¯¡1:96se(^¯);e^¯+1:96se(^¯)]
Whichapproachdoyouthinkwouldbethemost
preferable?
210
HypothesisTests:
Foreachcovariateofinterest,thenullhypothesisis
Ho:HRj=1,¯j=0
AWaldtest2oftheabovehypothesisisconstructed as:
Z=^¯j
se(^¯j)orÂ2=0
BB@^¯j
se(^¯j)1
CCA2
Thistestfor¯j=0assumes thatallothertermsinthe
modelareheld¯xed.
Note:ifwehaveafactorAwithalevels,thenwewouldneed
toconstruct aÂ2testwith(a¡1)df,usingateststatistic
basedonaquadratic form:
Â2
(a¡1)=c¯0
AVar(c¯A)¡1c¯A
where¯A=(¯2;:::;¯a)0arethe(a¡1)coe±cientscor-
respondingtoZ2;:::;Za(orZ1;:::;Za¡1,dependingonthe
reference group).
2The¯rstfollowsanormaldistribution, andthesecondfollowsaÂ2with1df.
STATAgivestheZstatistic,whileSASgivestheÂ2
1teststatistic(thep-values
arealsogiven,anddon'tdependonwhichform,ZorÂ2,isprovided)
211
LikelihoodRatioTests:
Supposethereare(p+q)explanatory variablesmeasured:
Z1;:::;Zp;Zp+1;:::;Zp+q
andproportionalhazardsareassumed.
Considerthefollowingmodels:
²Model1:(containsonlythe¯rstpcovariates)
¸i(t;Z)
¸0(t)=exp(¯1Z1+¢¢¢+¯pZp)
²Model2:(containsall(p+q)covariates)
¸i(t;Z)
¸0(t)=exp(¯1Z1+¢¢¢+¯p+qZp+q)
Thesearenestedmodels.Forsuchnestedmodels,wecan
construct alikelihoodratiotestof
H0:¯p+1=¢¢¢=¯p+q=0
as:
Â2
LR=¡2·
log(^L(1))¡log(^L(2))¸
UnderHo,thisteststatistic isapproximately distributed as
Â2withqdf.
212
SomeexamplesusingtheStatastcoxcommand:
Model1:
.usemac
.stsetmactime macstat
.stcoxkarnofrifclari,nohr
failure _d:macstat
analysis time_t:mactime
Coxregression --Breslow methodforties
No.ofsubjects = 1151 Numberofobs=1151
No.offailures = 121
Timeatrisk = 489509
LRchi2(3) =32.01
Loglikelihood =-754.52813 Prob>chi2 =0.0000
------------------------------------------------------------------- ----
_t|
_d|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- ----
karnof|-.0448295 .0106355 -4.215 0.000-.0656747 -.0239843
rif|.8723819 .2369497 3.6820.000 .4079691 1.336795
clari|.2760775 .2580215 1.0700.285-.2296354 .7817903
------------------------------------------------------------------- ----
213
Model2:
.stcoxkarnofrifclaricd4,nohr
failure _d:macstat
analysis time_t:mactime
Coxregression --Breslow methodforties
No.ofsubjects = 1151 Numberofobs=1151
No.offailures = 121
Timeatrisk = 489509
LRchi2(4) =63.74
Loglikelihood =-738.66225 Prob>chi2 =0.0000
------------------------------------------------------------------- ------
_t|
_d|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- ------
karnof|-.0368538 .0106652 -3.456 0.001 -.0577572 -.0159503
rif|.880338 .2371111 3.713 0.000 .4156089 1.345067
clari|.2530205 .2583478 0.979 0.327 -.253332 .7593729
cd4|-.0183553 .0036839 -4.983 0.000 -.0255757 -.0111349
------------------------------------------------------------------- ------
214
Notes:
²Ifweomitthenohroption,wewillgettheestimated
hazardratioalongwith95%con¯dence intervalsusing
MethodII(i.e.,formingaCIforthelogHR(beta),and
thenexponentiatingthebounds)
------------------------------------------------------------------- -----
_t|
_d|Haz.Ratio Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----
karnof|.9638171 .0102793 -3.456 0.001 .9438791 .9841762
rif|2.411715 .5718442 3.713 0.000 1.515293 3.838444
clari|1.28791 .3327287 0.979 0.327 .7762102 2.136936
cd4|.9818121 .0036169 -4.983 0.000 .9747486 .9889269
------------------------------------------------------------------- -----
²Wecanalsocompute thehazardratioourselves,byex-
ponentiatingthecoe±cients:
HRcd4=exp(¡0:01835)=0:98
WhyisthisHRsocloseto1,andyetstill
highlysigni¯cant?
WhatistheinterpretationofthisHR?
²Thelikelihoodratiotestforthee®ectofCD4istwice
thedi®erence inminuslog-likelihoodsbetweenthetwo
models:
Â2
LR=2¤(754:533¡(738:66))=31:74
Howdoesthisteststatisticcompare totheWaldÂ2test?
215
²Inthemacstudy,therewerethreetreatmen tarms(rif,
clari,andtherif+clari combination). Because wehave
onlyincluded therifandclarie®ectsinthemodel,
thecombination therapyisthe\reference" group.
²Wecanconduct anoveralltestoftreatmen tusingthe
testcommand inStata:
.testrifclari
(1)rif=0.0
(2)clari=0.0
chi2(2)=17.01
Prob>chi2=0.0002
fora2dfWaldchi-square testofwhetherbothtreatmen t
coe±cientsareequalto0.Thistestcommand canbe
usedtoconductanoveralltestforanynumberofe®ects.
²Thetestcommand canalsobeusedtotestwhether
thereisadi®erence betweentherifandclaritreat-
mentarms:
.testrif=clari
(1)rif-clari=0.0
chi2(1)=8.76
Prob>chi2=0.0031
216
SomeexamplesusingSASPROCPHREG
procphregdata=alloi;
modeldthtime*dthstat(0)=mlogrna cd4grp1 cd4grp2 combther
/risklimits;
cd4level: testcd4grp1, cd4grp2;
title1'Proportional hazards regression modelfortimetoDeath';
title2'Baseline viralloadandCD4predictors';
procphregdata=alloi;
modeldthtime*dthstat(0)=mlogrna cd4grp1 cd4grp2 combther decrs8incrs8
/risklimits;
cd4level: testcd4grp1, cd4grp2;
wk8resp: testdecrs8, incrs8;
Notes:
²The\risklimits"optiononthemodelstatementprovides95%
con¯denceintervalsusingMethodIIfrompage2.(i.e.,forming
aCIforthelogHR(beta),andthenexponentiatingthebounds)
²The\test"statementhasthefollowingform:
Label:testvarname1, varname2, ...,varnamek;
forakdfWaldchi-squaretestofwhetherthekcoe±cientsare
allequalto0.
²WecanusethesameapproachdescribedbyFreedmantoassess
thee®ectsofintermediateendpoints(incrs8,decrs8)onthe
treatmente®ect(i.e.,assesstheiruseassurrogatemarkers).
Thepercentageoftreatmente®ectexplained, °,isestimated
by:
^°=1¡^¯trt;M2
^¯trt;M1
whereM1isthemodelwithouttheintermediateendpointand
M2isthemodelwiththemarker.
217
OUTPUTFROMPROCPHREG (Model1)
Proportional hazards regression modelfortimetoDeath
Baseline viralloadandCD4predictors
DataSet:WORK.ALLOI
Dependent Variable: DTHTIME Timetodeath(days)
Censoring Variable: DTHSTAT Deathstatus(1=died,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
690 89 601 87.10
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1072.543 924.167 148.376 with4DF(p=0.0001)
Score . . 189.702 with4DF(p=0.0001)
Wald . . 127.844 with4DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>
Variable DF Estimate Error Chi-Square Chi-Square
MLOGRNA 1 0.833237 0.17808 21.89295 0.0001
CD4GRP1 1 2.364612 0.32436 53.14442 0.0001
CD4GRP2 1 1.171137 0.34434 11.56739 0.0007
COMBTHER 1 -0.497161 0.24389 4.15520 0.0415
218
OUTPUTFROMPROCPHREG ,continued
Outputfrom\risklimits" and\test"statements
Analysis ofMaximum Likelihood Estimates
Conditional RiskRatioand
95%Confidence Limits
Risk
Variable Ratio Lower UpperLabel
MLOGRNA 2.301 1.623 3.262logbaseline rna(rocheassay)
CD4GRP1 10.640 5.634 20.093CD4<=100
CD4GRP2 3.226 1.643 6.335100<CD4<=200
COMBTHER 0.608 0.377 0.981Combination therapy withAZT/ddI/ddC/Nvp
LinearHypotheses Testing
Wald Pr>
Label Chi-Square DFChi-Square
CD4LEVEL 55.0794 2 0.0001
219
OUTPUTFROMPROCPHREG ,(Model2)
Proportional hazards regression modelfortimetoDeath
Baseline viralloadandCD4predictors
DataSet:WORK.ALLOI
Dependent Variable: DTHTIME Timetodeath(days)
Censoring Variable: DTHSTAT Deathstatus(1=died,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
690 89 601 87.10
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1072.543 912.009 160.535 with6DF(p=0.0001)
Score . . 198.537 with6DF(p=0.0001)
Wald . . 132.091 with6DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>
Variable DF Estimate Error Chi-Square Chi-Square
MLOGRNA 1 0.893838 0.18062 24.48880 0.0001
CD4GRP1 1 2.023005 0.33594 36.26461 0.0001
CD4GRP2 1 1.001046 0.34907 8.22394 0.0041
COMBTHER 1 -0.456506 0.24687 3.41950 0.0644
DECRS8 1 -0.410919 0.26383 2.42579 0.1194
INCRS8 1 -0.834101 0.32884 6.43367 0.0112
220
OUTPUTFROMPROCPHREG ,continued
Outputfrom\risklimits" and\test"statements
Analysis ofMaximum Likelihood Estimates
Conditional RiskRatioand
95%Confidence Limits
Risk
Variable Ratio Lower UpperLabel
MLOGRNA 2.444 1.716 3.483logbaseline rna(rocheassay)
CD4GRP1 7.561 3.914 14.606CD4<=100
CD4GRP2 2.721 1.373 5.394100<CD4<=200
COMBTHER 0.633 0.390 1.028Combination therapy withAZT/ddI/ddC/Nvp
DECRS8 0.663 0.395 1.112Decrease>=0.5 logrnaatweek8?
INCRS8 0.434 0.228 0.827Increase>=50 CD4cells,week8?
LinearHypotheses Testing
Wald Pr>
Label Chi-Square DFChi-Square
CD4LEVEL 37.6833 20.0001
WK8RESP 10.4312 20.0054
Thepercentageoftreatmen te®ectexplained byincluding
theRNAandCD4responsetotreatmen tbyWeek8is:
^°=1¡¡0:456
¡0:497¼0:08
or8%.Thepercentageoftreatmen te®ectontimeto¯rst
opportunistic infection ordeathismuchhigher(about24%).
221
II.PredictedSurvivalusingPH
TheCoxPHmodelsaysthat¸i(t;Z)=¸0(t)exp(¯Z).
Whatdoesthisimplyaboutthesurvivalfunction,Sz(t),for
thei-thindividual withcovariatesZi?
Forthebaseline (reference) group,wehave:
S0(t)=e¡Rt
0¸0(u)du=e¡¤0(t)
Thisisbyde¯nition ofasurvivalfunction (seeintronotes).
Forthei-thpatientwithcovariatesZi,wehave:
Si(t)=e¡Rt0¸i(u)du=e¡¤i(t)
=e¡Rt0¸0(u)exp(¯Zi)du
=e¡exp(¯Zi)Rt0¸0(u)du
="
e¡Rt0¸0(u)du#exp(¯Zi)
=[S0(t)]exp(¯Zi)
(Thisusesthemathematical relationship [eb]a=eab)
222
Sayweareinterestedinthesurvivalpatternforsinglemales
inthenursinghomestudy.Basedontheprevious formula,
ifwehadanestimate forthesurvivalfunction intherefer-
encegroup,i.e.,^S0(t),wecouldgetestimates ofthesurvival
function foranysetofcovariatesZi.
Howcanweestimatethesurvivalfunction, S0(t)?
WecouldusetheKMestimator, butthereareafewdisad-
vantagesofthatapproach:
²Itwouldonlyusethesurvivaltimesforobservationscon-
tainedinthereference group,andnotalltherestofthe
survivaltimes.
²Itwouldtendtobesomewhat choppy,sinceitwould
re°ectthesmallersamplesizeofthereference group.
²It'spossiblethattherearenosubjectsinthedataset
whoareinthe\reference" group(ex.saycovariatesare
ageandsex;thereisnooneofage=0inourdataset).
223
Instead, wewilluseabaselinehazardestimator whichtakes
advantageoftheproportional hazardsassumption togeta
smootherestimate.
^Si(t)=[^S0(t)]exp(c¯Zi)
Usingtheaboveformula,wesubstitutec¯basedon¯ttingthe
CoxPHmodel,andcalculate ^S0(t)byoneofthefollowing
approaches:
²Breslowestimator (Stata)
²Kalb°eisc h/Prenticeestimator (SAS)
224
(1)BreslowEstimator:
^S0(t)=exp¡^¤0(t)
where^¤0(t)istheestimated cumulativebaselinehazard:
^¤(t)=X
j:¿j<t0
BB@dj
P
k2R(¿j)exp(¯1Z1k+:::¯pZpk)1
CCA
(2)Kalb°eisch/PrenticeEstimator
^S0(t)=Y
j:¿j<t^®j
where^®j;j=1;:::daretheMLE'sobtained byassum-
ingthatS(t;Z)satis¯es
S(t;Z)=[S0(t)]e¯Z=2
64Y
j:¿j<t®j3
75e¯Z
=Y
j:¿j<t®e¯Z
j
225
BreslowEstimator:furthermotivation
TheBreslowestimator isbasedonextending theconcept
oftheNelson-Aalen estimator totheproportional hazards
model.
Recallthatforasinglesamplewithnocovariates,theNelson-
AalenEstimator ofthecumulativehazardis:
^¤(t)=X
j:¿j<tdj
rj
wheredjandrjarethenumberofdeathsandthenumber
atrisk,respectively,atthej-thdeathtime.
Whentherearecovariatesandassuming thePHmodelabove,
onecangeneralize thistoestimate thecumulativebaseline
hazardbyadjusting thedenominator:
^¤(t)=X
j:¿j<t0
BB@dj
P
k2R(¿j)exp(¯1Z1k+:::¯pZpk)1
CCA
Heuristic: Theexpectednumberoffailuresin(t;t+±t)is
dj¼±t£X
k2R(t)¸0(t)exp(zk^¯)
Hence,
±t£¸0(tj)¼dj
P
k2R(t)exp(zk^¯)
226
Kalb°eisch/PrenticeEstimator:furthermotivation
Thismethodisanalogous totheKaplan-Meier Estimator.
Consider adiscretetimemodelwithhazard(1¡®j)atthe
j-thobserveddeathtime.
(Note:weuse®j=(1¡¸j)tosimplifythealgebra!)
Thus,forsomeone withz=0,thesurvivorshipfunction is
S0(t)=Y
j:¿j<t®j
andforsomeone withZ6=0,itis:
S(t;Z)=S0(t)e¯Z=2
64Y
j:¿j<t®j3
75e¯Z
=Y
j:¿j<t®e¯Z
j
Thelikelihoodcontributions underthismodelare:
²forsomeone censored att:S(t;Z)
²forsomeone whofailsattj:
S(t(j¡1);Z)¡S(tj;Z)=2
64Y
k<j®j3
75e¯z
[1¡®e¯Z
j]
Thesolution for®jsatis¯es:
X
k2Djexp(Zk¯)
1¡®exp(Zk¯)
j=X
k2Rjexp(Zk¯)
(NotewhathappenswhenZ=0)
227
Obtaining ^S0(t)fromsoftwarepackages
²StataprovidestheBreslowestimator ofS0(t;Z),butnot
predicted survivalsatspeci¯edcovariatevalues..... you
havetoconstruct theseyourself
²SASusestheKalb°eisc h/Prenticeestimator ofthebase-
linehazard,andcanprovideestimates ofsurvivalatar-
bitraryvaluesofthecovariateswithalittlebitofpro-
gramming.
Inpractice, theyareincredibly close!(seeFleming and
Harrington 1984,Communic ationsinStatistics )
228
UsingStatatoPredictSurvival
TheStatacommand basesurv calculates thepredicted sur-
vivalvaluesforthereference group,i.e.,thosesubjectswith
allcovariates=0.
(1)BaselineSurvival:
Toobtaintheestimated baseline survival^S0(t),follow
theexample below(forthenursinghomedata):
.usenurshome
.stsetlosfail
.stcoxmarried health, basesurv(prsurv)
.sortlos
.listlosprsurv
229
EstimatingtheBaselineSurvivalwithStata
los prsurv
1. 1.99252899
2. 1.99252899
3. 1.99252899
4. 1.99252899
5. 1.99252899
.
.
.
22. 1.99252899
23. 2.98671824
24. 2.98671824
25. 2.98671824
26. 2.98671824
27. 2.98671824
28. 2.98671824
29. 2.98671824
30. 2.98671824
31. 2.98671824
32. 2.98671824
33. 2.98671824
34. 2.98671824
35. 2.98671824
36. 2.98671824
37. 2.98671824
38. 2.98671824
39. 2.98671824
40. 3.98362595
41. 3.98362595
.
.
.
Statacreatesapredicted baseline survivalestimate for
everyobservedeventtimeinthedataset, evenifthere
areduplicates.
230
(2)PredictedSurvivalforSubgroups
Toobtaintheestimated survival^Si(t)foranyothersub-
group(i.e.,notthereference orbaseline group),follow
theStatacommands below:
.predict betaz,xb
.gennewterm=exp(betaz)
.genpredsurv=prsurv^newterm
.sortmarried healthlos
.listmarried healthlospredsurv
231
PredictingSurvivalforSubgroupswithStata
married health lospredsurv
1. 0 2 1.9896138
8. 0 2 2.981557
11. 0 2 3.9772769
13. 0 2 4.9691724
16. 0 2 5.9586483
........................................................... .....
300. 0 3 1.9877566
302. 0 3 2.9782748
304. 0 3 3.9732435
305. 0 3 4.9637272
312. 0 3 5.9513916
........................................................... .....
768. 0 4 1.9855696
777. 0 4 2.9744162
779. 0 4 3.9685058
781. 0 4 4.9573418
785. 0 4 5.9428996
.
.
.
1468. 1 4 1.9806339
1469. 1 4 2.9657326
1472. 1 4 3.9578599
1473. 1 4 5.9239448
........................................................... .....
1559. 1 5 1.9771894
1560. 1 5 2.9596928
1562. 1 5 3.9504684
1564. 1 5 4.9331349
232
UsingSAStoPredictSurvival
TheSAScommand BASELINE calculates thepredicted sur-
vivalvaluesattheeventtimesforagivensetofcovariate
values.
(1)Togettheestimated baseline survival^S0(t),createa
datasetwith0'sforvaluesofallcovariatesinthemodel
(2)Togettheestimated survival^Si(t)foranyothersub-
group(i.e.,notthereference orbaselinegroup),createa
datasetwhichinputsthebaseline valuesofthecovari-
atesforthesubgroup ofinterest.
Foreithercase,wethensupplythecorrespondingdataset
nametotheBASELINE command underPROCPHREG.
Bygivingtheinputdatasetseverallines,eachcorresponding
toadi®erentcombination ofcovariatevalues,wecancom-
putepredicted survivalvaluesformorethanonegroupat
once.
233
(1)BaselineSurvivalEstimate
(notethatthebaselinesurvivalfunction doesnotcorrespond
toanyobservationsinoursample,sincehealthstatusvalues
rangefrom2-5)
***Estimating Baseline Survival Function underPH;
datainrisks;
inputmarried health;
cards;
00
;
procphregdata=pop out=survres;
modellos*fail(0)=married health;
baseline covariates=inrisks out=outph survival=ps/nomean;
procprintdata=outph;
title1'Nursinghome data:Baseline Survival Estimate';
234
EstimatingtheBaselineSurvivalwithSAS
Nursinghome data:Baseline Survival Estimate
OBSMARRIED HEALTH LOS PS
1 0 0 01.00000
2 0 0 10.99253
3 0 0 20.98672
4 0 0 30.98363
5 0 0 40.97776
6 0 0 50.97012
7 0 0 60.96488
8 0 0 70.95856
9 0 0 80.95361
10 0 0 90.94793
11 0 0 100.94365
12 0 0 110.93792
13 0 0 120.93323
14 0 0 130.92706
15 0 0 140.92049
16 0 0 150.91461
17 0 0 160.91017
18 0 0 170.90534
19 0 0 180.90048
20 0 0 190.89635
21 0 0 200.89220
22 0 0 210.88727
23 0 0 220.88270
.
.
.
235
(2)PredictedSurvivalEstimateforSubgroup
ThefollowingSAScommands willgenerate thepredicted
survivalprobabilit yforeachcombination ofcovariates,at
everyobservedeventtimeinthedataset.
***Estimating Baseline Survival Function underPH;
datainrisks;
inputmarried health;
cards;
02
05
12
15
;
procphregdata=pop out=survres;
modellos*fail(0)=married health;
baseline covariates=inrisks out=outph survival=ps/nomean;
procprintdata=outph;
title1'Nursinghome data:predicted survival bysubgroup';
236
SurvivalEstimatesbyMaritalandHealthStatus
Nursinghome data:Predicted Survival bySubgroup
OBSMARRIED HEALTH LOS PS
1 0 2 01.00000
2 0 2 10.98961
3 0 2 20.98156
4 0 2 30.97728
........................................................... .....
171 0 21840.50104
172 0 21850.49984
........................................................... .....
396 0 5 01.00000
397 0 5 10.98300
398 0 5 20.96988
399 0 5 30.96295
........................................................... .....
474 0 5 780.50268
475 0 5 800.49991
........................................................... .....
791 1 2 01.00000
792 1 2 10.98605
793 1 2 20.97527
794 1 2 30.96955
........................................................... .....
897 1 21080.50114
898 1 21090.49986
........................................................... .....
1186 1 5 01.00000
1187 1 5 10.97719
1188 1 5 20.95969
1189 1 5 30.95047
........................................................... .....
1233 1 5 470.50519
1234 1 5 480.49875
237
Wecangetavisualpictureofwhatthepropor-
tionalhazardsassumptionimpliesbylookingat
thesefoursubgroups
Subgroup Single, healthy Single, unhealth
Married, healthy Married, unhealt0.00.10.20.30.40.50.60.70.80.91.0
LOS01002003004005006007008009001000
238
III.PredictedmediansandP-yearsurvival
PredictedMedians
Supposewewantto¯ndthepredicted mediansurvivalforan
individual withaspeci¯edcombination ofcovariates(e.g.,a
singlepersonwithhealthstatus5).
Threepossibleapproaches:
(1)Calculate themedianfromthesubsetofindividuals with
thespeci¯edcovariatecombination (usingKMapproach)
(2)Generate predicted survivalcurvesforeachcombination
ofcovariates,andobtainthemedians directly
OBSMARRIED HEALTH LOSPREDSURV
171 0 21840.50104
172 0 21850.49984
474 0 5 780.50268
475 0 5 800.49991
897 1 21080.50114
898 1 21090.49986
1233 1 5 470.50519
1234 1 5 480.49875
Recallthatpreviously wede¯nedthemedianasthe
smallestvalueoftforwhich^S(t)·0:5,sothemedians
fromabovewouldbe185,80,109,and48daysforsingle
healthy,singleunhealth y,marriedhealthy,andmarried
unhealth y,respectively.
239
(3)Generate thepredicted survivalcurvefromtheestimated
baseline hazard,asfollows:
Wewanttheestimated median(M)foranindividual
withcovariatesZi.Weknow
S(M;Z)=[S0(M)]e¯Zi=0:5
Hence,Msatis¯es(multiplying bothsidesbye¡¯Zi):
S0(M)=[0:5]e¡¯Z
Ex.Supposewewanttoestimate themediansurvival
forasingleunhealth ysubjectfromthenursinghome
data.Thereciprocalofthehazardratioforunhealth y
(health=5) is:e¡0:165¤5=0:4373,(where^¯=0:165for
healthstatus)
So,wewantMsuchthatS0(M)=(0:5)0:4373=0:7385
Sothemedianforsingleunhealth ysubjectisthe73.8th
percentileofthebaseline group.
OBSMARRIED HEALTH LOSPREDSURV
79 0 0 780.74028
80 0 0 800.73849
81 0 0 810.73670
Sotheestimated medianwouldstillbe80days.Note:simi-
larlogiccanbefollowedtoestimate otherquantilesbesides
themedian.
240
EstimatingP-yearsurvival
Supposewewantto¯ndtheP-yearsurvivalrateforanindi-
vidualwithaspeci¯edcombination ofcovariates,^S(P;Zi)
Foranindividual withZi=0,theP-yearsurvivalcanbe
obtained fromthebaseline survivorshipfunction, ^S0(P)
Forindividuals withZi6=0,itcanbeobtained as:
^S(P;Zi)=[^S0(P)]ec¯Zi
Notes:
²Although Isay\P-year"survival,theunitsoftimeina
particular datasetmaybedays,weeks,ormonths.The
answerherewillbeinthesameunitsoftimeasthe
originaldata.
²Ifc¯Ziispositive,thentheP-yearsurvivalrateforthei-
thindividual willbelowerthanforabaselineindividual.
Whyisthistrue?
241
ModelSelection inSurvivalAnalysis
Supposewehaveacensored survivaltimethatwewantto
modelasafunction ofa(possiblylarge)setofcovariates.
Twoimportantquestions are:
²Howtodecidewhichcovariatestouse
²Howtodecideifthe¯nalmodel¯tswell
Toaddressthesetopics,we'llconsider anewexample:
SurvivalofAtlanticHalibut-Smithetal
Survival TowDi® Length HandlingTotal
Obs Time CensoringDurationinofFishTime log(catch)
#(min) Indicator (min.) Depth(cm) (min.) ln(weight)
100353.0 1 301539 5 5.685
109111.0 1 100 544 29 8.690
11364.0 0 100 1053 4 5.323
116500.0 1 100 1044 4 5.323
...
Hosmer&Lemesho w
Chapter 5: ModelDevelopment
Chapter 6: Assessmen tofModelAdequacy
(sections 6.1-6.2)
242
ProcessofModelSelection
Collett(Section 3.6)hasanexcellentdiscussion ofvarious
approachesformodelselection. Inpractice, modelselection
proceedsthrough acombination of
²knowledgeofthescience
²trialanderror,common sense
²automatic variableselection procedures
{forwardselection
{backwardselection
{stepwiseselection
Manyadvocatetheapproachof¯rstdoingaunivariateanal-
ysisto\screen" outpotentiallysigni¯can tvariablesforcon-
sideration inthemultivariatemodel(seeCollett).
Let'sstartwiththisapproach.
243
UnivariateKMplotsofAtlanticHalibutsurvival
(continuousvariableshavebeendichotomized)Survival Distribution Function
0.00.10.20.30.40.50.60.70.80.91.0
SURVTIME0100200300400500600700800900100011001200
STRATA:TOWDUR=0TOWDUR=1
Survival Distribution Function
0.00.10.20.30.40.50.60.70.80.91.0
SURVTIME0100200300400500600700800900100011001200
STRATA:LENGTHGP=0LENGTHGP=1Survival Distribution Function
0.00.10.20.30.40.50.60.70.80.91.0
SURVTIME0100200300400500600700800900100011001200
STRATA:DEPTHGP=0DEPTHGP=1
Survival Distribution Function
0.00.10.20.30.40.50.60.70.80.91.0
SURVTIME0100200300400500600700800900100011001200
STRATA:HANDLGP=0HANDLGP=1Survival Distribution Function
0.00.10.20.30.40.50.60.70.80.91.0
SURVTIME0100200300400500600700800900100011001200
STRATA:LOGCATGP=0LOGCATGP=1
Whichcovariateslookliketheymightbeimportant?
244
AutomaticVariableselectionprocedures
inStataandSAS
StatisticalSoftware:
²Stata:swcommand beforecoxcommand
²SAS:selection= optiononmodelstatemen tof
procphreg
Options:
(1)forward
(2)backward
(3)stepwise
(4)bestsubset(SASonly,usingscoreoption)
Onedrawbackoftheseoptionsisthattheycanonlyhandle
variablesoneatatime.Whenmightthatbeadisadvantage?
245
Collett'sModelSelectionApproach
Section3.6.1
Thisapproachassumes thatallvariablesareconsidered to
beonanequalfooting,andthereisnoapriorireasonto
includeanyspeci¯cvariables(liketreatmen t).
Approach:
(1)Fitaunivariatemodelforeachcovariate,andidentify
thepredictors signi¯can tatsomelevelp1,say0:20.
(2)Fitamultivariatemodelwithallsigni¯can tunivariate
predictors, andusebackwardselection toeliminate non-
signi¯can tvariablesatsomelevelp2,say0.10.
(3)Starting with¯nalstep(2)model,consider eachofthe
non-signi¯can tvariablesfromstep(1)usingforwardse-
lection,withsigni¯cance levelp3,say0.10.
(4)Do¯nalpruning ofmain-e®ects model(omitvariables
thatarenon-signi¯can t,addanythataresigni¯can t),
usingstepwise regression withsigni¯cance levelp4.At
thisstage,youmayalsoconsider addinginteractions be-
tweenanyofthemaine®ectscurrentlyinthemodel,
underthehierarchicalprinciple.
Collettrecommends usingalikelihoodratiotestforallvari-
ableinclusion/exclusion decisions.
246
StataCommandforForwardSelection:
ForwardSelection =)usepe(®)option,where®isthe
signi¯cance levelforenteringavariableintothemodel.
.usehalibut
.stsetsurvtime censor
.swcoxsurvtime towdurdepthlengthhandling logcatch,
>dead(censor) pe(.05)
beginwithemptymodel
p=0.0000<0.0500 adding handling
p=0.0000<0.0500 adding logcatch
p=0.0010<0.0500 adding towdur
p=0.0003<0.0500 adding length
CoxRegression --entrytime0 Numberofobs=294
chi2(4) =84.14
Prob>chi2=0.0000
LogLikelihood =-1257.6548 PseudoR2=0.0324
------------------------------------------------------------------- --------
survtime |
censor|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- --------
handling |.0548994 .0098804 5.556 0.000 .0355341 .0742647
logcatch |-.1846548 .051015 -3.620 0.000 .2846423 -.0846674
towdur|.5417745 .1414018 3.831 0.000 .2646321 .818917
length|-.0366503 .0100321 -3.653 0.000 -.0563129 -.0169877
------------------------------------------------------------------- --------
247
StataCommandforBackwardSelection:
BackwardSelection =)usepr(®)option,where®is
thesigni¯cance levelforavariabletoremaininthemodel.
.swcoxsurvtime towdurdepthlengthhandling logcatch,
>dead(censor) pr(.05)
beginwithfullmodel
p=0.1991>=0.0500 removing depth
CoxRegression --entrytime0 Numberofobs=294
chi2(4) =84.14
Prob>chi2=0.0000
LogLikelihood =-1257.6548 PseudoR2=0.0324
------------------------------------------------------------------- -------
survtime |
censor|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -------
towdur|.5417745 .1414018 3.831 0.000 .2646321 .818917
logcatch |-.1846548 .051015 -3.620 0.000-.2846423 -.0846674
length|-.0366503 .0100321 -3.653 0.000-.0563129 -.0169877
handling |.0548994 .0098804 5.556 0.000 .0355341 .0742647
------------------------------------------------------------------- -------
248
StataCommandforStepwiseSelection:
StepwiseSelection =)usebothpe(:)andpr(:)options,
withpr(:)>pe(:)
.swcoxsurvtime towdurdepthlengthhandling logcatch,
>dead(censor) pr(0.10) pe(0.05)
beginwithfullmodel
p=0.1991>=0.1000 removing depth
CoxRegression --entrytime0 Numberofobs=294
chi2(4) =84.14
Prob>chi2=0.0000
LogLikelihood =-1257.6548 PseudoR2=0.0324
------------------------------------------------------------------- ------
survtime |
censor|Coef. Std.Err. zP>|z|[95%Conf.Interval]
---------+--------------------------------------------------------- ------
towdur|.5417745 .1414018 3.831 0.000.2646321 .818917
handling |.0548994 .0098804 5.556 0.000.0355341 .0742647
length|-.0366503 .0100321 -3.653 0.000-.0563129 -.0169877
logcatch |-.1846548 .051015 -3.620 0.000-.2846423 -.0846674
------------------------------------------------------------------- ------
Itisalsopossibletodoforwardstepwiseregression byin-
cludingbothpr(:)andpe(:)optionswithforward option
249
SASprogrammingstatementsformodelselection
datafish;
infile'fish.dat';
inputIDSURVTIME CENSORTOWDURDEPTHLENGTHHANDLING LOGCATCH;
run;
title'Survival ofAtlantic Halibut';
***automatic variable selection procedures;
procphregdata=fish;
modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch
/selection=stepwise slentry=0.1 slstay=0.1 details;
title2'Stepwise selection';
run;
procphregdata=fish;
modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch
/selection=forward slentry=0.1 details;
title2'Forward selection';
run;
procphregdata=fish;
modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch
/selection=backward slstay=0.1 details;
title2'Backward selection';
run;
procphregdata=fish;
modelsurvtime*censor(0)= towdurdepthlengthhandling logcatch
/selection=score;
title2'Bestsubsets selection';
run;
250
Finalmodelforstepwiseselectionapproach
Survival ofAtlantic Halibut
Stepwise selection
ThePHREGProcedure
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
TOWDUR 10.007740 0.00202 14.68004 0.0001 1.008
LENGTH 1-0.036650 0.01003 13.34660 0.0003 0.964
HANDLING 10.054899 0.00988 30.87336 0.0001 1.056
LOGCATCH 1-0.184655 0.05101 13.10166 0.0003 0.831
Analysis ofVariables NotintheModel
Score Pr>
Variable Chi-Square Chi-Square
DEPTH 1.6661 0.1968
Residual Chi-square =1.6661 with1DF(p=0.1968)
NOTE:No(additional) variables metthe0.1levelforentryintothe
model.
Summary ofStepwise Procedure
Variable Number Score Wald Pr>
StepEntered Removed InChi-Square Chi-Square Chi-Square
1HANDLING 147.1417 .0.0001
2LOGCATCH 218.4259 .0.0001
3TOWDUR 311.0191 .0.0009
4LENGTH 413.4222 .0.0002
251
OutputfromPROCSAS\score"option
NUMBEROF SCORE VARIABLES INCLUDED
VARIABLES VALUE INMODEL
147.1417 HANDLING
129.9604 TOWDUR
112.0058 LENGTH
1 4.2185 DEPTH
1 1.4795 LOGCATCH
---------------------------------
265.6797 HANDLING LOGCATCH
259.9515 TOWDURHANDLING
256.1825 LENGTHHANDLING
251.6736 TOWDURLENGTH
247.2229 DEPTHHANDLING
232.2509 TOWDURLOGCATCH
230.6815 TOWDURDEPTH
216.9342 DEPTHLENGTH
214.4412 LENGTHLOGCATCH
2 9.1575 DEPTHLOGCATCH
-------------------------------------
376.8829 LENGTHHANDLING LOGCATCH
376.3454 TOWDURHANDLING LOGCATCH
375.5291 TOWDURLENGTHHANDLING
369.0334 DEPTHHANDLING LOGCATCH
360.0340 TOWDURDEPTHHANDLING
356.4207 DEPTHLENGTHHANDLING
355.8374 TOWDURLENGTHLOGCATCH
352.4130 TOWDURDEPTHLENGTH
334.7563 TOWDURDEPTHLOGCATCH
324.2039 DEPTHLENGTHLOGCATCH
--------------------------------------------
494.0062 TOWDURLENGTHHANDLING LOGCATCH
481.6045 DEPTHLENGTHHANDLING LOGCATCH
477.8234 TOWDURDEPTHHANDLING LOGCATCH
475.5556 TOWDURDEPTHLENGTHHANDLING
459.1932 TOWDURDEPTHLENGTHLOGCATCH
-------------------------------------------------
596.1287 TOWDURDEPTHLENGTHHANDLING LOGCATCH
------------------------------------------------------
252
Bestmultivariatemodelforall3options
Survival ofAtlantic Halibut
BestMultivariate Model
ThePHREGProcedure
DataSet:WORK.FISH
Dependent Variable: TIME
Censoring Variable: CENSOR
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
294 273 21 7.14
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 2599.449 2515.310 84.140with4DF(p=0.0001)
Score . . 94.006with4DF(p=0.0001)
Wald . . 90.247with4DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
TOWDUR 10.007740 0.00202 14.68004 0.0001 1.008
LENGTH 1-0.036650 0.01003 13.34660 0.0003 0.964
HANDLING 10.054899 0.00988 30.87336 0.0001 1.056
LOGCATCH 1-0.184655 0.05101 13.10166 0.0003 0.831
253
Notes:
²Whenthehalibutdatawasanalyzed withtheforward,
backwardandstepwiseoptions,thesame¯nalmodelwas
reached.However,thiswillnotalwaysbethecase.
²Variablescanbeforcedintothemodelusingthelockterm
optioninStataandtheinclude optioninSAS.Any
variables thatyouwanttoforceinclusion ofmustbe
listed¯rstinyourmodelstatemen t.
²StatausestheWaldtestforbothforwardandbackward
selection, although ithasanoptiontousethelikelihood
ratiotestinstead(lrtest).SASusesthescoretestto
decidewhatvariablestoaddandtheWaldtestforwhat
variablestoremove.
²Ifyou¯tarangeofmodelsmanually,youcanapplythe
AICcriteriadescribedbyCollett:
minimize AIC=¡2log(^L)+(®¤q)
whereqisthenumberofunknownparameters inthe
modeland®istypicallybetween2and6(theysuggest
®=3).
Themodelisthenchosenwhichminimizes theAIC(sim-
ilartomaximizing log-likelihood,butwithapenaltyfor
numberofvariablesinthemodel)
254
Questions:
²Whenmightwewanttoforcecertainvariablesintothe
model?
(1)toexamine interactions
(2)tokeepmaine®ectsinthemodel
(3)tocalculate ascoretestforaparicular e®ect
²Woulditbepossibletogetdi®erent¯nalmodelsfrom
SASandStata?
²Basedonwhatwe'veseeninthebehaviorofWaldtests,
wouldSASorStatabemorelikelytoaddacovariateto
amodelinaforwardselection model?
²IfweusetheAICcriteriawith®=3,howdoesthat
compare tothelikelihoodratiotest?
255
Assessingoverallmodel¯t
Howdoweknowifthemodel¯tswell?
²Alwayslookatunivariateplots(Kaplan-Meiers)
ConstructaKaplan-Meier survivalplotforeachoftheimpor-
tantpredictors,liketheonesshownatthebeginningofthese
notes.
²Checkproportionalit yassumption (thiswillbethetopic
ofthenextlecture)
²Checkresiduals!
(a)generalized (Cox-Snell)
(b)martingale
(c)deviance
(d)Schoenfeld
(e)weightedSchoenfeld
256
Residuals forsurvivaldataareslightlydi®erentthanfor
othertypesofmodels,duetothecensoring. Beforewestart
talkingaboutresiduals, weneedanimportantbasicresult:
InverseCDF:
IfTi(thesurvivaltimeforthei-thindividual)has
survivorshipfunctionSi(t),thenthetransformed
randomvariableSi(Ti)(i.e.,thesurvivalfunction
evaluatedattheactualsurvivaltimeTi)should
befromauniformdistributionon[0;1],andhence
¡log[Si(Ti)]shouldbefromaunitexponentialdis-
tribution
Moremathematically:
IfTi»Si(t)
thenSi(Ti)»Uniform[0;1]
and¡logSi(Ti)»Exponential (1)
257
(a)Generalized(Cox-Snell)Residuals :
Theimplication ofthelastresultisthatifthemodeliscor-
rect,theestimated cumulativehazardforeachindividual at
thetimeoftheirdeathorcensoring shouldbelikeacensored
samplefromaunitexponential.Thisquantityiscalledthe
generalizedorCox-Snellresidual.
Hereishowthegeneralized residualmightbeused.Suppose
we¯taPHmodel:
S(t;Z)=[S0(t)]exp(¯Z)
or,intermsofhazards:
¸(t;Z)=¸0(t)exp(¯Z)
=¸0(t)exp(¯1Z1+¯2Z2+¢¢¢+¯kZk)
After¯tting,wehave:
²^¯1;:::;^¯k
²^S0(t)
258
So,foreachpersonwithcovariatesZi,wecanget
^S(t;Zi)=[^S0(t)]exp(¯Zi)
Thisgivesapredicted survivalprobabilit yateachtimetin
thedataset(seenotesfromtheprevious lecture).
Thenwecancalculate
^¤i=¡log[^S(Ti;Zi)]
Inotherwords,¯rstwe¯ndthepredictedsur-
vivalprobabilityattheactualsurvivaltimefor
anindividual,thenlog-transformit.
259
Example:Nursinghomedata
Saywehave
²asinglemale
²withactualduration ofstayof941days(Xi=941)
Wecompute theentiredistribution ofsurvivalprobabilities
forsinglemales,andobtain^S(941)=0:260.
¡log[^S(941;singlemale)]=¡log(0:260)=1:347
Werepeatthisforeveryoneinourdataset. Theseshouldbe
likeacensored samplefromanexponential(1)distribution
ifthemodel¯tsthedatawell.
Basedonthepropertiesofaunitexponentialmodel
²plotting¡log(^S(t))vstshouldyieldastraightline
²plotting log[¡logS(t)]vslog(t)shouldyieldastraight
linethrough theoriginwithslope=1.
Toconvinceyourselfofthis,startwithS(t)=e¡¸tand
calculate log[¡logS(t)].Whatdoyougetfortheslopeand
intercept?
(Note:thisdoesnotnecessarily meanthattheunderlying
distribution oftheoriginalsurvivaltimesisexponential!)
260
ObtainingthegeneralizedresidualsfromStata
²FitaCoxPHmodelwiththestcoxcommand, along
withthemgale(newvar)option
²Usethepredict command withthecsnelloption
²De¯neasurvivaldatasetusingtheCox-Snellresiduals
asthe\pseudo" failuretimes
²Calculate theestimated KMsurvival
²Takethelog[¡log(S(t))]basedontheabove
²Generate thelogoftheCox-Snellresiduals
²Graphlog[¡logS(t)]vslog(t)
.stcoxtowdurhandling lengthlogcatch, mgale(mg)
.predict csres,csnell
.stsetcsrescensor
.stslist
.stsgensurvcs=s
.genlls=log(-log(survcs))
.genloggenr=log(csres)
.graphllsloggenr
261
LLS
-6-5-4-3-2-1012
Log of SURVIVAL-5 -4 -3 -2 -1 0 1
Allisonstates\Cox-Snellresiduals...arenotveryinformativefor
Coxmodelsestimatedbypartiallikelihood."Heinsteadprefers
devianceresiduals(later).
262
ObtainingthegeneralizedresidualsfromSAS
Thegeneralizedresiduals canbeobtained fromSAS
after¯ttingaPHmodelusingtheoutputstatemen twith
thelogsurv option.
procphregdata=fish;
modelsurvtime*censor(0) =towdurhandling logcatch length;
outputout=phres logsurv=genres;
***takenegative logPr(survival) ateachpersons survtime;
dataphres;
setphres;
genres=-genres;
***Nowwetreatthegeneralized residuals astheinputdataset;
***toevaluate whether theassumption ofanexponential;
***distribution isappropriate;
proclifetest data=phres outsurv=survres;
timegenres*censor(0);
datasurvres;
setsurvres;
lls=log(-log(survival));
loggenr=log(genres);
procgplotdata=survres;
plotlls*loggenr;
run;
263
(b)MartingaleResiduals
(seeFleming andHarrington, p.164)
Martingale residuals arede¯nedforthei-thindividual as:
ri=±i¡^¤(Ti)
Properties:
²ri'shavemean0
²rangeofri'sisbetween¡1and1
²approximately uncorrelated (inlargesamples)
²Interpretation: -theresidualricanbeviewedasthe
di®erence betweentheobservednumberofdeaths(0or
1)forsubjectibetweentime0andTi,andtheexpected
numbersbasedonthe¯ttedmodel.
264
Themartingaleresiduals canbeobtained fromStata
usingthemgaleoptionshownpreviously .
Oncethemartingale residualiscreated,youcanplotitversus
thepredicted logHR(i.e.,¯Zi),oranyoftheindividual
covariates.
.stcoxtowdurhandling lengthlogcatch, mgale(mg)
.predict betaz=xb
.graphmgbetaz
.graphmglogcatch
.graphmgtowdur
.graphmghandling
.graphmglength
265
Themartingaleresiduals canbeobtained fromSAS
after¯ttingaPHmodelusingtheoutputstatemen twith
theresmart option.
Onceyouhavethem,youcan
²plotagainstpredicted values
²plotagainstcovariates
procphregdata=fish;
modelsurvtime*censor(0) =towdurhandling logcatch length;
outputout=phres resmart=mres xbeta=xb;
procgplotdata=phres;
plotmres*xb; /*predicted values*/
plotmres*towdur;
plotmres*handling;
plotmres*logcatch;
plotmres*length;
run;
Allisonstillprefersthedeviance residuals (next)
266
MartingaleResiduals
M
a
r
t
i
n
g
a
l
e
R
e
s
i
d
u
a
l-6-5-4-3-2-101
Towing Duration0102030405060708090100110120M
a
r
t
i
n
g
a
l
e
R
e
s
i
d
u
a
l-6-5-4-3-2-101
Length of fish (cm)20 30 40 50 60
M
a
r
t
i
n
g
a
l
e
R
e
s
i
d
u
a
l-6-5-4-3-2-101
Log(weight) of catch2 3 4 5 6 7 8M
a
r
t
i
n
g
a
l
e
R
e
s
i
d
u
a
l-6-5-4-3-2-101
Handling time0 10 20 30 40
M
a
r
t
i
n
g
a
l
e
R
e
s
i
d
u
a
l-6-5-4-3-2-101
Linear Predictor-3 -2 -1 0
267
(c)DevianceResiduals
Oneproblem withthemartingale residuals isthattheytend
tobeasymmetric.
Asolution istousedevianceresiduals .Forpersoni,
thesearede¯nedasafunction ofthemartingale residuals
(ri):
^Di=sign(^ri)r
¡2[^ri+±ilog(±i¡^ri)]
InStata,thedeviance residuals aregenerated usingthesame
approachastheCox-Snellresiduals.
.stcoxtowdurhandling lengthlogcatch, mgale(mg)
.predict devres, deviance
andthentheycanbeplottedversusthepredicted log(HR)
ortheindividual covariates,asshownfortheMartingale
residuals.
InSAS,justuseresdev optioninsteadofresmart.
Deviance residuals behavemuchlikeresiduals fromOLSre-
gression(i.e.,mean=0, s.d.=1). Theyarenegativeforobser-
vationswithsurvivaltimesthataresmallerthanexpected.
268
DevianceResiduals
D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-3-2-10123
Towing Duration0102030405060708090100110120D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-3-2-10123
Length of fish (cm)20 30 40 50 60
D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-3-2-10123
Log(weight) of total catch2 3 4 5 6 7 8D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-3-2-10123
Handling time0 10 20 30 40
D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-101234
Linear Predictor0.00.10.20.30.40.50.60.70.80.9
269
(d)SchoenfeldResiduals
Thesearede¯nedateachobservedfailuretimeas:
rs
ij=Zij(ti)¡¹Zj(ti)
Notes:
²representthedi®erence betweentheobservedcovariate
andtheaverageovertherisksetatthattime
²calculated foreachcovariate
²notde¯nedforcensored failuretimes.
²usefulforassessing timetrendorlackorproportionalit y,
basedonplotting versuseventtime
²sumtozero,haveexpectedvaluezero,andareuncorre-
lated(inlargesamples)
InStata,theSchoenfeldresiduals aregenerated inthestcox
command itself,usingtheschoenf( newvar(s) )option:
.stcoxtowdurhandling lengthlogcatch, schoenf(towres handres lenres
logres)
.graphtowressurvtime
InSAS,addtotheoutputline
RESSCH=name1 name2...namek
foruptokregressors inthemodel.
270
SchoenfeldResiduals
-60-50-40-30-20-1001020304050
Survival Time0100200300400500600700800900100011001200-20-1001020
Survival Time0100200300400500600700800900100011001200
-3-2-1012345
Survival Time0100200300400500600700800900100011001200-20-100102030
Survival Time0100200300400500600700800900100011001200
271
(e)WeightedSchoenfeldResiduals
Theseareactually usedmoreoftenthantheprevious un-
weightedversion,becausetheyaremorelikethetypicalOLS
residuals (i.e.,symmetric around0).
Theyarede¯nedas:
rw
ij=ncVrs
ij
wherecVistheestimated varianceof^¯.Theweightedresid-
ualscanbeusedinthesamewayastheunweightedonesto
assesstimetrendsandlackofproportionalit y.
InStata,usethecommand:
.stcoxtowdurlengthlogcatch handling depth,scaledsch(towres2
>lenres2 logres2 handres2 depres2)
.graphlogres2 survtime
InSAS,addtotheoutputline
WTRESSCH=name1 name2...namek
foruptokregressors inthemodel.
272
WeightedSchoenfeldResiduals
-0.08-0.07-0.06-0.05-0.04-0.03-0.02-0.010.000.010.020.030.040.050.060.07
Survival Time0100200300400500600700800900100011001200-0.4-0.3-0.2-0.10.00.10.20.30.4
Survival Time0100200300400500600700800900100011001200
-2-10123
Survival Time0100200300400500600700800900100011001200-0.4-0.3-0.2-0.10.00.10.20.30.40.5
Survival Time0100200300400500600700800900100011001200
273
UsingResidualplotstoexplorerelationships
Ifyoucalculate martingale ordeviance residuals withoutany
covariatesinthemodelandthenplotagainstcovariates,you
obtainagraphical impression oftherelationship betweenthe
covariateandthehazard.
InSplus,itiseasytodothis(alsopossibleinstatausingthe
\estimate" option)
**readinthedataset andfitacoxPHmodel
fish_read.table('fish.data',header=T)
x_fish$towdur
fishres_coxreg(fish$time, fish$censor, x,resid="martingale",iter.max=0)
**the2commands belowsetupthepostscript file,with4graphs
postscript("fishres.plt",horizontal=F,height=10,width=7)
par(mfrow=c(2,2),oma=c(0,0,2,0))
**plotthemartingale residuals vseachoftheothercovariates
**andaddalowesssmoothed fittotheplot
plot(fish$depth, fishres$resid, xlab="depth")
lines(lowess(fish$depth,fishres$resid,iter=0))
plot(fish$length, fishres$resid, xlab="length")
lines(lowess(fish$length,fishres$resid,iter=0))
plot(fish$handling, fishres$resid, xlab="handling")
lines(lowess(fish$handling,fishres$resid,iter=0))
plot(fish$logcatch, fishres$resid, xlab="logcatch")
lines(lowess(fish$logcatch,fishres$resid,iter=0))
274
SplusPlotsofMartingaleResidualsforCoxModel
containingonlytowingdurationasapredictor,
vsothercovariates
··
···
·····
·
··
···
···
···
··········
···
···
··
···
···
·
······
··
····
·····················
·
···
··
··
·
····
··
·····
··
··
·
··
······
··
·
·
···
··
····
····
··
···
··
··
·······
····
····
·····
·
··
··
·
··
··
·
··
·······················
·····
···
······
········
·············
··
··
·······
······
······
······
····
··
···
·
···
depthfishres$resid
0103050-4-3-2-101
··· ··
··· ·
·····
·
··
·····
···
···
········· ·
···
···
··
···
···
·
······
··
····
········ ·· ··· ··········
·
···
··
···
·
·····
··
·····
··
··
·
··
·······
··
··
·
···
··
····
····
··
···
··
··
·······
····
····
·····
·
··
··
·
··
··
·
···
············· ······ ·······
·····
···
········
········ ·
···· ·········
··
··
····· ····
······
······
···· ···
····
··
···
·
···
Lengthfishres$resid
304050-4-3-2-101
··
·· ·
·····
·
··
···
···
···
··········
···
···
··
···
···
·
······
··
····
·····················
·
···
··
··
·
····
··
·····
··
··
·
··
······
··
·
·
···
··
····
····
··
···
··
··
·······
····
····
·····
·
··
··
·
··
··
·
··
·······················
·····
···
······
········
·············
··
··
·······
······
······
······
····
··
···
·
···
logcatchfishres$resid
345678-4-3-2-101
···
···
·····
·
··
···
···
···
··········
···
···
··
···
···
·
······
··
····
······················
·
···
··
··
·
····
··
·····
··
··
·
··
······
··
··
·
···
··
····
····
··
···
··
··
·······
····
····
·····
·
··
··
·
··
··
·
··
························
·····
···
······
········
·············
··
··
·······
······
······
·······
····
··
·· ·
·
···
handlingfishres$resid
0102030-4-3-2-101
275
(f)Deletiondiagnostics
Deletion diagnostics arede¯nedgenerally as:
±i=^¯¡^¯(i)
Inotherwords,theyarethedi®erence betweentheestimated
regression coe±cientusingallobservationsandthatwithout
thei-thindividual. Thiscanbeusefulforassessing thein-
°uence ofanindividual.
InSASPROCPHREG, weusethedfbetaoption:
(Notethatthereisaseparatedfbetacalculated foreachof
thepredictors.)
procphregdata=fish;
modelsurvtime*censor(0)=towdur handling logcatch length;
idid;
outputout=phinfl dfbeta=dtow dhanddlogcdlength ld=lrchange;
procunivariate data=phinfl;
vardtowdhanddlogcdlength lrchange;
idid;
run;
Theprocunivariateprocedurewillsupplythe5smallest val-
uesandthe5largestvalues.The\id"statemen tmeansthat
thesewillbelabeledwiththevalueofidfromthedataset.
276
(g)OtherIn°uencediagnostics
Otherin°uencediagnostics:
TheLDoptionisanothermethodforcheckingin°uence. It
calculates howmuchthelog-likelihood(x2)wouldchangeif
thei-thpersonwasremovedfromthesample.
LDi=2·
logL(c¯)¡logL(c¯¡i)¸
c¯=MLEforallparameters witheveryoneincluded
c¯¡i=MLEwithi-thsubjectomitted
Again,theprocunivariateprocedureinSASwillidentify
theobservationswiththelargestandsmallest valuesofthe
lrchange diagnostic measure.
277
Canweimprovethemodel?
Theplotsappeartohavesomestructure, whichindicatethat
wecouldbeleavingsomething out.Itisalwaysagoodidea
tocheckforinteractions:
Inthiscase,thereareseveralimportantinteractions. Iused
abackwardselection modelforcingallmaine®ectstobe
included, andconsidering allpairwise interactions. Hereare
theresults:
Parameter Standard Wald Pr>Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
TOWDUR 1-0.075452 0.01740 18.79679 0.0001 0.927
DEPTH 10.123293 0.06400 3.71107 0.0541 1.131
LENGTH 1-0.077300 0.02551 9.18225 0.0024 0.926
HANDLING 10.004798 0.03221 0.02219 0.8816 1.005
LOGCATCH 1-0.225158 0.07156 9.89924 0.0017 0.798
TOWDEPTH 10.002931 0.0004996 34.40781 0.0001 1.003
TOWLNGTH 10.001180 0.0003541 11.10036 0.0009 1.001
TOWHAND 10.001107 0.0003558 9.67706 0.0019 1.001
DEPLNGTH 1-0.006034 0.00136 19.77360 0.0001 0.994
DEPHAND 1-0.004104 0.00118 12.00517 0.0005 0.996
Interpretation:
Handling alonedoesn'tseemtoa®ectsurvival,unlessitis
combinedwithalongertowingduration orshallowertrawl-
ingdepths.
278
Analternativemodelingstrategywhenwehave
fewercovariates
Withadatasetwithonly5maine®ects,itwouldmakesense
toconsider interactions fromthestart.Howmanywould
therebe?
²Fitmodelwithallmaine®ectsandpairwise interactions
²Thenusebackwardselection toeliminate non-signi¯can t
pairwise interactions (remembertoforcethemaine®ects
intothemodelatthisstage)
²Oncenon-signi¯can tpairwise interactions havebeenelim-
inated,youcouldconsider backwardsselection toelim-
inateanynon-signi¯can tmaine®ectsthatarenotin-
volvedinremaining interaction terms
²Afterobtaining ¯nalmodel,useresiduals tocheck¯tof
model.
279
Assessing thePHAssumption
Sofar,we'vebeenconsidering thefollowingCoxPHmodel:
¸(t;Z)=¸0(t)exp(¯Z)
=¸0(t)exp(X¯jZj)
where¯jistheparameter forthethej-thcovariate(Zj).
Importantfeaturesofthismodel:
(1)thebaselinehazarddependsont,butnotonthecovari-
atesZ1;:::;Zp
(2)thehazardratio,i.e.,exp(¯Z),dependsonthecovariates
Z=(Z1;:::;Zp),butnotontimet.
Assumption (2)iswhatledustocallthisaproportional
hazards model.That'sbecausewecouldtaketheratioof
thehazardsfortwoindividuals withcovariatesZiandZi0,
andwriteitasaconstantintermsofthecovariates.
280
ProportionalHazardsAssumption
HazardRatio:
¸(t;Zi)
¸(t;Zi0)=¸0(t)exp(¯Zi)
¸0(t)exp(¯Zi0)
=exp(¯Zi)
exp(¯Zi0)
=exp[¯(Zi¡Zi0)]
=exp[X¯j(Zij¡Zi0j)]=µ
Inthelastformula,Zijisthevalueofthej-thcovariatefor
thei-thindividual. Forexample,Z42mightbethevalueof
gender (0or1)forthethe4-thperson.
Wecanalsowritethehazardforthei-thpersonasaconstant
timesthehazardforthei0-thperson:
¸(t;Zi)=µ¸(t;Zi0)
Thus,theHRbetweentwotypesofindividuals isconstant
(i.e.,=µ)overtime.Thesearemathematical waysofstating
theproportionalhazardsassumption.
281
Thereareseveraloptionsforcheckingtheassumption ofpro-
portionalhazards:
I.Graphical
(a)Plotsofsurvivalestimates fortwosubgroups
(b)Plotsoflog[¡log(^S)]vslog(t)fortwosubgroups
(c)PlotsofweightedSchoenfeldresiduals vstime
(d)Plotsofobservedsurvivalprobabilities versusex-
pectedunderPHmodel(seeKleinbaum,ch.4)
II.Useofgoodnessof¯ttests-wecanconstruct
agoodness-of-¯t testbasedoncomparing theobserved
survivalprobabilit y(fromstslist)withtheexpected
(fromstcox)undertheassumption ofproportionalhaz-
ards-seeKleinbaumch.4
III.Includinginteractiontermsbetweenacovari-
ateandt(time-dep endentcovariates)
282
Howdoweinterprettheabove?
Kleinbaum(andothertexts)suggestastrategy ofassuming
thatPHholdsunlessthereisverystrongevidence tocounter
thisassumption:
²estimated survivalcurvesarefairlyseparated, thencross
²estimated logcumulativehazardcurvescross,orlook
veryunparallel overtime
²weightedSchoenfeldresiduals clearlyincreaseordecrease
overtime(youcould¯taOLSregression lineandseeif
theslopeissigni¯can t)
²testfortime£covariateinteraction termissigni¯can t
(thisrelatestotime-dep endentcovariates)
IfPHdoesn'texactlyholdforaparticular covariatebutwe
¯tthePHmodelanyway,thenwhatwearegettingissort
ofanaverageHR,averagedovertheeventtimes.
Inmostcases,thisisnotsuchabadestimate. Allisonclaims
thattoomuchemphasis isputontestingthePHassumption,
andnotenoughtootherimportantaspectsofthemodel.
283
Implicationsofproportionalhazards
Consider aPHmodelwithasinglecovariate,Z:
¸(t;Z)=¸0(t)e¯Z
Whatdoesthisimplyfortherelationbetweenthesurvivor-
shipfunctions atvariousvaluesofZ?
UnderPH,
log[¡log[S(t;Z)]]=log[¡log[S0(t)]]+¯Z
Ingeneral,wehavethefollowingrelationship:
¤i(t)=Zt
0¸i(u)du
=Zt
0¸0(u)exp(¯Zi)du
=exp(¯Zi)Zt
0¸0(u)du
=exp(¯Zi)¤0(t)
Thismeansthattheratioofthecumulativehazardsisthe
sameastheratioofhazardrates:
¤i(t)
¤0(t)=exp(¯Zi)=exp(¯1Z1i+¢¢¢+¯pZpi)
284
Usingtheaboverelationship, wecanshowthat:
¯Zi=log0
B@¤i(t)
¤0(t)1
CA
=log¤i(t)¡log¤0(t)
=log[¡logSi(t)]¡log[¡logS0(t)]
solog[¡logSi(t)]=log[¡logS0(t)]+¯Zi
Thus,toassessifthehazards areactually proportional to
eachotherovertime(usinggraphical optionI(b))
²calculate KaplanMeierCurvesforvariouslevelsofZ
²compute log[¡log(^S(t;Z))](i.e.,logcumulativehazard)
²plotvslog-time toseeiftheyareparallel(linesorcurves)
Note:IfZiscontinuous,breakintocategories.
285
Question:Whynotjustcomparetheunderlying
hazardratestoseeiftheyareproportional?
Here'stwosimulatedexamples withhazardswhicharetruly
proportionalbetweenthetwogroups:
Weibull-typehazard:U-shapedhazard:
Plots of hazard function vs time
Simulated data with HR=2 for men vs women
GenderWomen MenHAZARD
0.0000.0020.0040.0060.0080.010
Length of Stay (days)010020030040050060070080090010001100Plots of hazard function vs time
Simulated data with HR=2 for men vs women
GenderWomen MenHAZARD
0.0000.0020.0040.0060.0080.010
Length of Stay (days)010020030040050060070080090010001100
Reason1:It'shardtoeyeballthese¯guresand
seethatthehazardratesareproportional-it
wouldbeeasiertolookforaconstantshiftbe-
tweenlines.
286
Reason2:Estimatedhazardratestendtobe
moreunstablethanthecumulativehazardrate
Consider thenursinghomeexample (wherewethinkPHis
reasonable). Ifwegroupthedataintointervalsandcalculate
thehazardrateusingactuarial method,wegettheseplots:
200dayintervals:100dayintervals:
Plots of hazard function vs time
GenderWomen Men0.0000.0010.0020.0030.0040.0050.006
Length of Stay (days)01002003004005006007008009001000Plots of hazard function vs time
GenderWomen Men0.0000.0010.0020.0030.0040.0050.0060.0070.0080.009
Length of Stay (days)01002003004005006007008009001000
50dayintervals:25dayintervals:
Plots of hazard function vs time
GenderWomen Men0.0000.0020.0040.0060.0080.0100.012
Length of Stay (days)010020030040050060070080090010001100Plots of hazard function vs time
GenderWomen Men0.0000.0020.0040.0060.0080.0100.0120.014
Length of Stay (days)010020030040050060070080090010001100
287
Incontrast,thelogcumulativehazardplotsare
easiertointerpretandtendtogivemorestable
estimates
Ex:NursingHome-genderandmaritalstatus
proclifetest data=pop outsurv=survres;
timelos*fail(0);
stratagender;
formatgendersexfmt.;
title'Duration ofLengthofStayinnursing homes';
datasurvres;
setsurvres;
labellog_los='Log(Length ofstayindays)';
iflos>0thenlog_los=log(los);
ifsurvival<1 thenlls=log(-log(survival));
procgplotdata=survres;
plotlls*log_los=gender;
formatgendersexfmt.;
title2'Plotsoflog-log KMversuslog-time';
run;
Thestatementsformaritalstatusaresimilar,substituting married
forgender.
Note:Thisisequivalenttocomparing plotsofthelogcumu-
lativehazard,log(^¤(t)),betweenthecovariatelevels,since
¤(t)=Zt
0¸(u;Z)du=¡log[S(t)]
288
Assessmentofproportionalhazardsforgender
andmaritalstatusinnursinghomedata(Mor-
ris)
Plots of log-log KM versus log-time
GenderWomen MenLLS
-6-5-4-3-2-101
Log(Length of stay in days)0 1 2 3 4 5 6
Plots of log-log KM versus log-time
Marital Status SingleMarriedLLS
-6-5-4-3-2-101
Log(Length of stay in days)0 1 2 3 4 5 6
289
Assessingproportionalitywithseveralcovariates
Ifthereisenoughdataandyouonlyhaveacoupleofcovari-
ates,createanewcovariatethattakesadi®erentvaluefor
everycombination ofcovariatevalues.
Example: Healthstatusandgenderfornursinghome
datapop;
infile'ch12.dat';
inputlosagerxgendermarried healthfail;
ifgender=0 andhealth=2 thenhlthsex=1;
ifgender=1 andhealth=2 thenhlthsex=2;
ifgender=0 andhealth=5 thenhlthsex=3;
ifgender=1 andhealth=5 thenhlthsex=4;
procformat;
valuehsfmt
1='Healthier Women'
2='Healthier Men'
3='Sicker Women'
4='Sicker Men';
proclifetest data=pop outsurv=survres;
timelos*fail(0);
stratahlthsex;
formathlthsex hsfmt.;
title'Length ofStayinnursing homes';
datasurvres;
setsurvres;
labellog_los='Log(Length ofstayindays)';
labelhlthsex='Health/Gender Status';
iflos>0thenlog_los=log(los);
ifsurvival<1 lls=log(-log(survival));
procgplotdata=survres;
plotlls*log_los=hlthsex;
formathlthsex hsfmt.;
title2'Plotsoflog-log KMversuslog-time';
run;
290
Log[-log(surviv al)]PlotsforHealthstatus*gender
Plots of log-log KM versus log-time
Health/Gender Status Healthier Women Healthier Men
Sicker Women Sicker MenLLS
-5-4-3-2-101
Log(Length of stay in days)0 1 2 3 4 5 6
Iftherearetoomanycovariates(ornotenoughdata)forthis,
thenthereisawaytotestproportionalit yforeachvariable,
oneatatime,usingthestrati¯cation option.
291
Whatifproportionalhazardsfails?
²doastrati¯ed analysis
²includeatime-varyingcovariatetoallowchanging haz-
ardratiosovertime
²includeinteractions withtime
Thesecondtwooptionsrelatetotime-dep endentcovariates,
whichwillbecoveredinfuturelectures.
Wewillfocusonthe¯rstalternativ e,andthenthesecond
twooptionswillbebrie°ydescribed.
292
Strati¯edAnalyses
Suppose:
²wearehappywiththeproportionalit yassumption onZ1
²proportionalit ysimplydoesnotholdbetweenvarious
levelsofasecondvariableZ2.
IfZ2isdiscrete(withalevels)andthereisenoughdata,¯t
thefollowingstrati¯edmodel:
¸(t;Z1;Z2)=¸Z2(t)e¯Z1
Forexample, anewtreatmen tmightleadtoa50%decrease
inhazardofdeathversusthestandard treatmen t,butthe
hazardforstandard treatmen tmightbedi®erentforeach
hospital.
Astrati¯edmodelcanbeusefulbothforprimary
analysisandforcheckingthePHassumption.
293
AssessingPHAssumptionforSeveralCovariates
Supposewehaveseveralcovariates(Z=Z1,Z2,...Zp),and
wewanttoknowifthefollowingPHmodelholds:
¸(t;Z)=¸0(t)e¯1Z1+:::+¯pZp
Tostart,we¯tamodelwhichstrati¯es byZk:
¸(t;Z)=¸0Zk(t)e¯1Z1+:::+¯k¡1Zk¡1+¯k+1Zk+1+:::+¯pZp
Sincewecanestimate thesurvivalfunction foranysubgroup,
wecanusethistoestimate thebaseline survivalfunction,
S0Zk(t),foreachlevelofZk.
Thenwecompute¡logS(t)foreachlevelofZk,controlling
fortheothercovariatesinthemodel,andgraphically check
whether thelogcumulativehazardsareparallelacrossstrata
levels.
294
Ex:PHassumptionforgender(nursinghomedata):
²includemarried andhealthascovariatesinaCoxPH
model,butstratifybygender.
²calculate thebaseline survivalfunction foreachlevelof
thevariablegender(i.e.,malesandfemales)
²plotthelog-cumulativehazards formalesandfemales
andevaluatewhether thelines(curves)areparallel
Intheaboveexample, wemakethePHassumption formarried
andhealth,butnotforgender.
ThisislikegettingaKMsurvivalestimate foreachgen-
derwithoutassuming PH,butismore°exiblesincewecan
controlforothercovariates.
Wewouldrepeatthestrati¯cation foreachvariableforwhich
wewantedtocheckthePHassumption.
295
SASCodeforAssessing PHwithinStrati¯ed Model:
datapop;
infile'ch12.dat';
inputlosagerxgendermarried healthfail;
iflos<=0thendelete;
datainrisks;
inputmarried health;
cards;
02
;
procformat;
valuesexfmt
1='Male'
0='Female';
procphregdata=pop;
modellos*fail(0)=married health;
stratagender;
baseline covariates=inrisks out=outph
loglogs=lls /nomean;
procprintdata=outph;
title'LogCumulative HazardEstimates byGender';
title2'Controlling forMarital andHealthStatus';
dataoutph;
setoutph;
iflos>0thenlog_los=log(los);
labellog_los='Log(LOS)'
lls='Log Cumulative Hazard';
procgplotdata=outph;
plotlls*log_los=gender;
formatgendersexfmt.;
title1'Log-log Survival versuslog-time byGender';
run;
296
Log[-log(surviv al)]PlotsforGender
ControllingforMaritalandHealthStatus
GENDERFemale Male-5-4-3-2-1012
Log(LOS)5.85.96.06.16.26.36.46.56.66.76.86.97.0
297
ModelswithTime-dependentInteractions
Consider aPHmodelwithtwocovariatesZ1andZ2.The
standard PHmodelassumes
¸(t;Z)=¸0(t)e¯1Z1+¯2Z2
However,ifthelog-hazards arenotreallyparallelbetween
thegroupsde¯nedbyZ2,thenyoucantryaddinganinter-
actionwithtime:
¸(t;Z)=¸0(t)e¯1Z1+¯2Z2+¯3Z2¤t
Atestofthecoe±cient¯3wouldbeatestoftheproportional
hazardsassumption forZ2.
If¯3ispositive,thenthehazardratiowouldbeincreasing
overtime;ifnegative,thendecreasing overtime.
Changes incovariatestatussometimes occurnaturally dur-
ingastudy(ex.patientgetsakidneytransplan t),andare
handled byintroducingtime-dependentcovariates .
298
Using StatatoAssessProportionalHazards
Statahastwocommands whichcanbeusedtographically
assesstheproportionalhazardsassumption, usinggraphical
options(b)and(d)describedpreviously:
²stphplot: plots¡log[¡log(¡(S(t))]curvesforeach
category ofanominal orordinalindependentvariable
versuslog(time). Optionally ,theseestimates canbead-
justedforothercovariates.
²stcoxkm: plotsKaplan-Meier observedsurvivalcurves
andcompares themtotheCoxpredicted curvesforthe
samevariable.(Noneedtorunstcoxpriortothiscom-
mand,itwillbedoneautomatically)
Foreithercommand, youmusthavestsetyourdata¯rst.
Youmustspecifyby()withstcoxkm andyoumustspecify
eitherby()orstrata() withstphplot .
299
AssessingPHAssumptionforaSingleCovariate
byComparing¡log[¡log(S(t))]Curves
.usenurshome
.stsetlosfail
.stphplot, by(gender)
Notethatthelineswillbegoingfromtoplefttobottomright,
ratherthanbottomlefttotopright,sinceweareplotting
¡log[¡log(S(t))]ratherthanlog[¡log(S(t))].
Thiswillgiveaplotsimilartothatonp.10(top).
Ofcourse,you'llwanttomakeyourplotprettierbyadding
titlesandlabels,asfollows:
.stphplot, by(gender) xlabylabb2(log(Length ofStay))
>title(Evaluation ofPHAssumption) saving(phplot)
300
AssessingPHAssumptionforSeveralCovariates
byComparing¡log[¡log(S(t))]Curves
.usenurshome
.stsetlosfail
.genhlthsex=1
.replace hlthsex=2 ifhealth==2 &gender==1
.replace hlthsex=3 ifhealth==5 &gender==0
.replace hlthsex=4 ifhealth==5 &gender==1
.tabhlthsex
.stphplot, by(hlthsex)
Thiswillgiveaplotsimilartothatonp.12.
301
AssessingPHAssumptionforaSingleCovariate
ControllingfortheLevelsofOtherCovariates
.usenurshome
.stsetlosfail
.stphplot, strata(gender) adjust(married health)
Thiswillproduceaplotsimilartothatonp.18.
302
AssessingPHAssumptionforaCovariate
ByComparingCoxPHSurvivaltoKMSurvival
Toconstruct plotsbasedonoptionI(d),usethestcoxkm
command, eitherforasinglecovariateorforanewlygen-
eratedcovariate(likehlthsex )whichrepresentscombined
levelsofmorethanonecovariate.
.usenurshome
.stsetlosfail
.stcoxkm, by(gender)
.stcoxkm, by(hlthsex)
Asusual,you'llwanttoaddtitles,labels,andsaveyour
graphforlateruse.
303
Timevarying(ortime-dependent)covariates
References:
Allison(*) p.138-153
Hosmer&Lemesho wChapter 7,Section3
Kalb°eisc h&Prentice Section5.3
Collett Chapter 7
Kleinbaum Chapter 6
Cox&Oakes Chapter 8
Andersen &Gill Page168(Advanced!)
Sofar,we'vebeenconsidering thefollowingCoxPHmodel:
¸(t;Z)=¸0(t)exp(¯Z)
=¸0(t)exp(X¯jZj)
where¯jistheparameter forthethej-thcovariate(Zj).
Importantfeaturesofthismodel:
(1)thebaselinehazarddependsont,butnotonthecovari-
atesZ1;:::;Zp
(2)thehazardratioexp(¯Z)dependsonthecovariatesZ1;:::;Zp,
butnotontimet.
Nowwewanttorelaxthesecondassumption, andallowthe
hazardratiotodependontimet.
304
Exampletomotivatetime-dependentcovariates
Stanford Hearttransplan texample:
Variables:
²survival-timefromprogramenrollmentuntildeathorcen-
soring
²dead-indicatorofdeath(1)orcensoring(0)
²transpl-whetherpatienteverhadtransplant
(1ifyes,2ifno)
²surgery-previousheartsurgerypriortoprogram
²age-ageattimeofacceptanceintoprogram
²wait-timefromacceptanceintoprogramuntiltransplant
surgery(=.forthosewithouttransplant)
Initially,aCoxPHmodelwas¯tforpredicting survivaltime:
¸(t;Z)=¸0(t)exp(¯1¤transpl+¯2¤surgery+¯3¤age)
However,thismodelcouldgivemisleading results,sincepa-
tientswhodiedmorequicklyhadlesstimeavailabletoget
transplan ts.Amodelwithatimedependentindicator of
whether apatienthadatransplan tateachpointintime
mightbemoreappropriate:
¸(t;Z)=¸0(t)exp(¯1¤trnstime+¯2¤surgery+¯3¤age)
wheretrnstime=1iftranspl=1andwait>t
305
SAScodeforthesetwomodels
Time-independentcovariatefortranspl :
procphregdata=stanford;
modelsurvival*dead(0)=transpl surgery age;
run;
Time-dependentcovariatefortranspl :
procphregdata=stanford;
modelsurvival*dead(0)=trnstime surgery age;
ifwait>survival orwait=.thentrnstime=0;
elsetrnstime=1;
run;
306
Ifweaddtime-dep endentcovariatesorinteractions withtime
totheCoxproportionalhazardsmodel,thenitisnot\pro-
portionalhazards" modelanylonger.
Werefertoitasan\extended Coxmodel".
Comparison withasinglebinarypredictor (likehearttrans-
plant):
²Astandard CoxPHmodelwouldcompare thesurvival
distributions betweenthosewithoutatransplan t(ever)
tothosewithatransplan t.Asubject'stransplan tstatus
attheendofthestudywoulddetermine whichcategory
theywereputintofortheentirestudyfollow-up.
²Anextended Coxmodelwouldcompare theriskofan
eventbetweentransplan tandnon-transplan tateach
eventtime,butwouldre-evaluatewhichriskgroupeach
personbelongedinbasedonwhether they'dhadatrans-
plantbythattime.
307
RecidivismExample: (seeAllison,p.42)
Recidivismstudy:
432maleinmates werefollowedforoneyearafterrelease
fromprison,toevaluateriskofre-arrest asfunction of¯nan-
cialaid(fin),ageatrelease(age),race(race),full-time
workexperiencepriorto¯rstarrest(wexp),maritalsta-
tus(mar),parolestatus(paro=1ifreleased withparole,
0otherwise), andnumberofpriorconvictions(prio).Data
werealsocollected onemploymentstatusovertimeduring
theyear.
Time-independentmodel:
Atimeindependentmodelmightincludetheemployment
statusoftheindividual atthebeginning ofthestudy(1if
employed,0ifunemployed),orperhapsatanypointduring
theyear.
Time-dependentmodel:
However,employmentstatuschangesovertime,anditmay
bethemorerecentemploymentstatusthatwoulda®ectthe
hazardforre-arrest. Forexample, wemightwanttode¯ne
atime-dep endentcovariateforeachmonthofthestudythat
indicates whether theindividual wasemployedduringthe
pastmonth.
308
ExtendedCoxModel
Framework:
Forindividuali,supposewehavetheirfailuretime,failure
indicator, andasummary oftheircovariatevaluesovertime:
(Xi;±i;fZi(t);t2[0;Xi]g);
fZi(t);t2[0;Xi]grepresentsthecovariatepathforthe
i-thindividual whiletheyareinthestudy,andthecovariates
cantakedi®erentvaluesatdi®erenttimes.
Assumptions:
²conditional onanindividual's covariatehistory,thehaz-
ardforfailureattimetdependsonlyonthevalueofthe
covariatesatthattime:
¸(t;fZi(u);u2[0;t]g)=¸(t;Zi(t))
²theCoxmodelforthehazardholds:
¸(t;Zi(t))=¸0(t)e¯Zi(t)
Survivorfunction:
S(t;Z)=expf¡Zt
0exp(¯Z(u))¸0(u)dug
anddependsonthevaluesofthetimedependentvariables
overtheintervalfrom0tot.
ThisistheclassicformulationofthetimevaryingCoxre-
gression survivalmodel.
309
Kindsoftime-varyingcovariates:
²internalcovariates:
variablesthatrelatetotheindividuals, andcanonlybe
measured whenanindividual isalive,e.g.whiteblood
cellcount,CD4count
²externalcovariates:
{variablewhichchangesinaknownway,e.g.age,dose
ofdrug
{variablethatexiststotallyindependentlyofallindi-
viduals,e.g.airtemperature
310
ApplicationsandExamples
Theextended Coxmodelisused:
I.Whenimportantcovariateschangeduringastudy
²FraminghamHeartstudy
5209subjectsfollowedsince1948toexamine relation-
shipbetweenriskfactorsandcardiovasculardisease. A
particular example:
Outcome: timetocongestiv eheartfailure
Predictors: age,systolicbloodpressure, #cigarettes
perday
²LiverCirrhosis (Andersen andGill,p.528)
Clinicaltrialcomparing treatmen ttoplaceboforcirrho-
sis.Theoutcome ofinterestistimetodeath.Patients
wereseenattheclinicafter3,6and12months,then
yearly.
Fixedcovariates: treatmen t,gender,age(atdiagno-
sis)
Time-varyingcovariates: alcoholconsumption, nu-
tritional status,bleeding, albumin, bilirubin, alkaline
phosphatase andprothrom bin.
²RecidivismStudy(Allison, p.42)
311
II.Forcross-overstudies,toindicate changeintreatmen t
²Stanfordheartstudy(CoxandOakesp.129)
Between1967and1980,249patientsenteredaprogram
atStanford Universitywheretheywereregistered tore-
ceiveahearttransplan t.Ofthese,184receivedtrans-
plants,57diedwhilewaiting,and8droppedoutofthe
program forotherreasons. Doesgettingahearttrans-
plantimprovesurvival?Hereisasampleofthedata:
Waiting transplant? survival post total final
time transplant survival status
------------------------------------------------------------
49 2 . . 1
5 2 . . 1
0 1 15 15 1
35 1 3 38 1
17 2 . . 1
11 1 46 57 1
etc
(survivalisnotindicated aboveforthosewithouttransplants,butwasavail-
ableinthedataset)
Naiveapproach:Compare thetotalsurvivaloftrans-
plantedandnon-transplan ted.
Problem: LengthBias!
312
III.ForCompetingRisksAnalysis
Forexample, incancerclinicaltrials,\tumorresponse"(or
shrinking ofthetumor)isusedasanoutcome. However,
clinicians wanttoknowwhether tumorresponsecorrelates
withsurvival.
Forthispurpose,wecan¯tanextended Coxmodelfortime
todeath,withtumorresponseasatimedependentcovariate.
IV.FortestingthePHassumption
Forexample, wecan¯tthesetwomodels:
(1)TimeindependentcovariateZ1
¸(t;Z)=¸0(t)exp(¯1¤Z1)
ThehazardratioforZ1isexp(¯1).
(2)TimedependentcovariateZ1
¸(t;Z)=¸0(t)exp(¯1¤Z1+¯2¤Z1¤t)
ThehazardratioforZ1isexp(¯1+¯2t).
(note:wemaywanttoreplacetby(t¡t0),sothatexp(¯1)
representsHRatsomeconvenienttime,likethemediansurvival
time.)
Atestoftheparameter¯2isatestofthePHassumption.
(howdowegetthetest?...usingtheWaldtestfromthe
outputofsecondmodel,orLRtestformedbycomparing
thelog-likelihoodsofthetwomodels)
313
Partiallikelihoodwithtime-varyingcovariates
Startingoutjustasbefore...
SupposethereareKdistinctfailure(ordeath)times,and
let(¿1;::::¿K)representtheKordered, distinctdeathtimes.
Fornow,assumetherearenotieddeathtimes.
RiskSet:LetR(t)=fi:xi¸tgdenotethesetof
individuals whoare\atrisk"forfailureattimet.
Failure: Letijdenotethelabeloridentityoftheindividual
whofailsattime¿j,including thevalueoftheirtime-varying
covariateduringtheirtimeinthestudy
fZij(t);t2[0;¿j]g
History: LetHjdenotethe\history" oftheentiredata
set,uptothej-thdeathorfailuretime,including thetime
ofthefailure,butnottheidentityoftheonewhofails,also
including thevaluesofallcovariatesforeveryoneuptoand
including time¿j.
PartialLikelihood:Wehaveseenpreviously thatthe
partiallikelihoodcanbewrittenas
L(¯)=dY
j=1P(ijjHj)
=dY
j=1¸(¿j;Zj(¿j))
P
`2R(¿j)¸(¿j;Z`(¿j))
314
UnderthePHassumption, thisis:
L(¯)=dY
j=1exp(¯Zjj)
P
`2R(¿j)exp(¯Z`j)
whereZ`jisashort-cut waytodenotethevalueoftheco-
variatevectorforthe`-thpersonatthej-thdeathtime,
ie:
Z`j=Z`(¿j)
WhatifZisnotmeasured forperson`attime¿j?
²usethemostrecentvalue(assumes stepfunction)
²interpolate
²imputebasedonsomemodel
Inference (i.e.estimating theregression coe±cients,con-
structing scoretests,etc.)proceedssimilarly tostandard
case.Themaindi®erence isthatthevaluesofZwillchange
ateachriskset.
AllisonnotesthatitisveryeasytowritedownaCoxmodel
withtime-dep endentcovariates,butmuchharderto¯t(com-
putationally) andinterpret.
315
OldExamplerevisited:
Group0:4+;7;8+;9;10+
Group1:3;5;5+;6;8+
LetZ1begroup,andaddanother¯xedcovariateZ2
IDfailcensorZ1Z2e(¯1Z1+¯2Z2)
13111e¯1+¯2
24001e¯2
35111e¯1+¯2
45010e¯1
56111e¯1+¯2
671001
78001e¯2
88010e¯1
99101e¯2
10100001
ordered Partial
failure Individuals Likelihood
time(¿j)atriskfailureIDcontribution
3
5
6
7
9
316
Examplecontinued:
NowsupposeZ2(acompletely di®erentcovariate)isatime
varyingcovariate:
Z2(t)
IDfailcensorZ13456789
13110
240011
3511111
4501000
56110000
671000011
7800000000
8801000011
99100001111
1010000111111
ordered Partial
failure Individuals Likelihood
time(¿j)atriskfailureIDcontribution
3
5
6
7
9
317
SASsolutiontopreviousexamples
Title'Phregression: smallclassexample';
dataph;
inputtimestatusgroupz3z4z5z6z7z8z9;
cards;
3110......
40011.....
511111....
501000....
6110000...
71000011..
800000000.
801000011.
9100001111
10000111111
run;
procphreg;
modeltime*status(0)=group z3;
run;
procphreg;
modeltime*status(0)=group z;
z=z3;
if(time>=4)thenz=z4;
if(time>=5)thenz=z5;
if(time>=6)thenz=z6;
if(time>=7)thenz=z7;
if(time>=8)thenz=z8;
if(time>=9)thenz=z9;
run;
318
SASoutputfrom¯ttingbothmodels
Modelwithz3:
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 16.953 13.699 3.254with2DF(p=0.1965)
Score . . 3.669with2DF(p=0.1597)
Wald . . 2.927with2DF(p=0.2315)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
GROUP 11.610529 1.21521 1.75644 0.1851 5.005
Z3 11.360533 1.42009 0.91788 0.3380 3.898
Modelwithtime-dependentZ:
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 16.953 14.226 2.727with2DF(p=0.2558)
Score . . 2.725with2DF(p=0.2560)
Wald . . 2.271with2DF(p=0.3212)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
GROUP 11.826757 1.22863 2.21066 0.1371 6.214
Z 10.705963 1.20630 0.34249 0.5584 2.026
319
TheStanfordHeartTransplantdata
Title'Stanford hearttransplant data:C&OTable8.1';
dataheart;
infile'heart.dat';
inputwaittranspostsurvstatus;
run;
dataheart;
setheart;
iftrans=2 thensurv=wait;
run;
***naiveanalysis;
procphreg;
modelsurv*status(2)=tstat;
tstat=2-trans;
***analysis withtime-dependent covariate;
procphreg;
modelsurv*status(2)=tstat;
tstat=0;
if(trans=1 andsurv>=wait)thentstat=1;
run;
Thesecondmodeltookabouttwiceaslongtorunasthe
¯rstmodel,whichisusuallythecaseformodelswithtime-
dependentcovariates.
320
RESULTSforStanfordHeartTransplantdata:
Naivemodelwith¯xedtransplantindicator:
Criterion Covariates Covariates ModelChi-Square
-2LOGL 718.896 674.699 44.198with1DF(p=0.0001)
Score . . 68.194with1DF(p=0.0001)
Wald . . 51.720with1DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
TSTAT 1-1.999356 0.27801 51.72039 0.0001 0.135
Modelwithtime-dependenttransplantindicator:
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1330.220 1312.710 17.510with1DF(p=0.0001)
Score . . 17.740with1DF(p=0.0001)
Wald . . 17.151with1DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
TSTAT 1-0.965605 0.23316 17.15084 0.0001 0.381
321
RecidivismExample:
Hazardforarrestwithinoneyearofreleasefromprison:
Modelwithoutemploymentstatus
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1350.751 1317.496 33.266with7DF(p=0.0001)
Score . . 33.529with7DF(p=0.0001)
Wald . . 32.113with7DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
FIN 1-0.379422 0.1914 3.931 0.0474 0.684
AGE 1-0.057438 0.0220 6.817 0.0090 0.944
RACE 10.313900 0.3080 1.039 0.3081 1.369
WEXP 1-0.149796 0.2122 0.498 0.4803 0.861
MAR 1-0.433704 0.3819 1.290 0.2561 0.648
PARO 1-0.084871 0.1958 0.188 0.6646 0.919
PRIO 10.091497 0.0287 10.200 0.0014 1.096
Whataretheimportantpredictorsofrecidivism?
322
RecidivismExample:(cont'd)
Now,weusetheindicators ofemploymentstatusforeachof
the52weeksinthestudy,recorded asemp1-emp52 .
Wecan¯tthemodelin2di®erentways:
procphregdata=recid;
modelweek*arrest(0)=fin ageracewexpmarparroprioemployed
/ties=efron;
arrayemp(*)emp1-emp52;
doi=1to52;
ifweek=ithenemployed=emp(i);
end;
run;
***ashortcut;
procphregdata=recid;
modelweek*arrest(0)=fin ageracewexpmarparroprioemployed
/ties=efron;
arrayemp(*)emp1-emp52;
employed=emp(week);
run;
Thesecondwaytakes23%lesstimethanthe¯rst
way,buttheresultsarethesame.
323
RecidivismExample:Output
ModelWITHemploymentastime-dependentcovariate
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
FIN 1-0.356722 0.1911 3.484 0.0620 0.700
AGE 1-0.046342 0.0217 4.545 0.0330 0.955
RACE 10.338658 0.3096 1.197 0.2740 1.403
WEXP 1-0.025553 0.2114 0.015 0.9038 0.975
MAR 1-0.293747 0.3830 0.488 0.4431 0.745
PARO 1-0.064206 0.1947 0.109 0.7416 0.938
PRIO 10.085139 0.0290 8.644 0.0033 1.089
EMPLOYED 1-1.328321 0.2507 28.070 0.0001 0.265
Iscurrentemploymentimportant?
Dotheothercovariateschangemuch?
Canyouthinkofanyproblemwithusingcurrent
employmentasapredictor?
324
Anotheroptionforassessingimpactofemploy-
ment
Allisonsuggests usingtheemploymentstatusofthepast
weekratherthanthecurrentweek,asfollows:
procphregdata=recid;
whereweek>1;
modelweek*arrest(0)=fin ageracewexpmarparroprioemployed
/ties=efron;
arrayemp(*)emp1-emp52;
employed=emp(week-1);
run;
Thecoe±cientforemplo yedchangesfrom-1.33
to-0.79,sotheriskratioisabout0.45insteadof
0.27.Itisstillhighlysigni¯cantwithÂ2=13:1.
Doesthismodelimprovethecausalinterpreta-
tion?
Otheroptionsfortime-dep endentcovariates:
²multiplelagsofemploymentstatus(week-1,week-2,etc.)
²cumulativeemploymentexperience(proportionofweeks
worked)
325
Somecautionarynotes
²Time-varyingcovariatesmustbecarefully constructed
toensureinterpretabilit y
²Thereisnopointaddingatime-varyingcovariatewhose
valuechangesthesameasstudytime.....youwillget
thesameanswerasusinga¯xedcovariatemeasured at
studyentry.Forexample, supposewewanttostudythe
e®ectofageontimetodeath.
Wecould
1.useageatstartofthestudyasa¯xedcovariate
2.ageasatimevaryingcovariate
However,theresultswillbethesame!Why?
326
Usingtime-varyingcovariatestoassessmodel¯t
Supposewehavejust¯tthefollowingmodel:
¸(t;Z)=¸0(t)exp(¯1Z1+¯2Z2+:::¯pZp)
E.g.,thenursinghomedatawithgender,maritalstatusand
health.
Supposewewanttotesttheproportionalit yassumption on
health(Zp)
Createanewvariable:
Zp+1(t)=Zp¤°(t)
where°(t)isaknownfunction oftime,suchas
°(t)=t
orlog(t)
ore¡½t
orIft>t¤g
ThentestingH0:¯p+1=0isatestfornon-prop ortionalit y
327
Illustration:ColonCancerdata
***modelwithout time*covariate interaction;
procphregdata=surv;
modelsurvtime*censs(1) =trtmstagen;
Modelwithouttime*stage interaction
EventandCensored Values
Percent
Total Event Censored Censored
274 218 56 20.44
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1959.927 1939.654 20.273with2DF(p=0.0001)
Score . . 18.762with2DF(p=0.0001)
Wald . . 18.017with2DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
TRTM 10.016675 0.13650 0.01492 0.9028 1.017
STAGEN 1-0.701408 0.16539 17.98448 0.0001 0.496
328
***modelWITHtime*covariate interaction;
procphregdata=surv ;
modelsurvtime*censs(1) =trtmstagentstage;
tstage=stagen*exp(-survtime/1000);
ModelWITHtime*stage interaction
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1959.927 1902.374 57.553with3DF(p=0.0001)
Score . . 35.960with3DF(p=0.0001)
Wald . . 19.319with3DF(p=0.0002)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr>Risk
Variable DFEstimate Error Chi-Square Chi-Square Ratio
TRTM 10.008309 0.13654 0.00370 0.9515 1.008
STAGEN 11.402244 0.45524 9.48774 0.0021 4.064
TSTAGE 1-8.322371 2.04554 16.55310 0.0001 0.000
LikeCoxandOakes,wecanrunafewdi®erentmodels
329
Time-varyingcovariatesinStata
CreateadatasetwithanIDcolumn,andonelineperperson
foreachdi®erentvalueofthetimevaryingcovariate.
.infileidtimestatusgroupzusingcox4_stata.dat
or
.inputidtimestatusgroup z
13 110
25 010
35 111
46 110
56 010
58 011
64 001
75 000
77 101
88 000
95 000
99 101
103 000
1010 001
.end
.stsettimestatus
.coxtimegroupz,dead(status) tvid(id)
------------------------------------------------------------------- -------- ---
time|
status|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -------- ---
group|1.826757 1.228625 1.487 0.137 -.5813045 4.234819
z|.7059632 1.206304 0.585 0.558 -1.65835 3.070276
------------------------------------------------------------------- -------- ---
330
Time-varyingcovariatesinSplus
Createadatasetwithstartandstopvaluesoftime:
idstartstopstatusgroup z
103110
205010
305111
406110
506010
568011
604001
705000
757101
808000
905000
959101
1003000
10310 001
331
ThentheSpluscommands andresultsare:
Commands:
y_read.table("cox4_splus.dat",header=T)
agreg(y$start,y$stop,y$status,cbind(y$group,y$z))
Results:
AliveDeadDeleted
95 0
coefexp(coef) se(coef) zp
[1,]1.827 6.21 1.231.4870.137
[2,]0.706 2.03 1.210.5850.558
exp(coef) exp(-coef) lower.95upper.95
[1,] 6.21 0.161 0.559 69.0
[2,] 2.03 0.494 0.190 21.5
Likelihood ratiotest=2.73on2df,p=0.256
Efficient scoretest=2.73on2df,p=0.256
332
PiecewiseCoxModel:(Collett,Chapter10)
Atimedependentcovariatecanbeusedtocreateapiecewise
PHcoxmodel.Supposeweareinterested incomparing two
treatmen ts,and:
²HR=µ1duringtheinterval(0;t1)
²HR=µ2duringtheinterval(t1;t2)
²HR=µ3duringtheinterval(t2;1)
De¯nethefollowingcovariates:
²X-treatmen tindicator
(X=0!standard,X=1!newtreatmen t)
²Z2-indicator ofchangeinHRduring2ndinterval
Z2(t)=8
><
>:1ift2(t1;t2)andX=1
0otherwise
²Z3-indicator ofchangeinHRduring3rdinterval
Z3(t)=8
><
>:1ift2(t2;1)andX=1
0otherwise
Themodelforthehazardforindividualiis:
¸i(t)=¸0(t)expf¯1xi+¯2z2i(t)+¯3z3i(t)g
Whataretheloghazardratiosforanindividual onthenew
treatmen trelativetooneonthestandard treatmen t?
333
Timevarying(ortime-dependent)covariates
CaseStudyofMACDiseaseTrial
ACTG196wasarandomizedclinicaltrialtostudythee®ects
ofcombinationregimensonpreventionofMAC(mycobacterium
aviumcomplex)disease,whichisoneofthemostcommonoppor-
tunisticinfectionsinAIDSpatientsandisassociatedwithhighmor-
talityandmorbidity.
Thetreatmentregimenswere:
²clarithromycin(new)
²rifabutin(standard)
²clarithromycinplusrifabutin
ThistrialenrolledpatientsbetweenApril1993andFebruary1994,
andfollowedpatientsthroughAugust1995.InFebruaryof1994,the
dosageofrifabutinwasreducedfrom3capsulesperday(450mg)
to2capsulesperday(300mg)duetoconcernoveruveitis,an
adverseexperienceresultinginin°ammationoftheuvealtractin
theeyes(about3-4%ofpatientsreporteduveitis).Allpatientswere
toreducetheirdosagebyMarch8,1994.However,somepatients
hadalreadydiscontinuedthetreatment,died,ordiscontinuedthe
study.
Themainintent-to-treatanalysiscomparedthe3treatmentarms
withoutadjustingforthischangeindosage.
Othersupportinganalysesattemptedtountanglethee®ectofthis
\studywidedosereduction" (SWDR).
334
ProportiononeachtreatmentarmwithSWDR
Treatment bystudywidedosereduction
TABLEOFTRTMTBYSWDRSTAT
TRTMT
SWDRSTAT(Study WideDoseReduction Status)
Frequency|
RowPct|No |Yes |Total
---------+--------+--------+
R |125|266|391
|31.97|68.03|
---------+--------+--------+
C+R |170|219|389
|43.70|56.30|
---------+--------+--------+
C |124|274|398
|31.16|68.84|
---------+--------+--------+
Total 419 759 1178
STATISTICS FORTABLEOFTRTMTBYSWDRSTAT
Statistic DFValue Prob
------------------------------------------------------
Chi-Square 216.820 0.001
Likelihood RatioChi-Square 216.610 0.001
Mantel-Haenszel Chi-Square 10.067 0.795
PhiCoefficient 0.119
Contingency Coefficient 0.119
Cramer's V 0.119
SampleSize=1178
335
OriginalLogranktestComparing 3TreatmentArms
(Howwouldyougetpairwisetests?)
Dependent Variable: MACTIME TimetoMACdisease (days)
Censoring Variable: MACSTAT MACstatus(1=yes,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
1178 121 1057 89.73
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1525.932 15.133with2DF(p=0.0005)
Score . . 15.890with2DF(p=0.0004)
Wald . . 15.209with2DF(p=0.0005)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.231842 0.25748 0.81074 0.3679 1.261
RIF 10.826883 0.23601 12.27480 0.0005 2.286
Variable Label
CLARI 1=Clarithromycin arm,0otherwise
RIF 1=Rifabutin arm,0otherwise
LinearHypotheses Testing
Wald Pr>
Label Chi-Square DFChi-Square
TEST_TRT 15.2094 2 0.0005
336
Kaplan-Meier SurvivalPlot
EstimatedProbabilitiesofRemainingMAC-freeSurvival Distribution Function
0.00.10.20.30.40.50.60.70.80.91.0
Time to MAC disease (days)0100200300400500600700800900
STRATA:TRTMT=Clar + Rif TRTMT=Clarithro TRTMT=Rifabutin
%ps(mactrt.ps,mode=replace);
proclifetest data=weighted noprint outsurv=survres
graphics nocensplots=(s);
timemactime*macstat(0);
stratatrtmt;
title'TimetoMACbyTreatment Regimen';
formattrtmttrtfmt.;
run;
337
Howwelldoesthismodel¯t?
Let'stakealookattheresidualplots...
First,thedeviance residuals:
D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-101234
1=Rifabutin arm, 0 otherwise-1 0 1D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-101234
1=Clarithromycin arm, 0 otherwise-1 0 1
D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-101234
Linear Predictor0.00.10.20.30.40.50.60.70.80.9
Plottingdevianceresidualsvsbinarycovariates
isnotveryuseful.
338
Howaboutthegeneralizedresiduals?
(Aretheylikeasamplefromacensored unitexponential?)
(i.e., is slope=1, intercept=0)
LLS
-8-7-6-5-4-3-2-1
Log(generalized residual)-8 -7 -6 -5 -4 -3 -2 -1
intercept=0.056
slope=1.028
(basedon¯ttingaregression linetoresiduals)
339
Wecanalsolookatthelogcumulativehazard
plots(i.e.,log[¡log(^S)])versuslogtimetoseewhether
thelinesareparallelforthethreetreatmentgroups.
Plot of log-log KM versus log-time
MAC Prophylaxis Therapy Rifabutin ClarithroClar + RifL
n
[
-
l
n
(
S
)
]
-6-5-4-3-2-1
Log(Time to MAC)3 4 5 6
(Ihavejoinedtheindividual pointsusingi=joininthesym-
bolstatemen t,tomakethemeasiertosee.)
340
Shouldn't weadjustforBaselineCD4count?
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1488.737 52.328with3DF(p=0.0001)
Score . . 43.477with3DF(p=0.0001)
Wald . . 43.680with3DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.198798 0.25747 0.59619 0.4400 1.220
RIF 10.837240 0.23598 12.58738 0.0004 2.310
CD4 1-0.019641 0.00367 28.59491 0.0001 0.981
Analysis ofMaximum Likelihood Estimates
Variable Label
CLARI 1=Clarithromycin arm,0otherwise
RIF 1=Rifabutin arm,0otherwise
CD4 CD4CellCount
IsCD4countaconfounder?
(Ananalysisstrati¯edbyCD4categorygavealmostidenticalre-
sults.OtherimportantcovariatesincludedCTG(clinicaltrials
group)andKarnofskystatus).
341
Whatdothedevianceresidualslooklikeversus
acontinuouscovariate,likeCD4?
D
e
v
i
a
n
c
e
R
e
s
i
d
u
a
l-2-101234
CD4 Cell Count0 100 200 300
Wemightwanttoconsidersomekindoftransformation ofCD4
count(likelogorsquareroot).Ifwedon'tfeelcomfortablewith
thelinearityofCD4count,wecanalsodichotomizeit(CD4CAT).
342
Anotherwayofcheckingtheproportionalityas-
sumptionisbyusingtheWeightedSchoenfeld
residualplotsforeachcovariate
RawCD4count logCD4count
-0.10-0.050.000.050.100.150.200.25
Time to MAC disease (days)0100200300400500600700800900-2-1012
Time to MAC disease (days)0 100200300400500600700800900
SquarerootCD4count
-0.8-0.6-0.4-0.20.00.20.40.60.81.01.2
Time to MAC disease (days)0100200300400500600700800900
343
Sofar,thegraphical techniqueshavenotindicated anyma-
jordeparture fromproportional hazards. However,wecan
testthisformally bycreating atimedependentcovariatefor
rifabutin andclarithrom ycin:
riftd=rif*((mactime-365)/30);
claritd=clari*((mactime-365)/30);
Eventhoughthedosereduction wasonlyforrifabutin, pa-
tientsonall3armshadtohavethedosereduction ...they
justtook2capsules oftheirplacebo,anddidn'tknowwhether
itwasplacebooractivedrug.
Ihavecenteredthetime-dep endentcovariatesat365days
(oneyear),sothattheHRforrifaloneandclarialonewill
applyatoneyear.ThenIhavedividedby30,sothatthe
resulting HRcanbeinterpreted asthechangeforeachmonth
awayfrom365days.
Question:Canwedothiswithinadatastepus-
ingtheabovestatements,ordothesestatements
needtobegiveninthePROCPHREGproce-
dure?
344
Time-dependentcovariatesforclariandrif
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1525.837 15.227with4DF(p=0.0043)
Score . . 16.033with4DF(p=0.0030)
Wald . . 15.327with4DF(p=0.0041)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.229811 0.25809 0.79287 0.3732 1.258
RIF 10.823227 0.23624 12.14274 0.0005 2.278
CLARITD 10.003065 0.04073 0.00566 0.9400 1.003
RIFTD 10.010627 0.03765 0.07965 0.7778 1.011
Analysis ofMaximum Likelihood Estimates
Variable Label
CLARI 1=Clarithromycin arm,0otherwise
RIF 1=Rifabutin arm,0otherwise
Neithertime-dependentcovariatewassigni¯cant.
345
Thisanalysis alsoindicated thattherearenomajorde-
partures fromproportionalhazardsforthethreetreatmen t
arms.
However,itmaystillbethecasethathavingthestudy-wide
dosereduction hadsomerelationship withMACdisease.
Wecanassessthisbycreating atimedependentvariablefor
theSWDR.
We'lllookatthefollowingmodels:
(1)SWDRST ATasasimpleindicator
(2)SWDRST ATandSWDRTD,with
swdrtd=swdrstat*((mactime-365)/30)
(3)SWDRastimedependentcovariate
346
Naivemodelwith¯xedSWDRindicator (SWDRST AT):
Dependent Variable: MACTIME TimetoMACdisease (days)
Censoring Variable: MACSTAT MACstatus(1=yes,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1495.857 45.208with3DF(p=0.0001)
Score . . 51.497with3DF(p=0.0001)
Wald . . 48.749with3DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.449936 0.26142 2.96236 0.0852 1.568
RIF 11.006639 0.23852 17.81114 0.0001 2.736
SWDRSTAT 1-1.125032 0.19283 34.04055 0.0001 0.325
Analysis ofMaximum Likelihood Estimates
Variable Label
CLARI 1=Clarithromycin arm,0otherwise
RIF 1=Rifabutin arm,0otherwise
SWDRSTAT StudyWideDoseReduction Status
Reduction ofdosagefrom450mgto300mgappearstobe
protective,whichseemscounter-intuitive
347
PredictedBaselineSurvivalCurves:
Another waytoseethisisthrough thepredicted baseline
survivalcurves.Thetwolinesareforthosenotonrifabutin,
whilethex'sand+'sareforthoseonrifabutin. Ineachcase,
thehigherline(betterprognosis) ofthepairisforthosewho
didhavetheSWDR.
SWDR/Rifabutin Status No SWDR/no RIF SWDR/no RIF
No SWDR/RIF SWDR/RIFP
r
(
s
u
r
v
i
v
a
l
)
0.00.20.40.60.81.0
Time to MAC (days)0 100200300400500600700800
348
Testforproportionality:
procphregdata=weighted;
modelmactime*macstat(0) =claririfswdrstat swdrtd;
***createtimebycovariate interaction forswdrstatus;
swdrtd=swdrstat*((mactime-365)/30);
test_trt: testclari,rif;
title'Testoftreatment Differences';
title2'andtestofproportionality att=365days';
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1492.692 48.372with4DF(p=0.0001)
Score . . 55.174with4DF(p=0.0001)
Wald . . 50.719with4DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.430051 0.26126 2.70947 0.0998 1.537
RIF 11.005416 0.23845 17.77884 0.0001 2.733
SWDRSTAT 1-1.126498 0.19752 32.52551 0.0001 0.324
SWDRTD 10.055550 0.03201 3.01112 0.0827 1.057
Variable Label
CLARI 1=Clarithromycin arm,0otherwise
RIF 1=Rifabutin arm,0otherwise
SWDRSTAT StudyWideDoseReduction Status
SWDRTD swdrstat*((mactime-365)/30)
349
InterpretationofHazardRatios
¯swdrstat=¡1:1265
¯swdrtd=0:0556
TimeTime Hazard
(months)(days)calculation Ratio
6182.5exp[¡1:1265+(¡6:08)(0:0556)]0.231
12365exp[¡1:1265+(0)(0:0556)]0.324
18547.5exp[¡1:1265+(6:08)(0:0556)]0.454
24730exp[¡1:1265+(12:17)(0:0556)]0.637
30912.5exp[¡1:1265+(18:25)(0:0556)]0.893
361095exp[¡1:1265+(24:33)(0:0556)]1.253
HR=exp[¯swdrstat+¯swdrtdÃmactime¡365)
30!
]
Intheearlyperiodafterrandomization totreatmen t,reduc-
tionofrandomized dosagefrom450mgto300mgisassoci-
atedwithadecreased riskofMACdisease.Aftertakingthe
higherdosageforabout32months,dropping tothelower
dosagehasnoimpact,andasthetreatmen ttimeincreases
beyond32months,alowerdosagetendstobeassociated
withincreased riskofMAC.
350
3di®erentwaystocodeSWDRastime-dependent
covariate
procphregdata=weighted;
modelmactime*macstat(0) =claririfswdr;
if(swdrtime>=mactime) thenswdr=0;
elsedo;
ifswdrstat=1 thenswdr=1;
elseswdr=0;
end;
test_trt: testclari,rif;
title2'I.Time-dependent indicator ofdosereduction';
procphregdata=weighted;
modelmactime*macstat(0) =claririfswdr;
ifswdrstat=0 or(swdrtime>=mactime) thenswdr=0;
elseswdr=1;
test_trt: testclari,rif;
title2'II.Time-dependent indicator ofdosereduction';
procphregdata=weighted;
modelmactime*macstat(0) =claririfswdr;
ifswdrstat=1 and(swdrtime<mactime) thenswdr=1;
elseswdr=0;
test_trt: testclari,rif;
title2'III.Time-dependent indicator ofdosereduction';
351
Outputisthesameforall3cases:
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
1178 121 1057 89.73
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1517.426 23.639with3DF(p=0.0001)
Score . . 24.844with3DF(p=0.0001)
Wald . . 24.142with3DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.328849 0.26017 1.59762 0.2062 1.389
RIF 10.905299 0.23775 14.49956 0.0001 2.473
SWDR 1-0.648887 0.21518 9.09389 0.0026 0.523
SWDRisstillprotective?Doesthismakesenseintuitively?
Whatothermethodscanweusetoaccountforchangein
dosage?
352
Weightedadjusteddose(WAD)analyses
Totrytogetabetterideaofthee®ectofchanging doses
ofrifabutin onthehazardforMACdisease,Icreatedthe
followingweighteddoseofrandomized rifabutin:
²Betweenrandomization dateandSWDRdate
=)#Daysat450mg
²BetweenSWDRdateando®-study date
=)#Daysat300mg
²Betweenrandomization dateandO®-study date
=)#TotalDays
²Weightedrandomized dose
rifwadr =(days450 +days300)/totdays
²Transformed tonumberofcapsules perday;
rifwadr=rifwadr/150;
²Alsocalculated weighteddosewhileontreatmentby
startingwithontreatmen tdate,stopping witho®-treatmen t
date,anddividing bythetotaldaysonstudy.
353
Weightedadjusteddose(WAD)analyses
Randomizedassignmenttorifabutin
Dependent Variable: MACTIME TimetoMACdisease (days)
Censoring Variable: MACSTAT MACstatus(1=yes,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
1178 121 1057 89.73
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1493.476 47.588with3DF(p=0.0001)
Score . . 52.770with3DF(p=0.0001)
Wald . . 50.295with3DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.453283 0.26119 3.01179 0.0827 1.573
RIF 11.004846 0.23826 17.78681 0.0001 2.731
RIFWADR 11.530462 0.25681 35.51502 0.0001 4.620
Foreachadditional capsuleofrifabutin speci¯edasran-
domizedtreatment,theHRforMACincreased by4.6
times
354
Weightedadjusteddose(WAD)analyses
Actualdosageofrifabutinduringthestudy
Dependent Variable: MACTIME TimetoMACdisease (days)
Censoring Variable: MACSTAT MACstatus(1=yes,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1489.993 51.071with3DF(p=0.0001)
Score . . 55.942with3DF(p=0.0001)
Wald . . 53.477with3DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.489583 0.26256 3.47693 0.0622 1.632
RIF 11.019675 0.23873 18.24291 0.0001 2.772
RIFWAD 1-0.664689 0.10686 38.69332 0.0001 0.514
Here,highervaluesofRIFWADprobably re°ectthatthe
patientwasabletostayontreatmentlonger,whichwas
protective.TheSWDRvariableisalsocapturing whethera
patienthadbeenabletotoleratethetreatmentlongenough
tohavethechancetohavetheprotocol-mandated dose
reduction.
355
Whathappensifweaddtreatmentdiscontinua-
tionasatimedependentcovariate?
Dependent Variable: MACTIME TimetoMACdisease (days)
Censoring Variable: MACSTAT MACstatus(1=yes,0=censored)
Censoring Value(s): 0
TiesHandling: BRESLOW
Summary oftheNumberof
EventandCensored Values
Percent
Total Event Censored Censored
1178 121 1057 89.73
Testing GlobalNullHypothesis: BETA=0
Without With
Criterion Covariates Covariates ModelChi-Square
-2LOGL 1541.064 1501.595 39.469with4DF(p=0.0001)
Score . . 42.817with4DF(p=0.0001)
Wald . . 41.027with4DF(p=0.0001)
Analysis ofMaximum Likelihood Estimates
Parameter Standard Wald Pr> Risk
Variable DF Estimate Error Chi-Square Chi-Square Ratio
CLARI 10.420447 0.26111 2.59284 0.1073 1.523
RIF 10.984114 0.23847 17.02975 0.0001 2.675
SWDR 1-0.139245 0.23909 0.33919 0.5603 0.870
RXSTOP 10.902592 0.21792 17.15473 0.0001 2.466
SWDRisnolongersigni¯cant!
356
Lastofall,acomparisonofsomeofthesemodels:
AIC
Modelterms q¡2logLCriterion
Clari, Rif 21525.931531.93
Clari, Rif,Cd4ca t 31497.571506.57
Clari, Rif,Cd4 31488.741497.74
Clari, Rif,Cd4ca t,Ctg,Karnof 51482.671497.67
Clari, Rif,Swdrst at 31495.861504.86
Clari, Rif,Rifwadr 31493.481502.48
Clari, Rif,Swdrst at,Rifwadr41493.441505.44
Clari, Rif,Rifwad 31489.991498.99
Modelswithtime-dependentcovariates
Clari, Rif,Claritd, Riftd 41525.841537.84
Clari, Rif,Swdrst at,Swdrtd41492.691504.69
Clari, Rif,Swdr 31517.431526.43
Clari, Rif,Swdr, Rxstop 41501.601513.60
Clari, Rif,Cd4ca t,Karnof, Rxstop51461.901476.90
Clari, Rif,Cd4ca t,Karnof, Rifwad51448.141463.14
357
Parametric SurvivalAnalysis
Sofar,wehavefocusedprimarily onnonparametric and
semi-parametric approachestosurvivalanalysis, withheavy
emphasis ontheCoxproportionalhazardsmodel:
¸(t;Z)=¸0(t)exp(¯Z)
Weusedthefollowingestimating approach:
²Weestimated¸0(t)nonparametrically ,usingtheKaplan-
Meierestimator, orusingtheKalb°eisc h/Prenticeesti-
matorunderthePHassumption
²Weestimated¯byassuming alinearmodelbetweenthe
logHRandcovariates,underthePHmodel
Bothestimates werebasedonmaximumlikelihoodtheory.
358
Thereareseveralreasonswhyweshouldconsider someal-
ternativeapproachesbasedonparametric models:
²Theassumption ofproportionalhazardsmightnotbe
appropriate (basedonmajordepartures)
²Ifaparametric modelactually holds,thenwewould
probably gaine±ciency
²Wemaywanttohandlenon-standard situations like
{intervalcensoring
{incorporatingpopulation mortality
²Wemaywanttomakesomeconnections withotherfa-
miliarapproaches(e.g.useofthePoissonlikelihood)
²Wemaywanttoobtainsomeestimates foruseindesign-
ingafuturesurvivalstudy.
359
Asimplestart:ExponentialRegression
²Observeddata: (Xi;±i;Zi)forindividuali,
Zi=(Zi1;Zi2;:::;Zip)representsasetofpcovariates.
²Rightcensoring: AssumethatXi=min(Ti;Ui)
²Survivaldistribution: AssumeTifollowsanexpo-
nentialdistribution withaparameter¸thatdependson
Zi,say¸i=ª(Zi).Thenwecanwrite:
Ti»exponential (ª(Zi))
First,let'sreviewsomefactsabouttheexponentialdistribu-
tion(fromour¯rstsurvivallecture):
f(t)=¸e¡¸tfort¸0
S(t)=P(T¸t)=Z1
tf(u)du=e¡¸t
F(t)=P(T<t)=1¡e¡¸t
¸(t)=f(t)
S(t)=¸constanthazard!
¤(t)=Zt
0¸(u)du=Zt
0¸du=¸t
360
Now,wesaythat¸isaconstantovertimet,butwewant
toletitdependonthecovariatevalues,sowearesetting
¸i=ª(Zi)
Thehazardratewouldtherefore bethesameforanytwo
individuals withthesamecovariatevalues.
Although therearemanypossiblechoicesforª,onesimple
andnaturalchoiceis:
ª(Zi)=exp[¯0+Zi1¯1+Zi2¯2+:::+Zip¯p]
WHY?
²ensuresapositivehazard
²foranindividual withZ=0,thehazardise¯0.
Themodeliscalledexponentialregression becauseof
thenaturalgeneralization fromregularlinearregression
361
Exponentialregressionforthe2-samplecase:
²AssumewehaveonlyasinglecovariateZ=Z,
i.e.,p=1.
HazardRate:
ª(Zi)=exp(¯0+Zi¯1)
²De¯ne:Zi=0ifindividualiisingroup0
Zi=1ifindividualiisingroup1
²Whatisthehazardforgroup0?
²Whatisthehazardforgroup1?
²Whatisthehazardratioofgroup1togroup
0?
²Whatistheinterpretationof¯1?
362
LikelihoodforExponentialModel
Undertheassumption ofrightcensored data,eachperson
hasoneoftwopossiblecontributions tothelikelihood:
(a)theyhaveaneventatXi(±i=1))contribution is
Li=S(Xi)
|{z}¢¸(Xi)
|{z}=e¡¸Xi¸
survivetoXifailatXi
(b)theyarecensored atXi(±i=0))contribution is
Li=S(Xi)
|{z}=e¡¸Xi
survivetoXi
Thelikelihoodistheproductoveralloftheindividuals:
L=Y
iLi
=Y
iµ
¸e¡¸Xi¶±i
|{z}µ
e¡¸Xi¶(1¡±i)
|{z}
eventscensorings
=Y
i¸±iµ
e¡¸Xi¶
363
MaximumLikelihoodforExponential
Howdoweusethelikelihood?
²¯rsttakethelog
²thentakethepartialderivativewithrespectto¯
²thensettozeroandsolveforc¯
²thisgivesusthemaximumlikelihoodestimators
Thelog-likelihoodis:
logL=log2
4Y
i¸±iµ
e¡¸Xi¶3
5
=X
i[±ilog(¸)¡¸Xi]
=X
i[±ilog(¸)]¡X
i¸Xi
Forthecaseofexponentialregression, wenowsubstitute the
hazard¸=ª(Zi)intheabovelog-likelihood:
logL=X
i[±ilog(ª(Zi))]¡X
iª(Zi)Xi(1)
364
GeneralFormofLog-likelihood
forRightCensored Data
Ingeneral,wheneverwehaverightcensored data,thelikeli-
hoodandcorrespondingloglikelihoodwillhavethefollowing
forms:
L=Y
i[¸i(Xi)]±iSi(Xi)
logL=X
i[±ilog(¸i(Xi))]¡X
i¤i(Xi)
where
²¸i(Xi)isthehazardfortheindividualiwhofailsatXi
²¤i(Xi)isthecumulativehazardforanindividual attheir
failureorcensoring time
Forexample, seethederivationofthelikelihoodforaCox
modelonp.11-13ofLecture4notes.Westartedwiththe
likelihoodabove,thensubstituted thespeci¯cformsfor¸(Xi)
underthePHassumption.
365
Consider ourmodelforthehazardrate:
¸=ª(Zi)=exp[¯0+Zi1¯1+Zi2¯2+:::+Zip¯p]
Wecanwritethisusingvectornotation, asfollows:
LetZi=(1;Zi1;:::Zip)T
and¯=(¯0;¯1;:::¯p)
(Since¯0istheintercept(i.e.,theloghazardrateforthe
baseline group),weputa\1"asthe¯rstterminthevector
Zi.)
Then,wecanwritethehazardas:
ª(Zi)=exp[¯Zi]
Nowwecansubstitute ª(Zi)=exp[¯Zi]inthelog-likelihood
shownin(1):
logL=nX
i=1±i(¯Zi)¡nX
i=1Xiexp(¯Zi)
366
ScoreEquations
Takingthederivativewithrespectto¯0,thescoreequation
is:
@logL
@¯0=nX
i=1[±i¡Xiexp(¯Zi)]
For¯k,k=1;:::p,theequations are:
@logL
@¯k=nX
i=1[±iZik¡XiZikexp(¯Zi)]
=nX
i=1Zik[±i¡Xiexp(¯Zi)]
To¯ndtheMLE's,wesettheaboveequations to0and
solve(simultaneously). Theequations aboveimplythat
theMLE'sareobtained bysettingtheweightednumberof
failures(P
iZik±i)equaltotheweightedcumulativehazard
(P
iZik¤(Xi)).
367
To¯ndthevarianceoftheMLE's,weneedtotakethesecond
derivatives:
¡@2logL
@¯k@¯j=nX
i=1ZikZijXiexp(¯Zi)
Somealgebra(seeCoxandOakessection6.2)revealsthat
Var(c¯)=I(¯)¡1=·
Z(I¡¦)ZT¸¡1
where
²Z=(Z1;:::;Zn)isa(p+1)£nmatrix
(pcovariatesplusthe\1"fortheintercept¯0)
²¦=diag(¼1;:::;¼n)(thismeansthat¦isadiagonal
matrix,withtheterms¼1;:::;¼nonthediagonal)
²¼iistheprobabilit ythatthei-thpersoniscensored, so
(1¡¼i)istheprobabilit ythattheyfailed.
²Note:Theinformation I(¯)(inverseofthevariance)
isproportionaltothenumberoffailures,notthesample
size.Thiswillbeimportantwhenwetalkaboutstudy
design.
368
TheSingleSampleProblem(Zi=1foreveryone):
First,whatistheMLEof¯0?
Weset@logL
@¯0=Pn
i=1[±i¡Xiexp(¯0Zi)]equalto0andsolve:
)nX
i=1±i=nX
i=1[Xiexp(¯0)]
d=exp(¯0)nX
i=1Xi
exp(d¯0)=d
Pni=1Xi
^¸=d
t
wheredisthetotalnumberofdeaths(orevents),andt=
PXiisthetotalperson-time contributed byallindividuals.
Ifd=tistheMLEfor¸,whatdoesthisimply
abouttheMLEof¯0?
369
Usingtheprevious formulaVar(^¯)=·
Z(I¡¦)ZT¸¡1,
whatisthevarianceofc¯0?:
Withsomematrixalgebra, youcanshowthatitis:
Var(c¯0)=1
Pni=1(1¡¼i)=1
d
Whatabout^¸=e^¯0?
Bythedeltamethod,
Var(^¸)=^¸2Var(c¯0)
=?
370
TheTwo-Sample Problem:
ZiSubjectsEventsFollow-up
Group0:Zi=0n0d0t0=Pn0i=1Xi
Group1:Zi=1n1d1t1=Pn1i=1Xi
Thelog-likelihood:
logL=nX
i=1±i(¯0+¯1Zi)¡nX
i=1Xiexp(¯0+¯1Zi)
so@logL
@¯0=nX
i=1[±i¡Xiexp(¯0+¯1Zi)]
=(d0+d1)¡(t0e¯0+t1e¯0+¯1)
@logL
@¯1=nX
i=1Zi[±i¡Xiexp(¯0+¯1Zi)]
=d1¡t1e¯0+¯1
Thisimplies: ^¸1=e^¯0+^¯1=?
^¸0=e^¯0=?
^¯0=?
^¯1=?
371
ImportantResult:
Themaximumlikelihoodestimates
(MLE's)ofthehazardratesunder
theexponentialmodelarethenum-
berofeventsdividedbytheperson-
yearsoffollow-up!
(thisresultwillbereliedonheavilywhenwedis-
cussstudydesign)
372
ExponentialRegression:
MeansandMedians
MeanSurvivalTime
Fortheexponentialdistribution, E(T)=1=¸.
²ControlGroup:
T0=1=^¸0=1=exp(^¯0)
²TreatmentGroup:
T1=1=^¸1=1=exp(^¯0+^¯1)
MedianSurvivalTime
ThisisthevalueMatwhichS(t)=e¡¸t=0:5,soM=
median=¡log(0:5)
¸
²ControlGroup:
^M0=¡log(0:5)
^¸0=¡log(0:5)
exp(^¯0)
²TreatmentGroup:
^M1=¡log(0:5)
^¸1=¡log(0:5)
exp(^¯0+^¯1)
373
ExponentialRegression:
VarianceEstimatesandTestStatistics
Wecanalsocalculate thevariances oftheMLE'sassimple
functions ofthenumberoffailures:
var(^¯0)=1
d0
var(^¯1)=1
d0+1
d1
Soourteststatistics areformedas:
FortestingHo:¯0=0:
Â2
w=µ^¯0¶2
var(^¯0)
=[log(d0=t0)]2
1=d0
FortestingHo:¯1=0:
Â2
w=µ^¯1¶2
var(^¯1)
=·
log(d1=t1
d0=t0)¸2
1
d0+1
d1
Howwouldweformcon¯dence intervalsforthehazard
ratio?
374
TheLikelihoodRatioTestStatistic:
(Analternativ etotheWaldtest)
Alikelihoodratiotestisbasedon2timesthelogoftheratio
ofthelikelihoodsunderthenullandalternativ e.Wereject
H0if2log(LR)>Â2
1;0:05,where
LR=L(H1)
L(H0)=L(b¸0;b¸1)
L(b¸)
Forasampleofnindependentexponentialrandomvariables
withparameter¸,theLikelihoodis:
L=nY
i=1[¸±iexp(¡¸xi)]
=¸dexp(¡¸Xxi)
=¸dexp(¡¸n¹x)
wheredisthenumberofdeathsorfailures.
Thelog-likelihoodis
`=dlog(¸)¡¸n¹x
andtheMLEis
b¸=d=(n¹x)
375
2-SampleCase:LRtestcalculations
Data:
Group0:d0failuresamongthen0females
meanfailuretimeis¹x0=(Pn0iXi)=n0
Group1:d1failuresamongthen1males
meanfailuretimeis¹x1=(Pn1iXi)=n1
Underthealternativehypothesis:
L=¸d11exp(¡¸1n1¹x1)£¸d00exp(¡¸0n0¹x0)
log(L)=d1log(¸1)¡¸1n1¹x1+d0log(¸0)¡¸0n0¹x0
TheMLE'sare:
b¸1=d1=(n1¹x1)formales
b¸0=d0=(n0¹x0)forfemales
Underthenullhypothesis:
L=¸d1+d0exp[¡¸(n1¹x1+n0¹x0)]
log(L)=(d1+d0)log(¸)¡¸[n1¹x1+n0¹x0]
ThecorrespondingMLEis
b¸=(d1+d0)=[n1¹x1+n0¹x0]
376
Alikelihoodratiotestcanbeconstructed bytakingtwicethe
di®erence ofthelog-likelihoodsunderthealternativ eandthe
nullhypotheses:
¡22
4(d0+d1)log0
@d0+d1
t0+t11
A¡d1log[d1=t1]¡d0log[d0=t0]3
5
Nursinghomeexample:
Forthefemales:
²n0=1173
²d0=902
²t0=310754
²¹x0=265
Forthemales:
²n1=418
²d1=367
²t1=75457
²¹x1=181
Plugging thesevaluesin,wegetaLRteststatisticof64.20.
377
HandCalculationsusingeventsandfollow-up:
Byaddingup\los"formalestogett1andforfemalesto
gett0,Iobtained:
²d0=902(females)
d1=367(males)
²t0=310754(femalefollow-up)
t1=75457(malefollow-up)
²ThisyieldsanestimatedlogHR:
^¯1=log2
4d1=t1
d0=t03
5=log2
4367=75457
902=3107543
5=log(1:6756)=0:5162
²Theestimatedstandarderroris:
r
var(^¯1)=vuut1
d1+1
d0=vuut1
902+1
367=0:06192
²SotheWaldtestbecomes:
Â2
W=^¯2
1
var(^¯1)=(0:51619)2
0:061915=69:51
²Wecanalsocalculate^¯0=log(d0=t0)=¡5:842,
alongwithitsstandarderrorse(^¯0)=q
(1=d0)=0:0333
378
ExponentialRegressioninSTATA
.usenurshome
.stsetlosfail
.streggender, dist(exp) nohr
failure _d:fail
analysis time_t:los
Iteration 0:loglikelihood =-3352.5765
Iteration 1:loglikelihood =-3321.966
Iteration 2:loglikelihood =-3320.4792
Iteration 3:loglikelihood =-3320.4766
Iteration 4:loglikelihood =-3320.4766
Exponential regression --logrelative-hazard form
No.ofsubjects = 1591 Numberofobs=1591
No.offailures = 1269
Timeatrisk = 386211
LRchi2(1) =64.20
Loglikelihood =-3320.4766 Prob>chi2 =0.0000
------------------------------------------------------------------- ------
_t|Coef.Std.Err. zP>|z| [95%Conf.Interval]
---------|--------------------------------------------------------- -----
gender|.516186 .0619148 8.337 0.000 .3948352 .6375368
_cons|-5.842142 .0332964 -175.459 0.000 -5.907402 -5.776883
------------------------------------------------------------------- ------
SinceZ=8:337,thechi-squaretestisZ2=69:51.
379
ExponentialRegressioninSAS-proclifereg
procformat;
valuecensfmt 1='Censored'
0='Dead';
valuegrpfmt 0='Group 0(F)'
1='Group 1(M)';
Title'Exponential HazardModelforNursing HomePatients';
datamorris;
infile'ch12.dat';
inputlosagetrtgendermarstat hltstat cens;
datamorris2;
setmorris;
iflos=0thendelete;
procfreqdata=morris2;
tablecens*gender/ norownocolnopercent;
formatcenscensfmt. gendergrpfmt.;
proclifereg data=pop covoutoutest=survres;
modellos*censor(1)=gender /dist=exponential;
run;
RESUL TS:
TABLEOFCENSBYGENDER
CENS GENDER
Frequency|Group 0|Group1|Total
|(F) |(M) |
---------+--------+--------+
Event |902|367|1269
---------+--------+--------+
Censored |271|51|322
---------+--------+--------+
Total 1173 418 1591
380
PROCLIFEREGRESULTS:
Exponential HazardModelforNursing HomePatients
Lifereg Procedure
DataSet =WORK.MORRIS2
Dependent Variable=Log(LOS)
Censoring Variable=CENS
Censoring Value(s)= 1
Noncensored Values= 1269RightCensored Values= 322
LeftCensored Values= 0Interval Censored Values= 0
LogLikelihood forEXPONENT -3320.476626
Lifereg Procedure
Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value
INTERCPT 15.84213388 0.033296 307860.0001Intercept
GENDER 1-0.5161878 0.061915 69.50734 0.0001
SCALE 0 1 0 Extreme valuescale
Notethattheestimatesfor¯0and¯1aboveare
theoppositesofwhatwecalculated.I'llexplain
whytheoutputhasthisformwhenwegetto
AFTmodels.
381
TheWeibullRegressionModel
Atthebeginning ofthecourse,wesawthatthesurvivorship
function foraWeibullrandomvariableis:
S(t)=exp[¡¸(t·)]
andthehazardfunction is:
¸(t)=·¸t(·¡1)
TheWeibullregression modelassumes thatforsomeone with
covariatesZi,thesurvivorshipfunction is
S(t;Zi)=exp[¡ª(Zi)(t·)]
whereª(Zi)isde¯nedasinexponentialregression tobe:
ª(Zi)=exp[¯0+Zi1¯1+Zi2¯2+:::Zip¯p]
Forthe2-sample problem, wehave:
ª(Zi)=exp[¯0+Zi1¯1]
382
WeibullMLEsforthe2-sampleproblem:
Log-likelihood:
logL=nX
i=1±ilogh
·exp(¯0+¯1Zi)X·¡1
ii
¡nX
i=1X·
iexp(¯0+¯1Zi)
)exp(^¯0)=d0=t0·
exp(^¯0+^¯1)=d1=t1·
wheretj·=njX
i=1X^·
iamongnjsubjects
^¸0(t)=^·exp(^¯0)t^·¡1
^¸1(t)=^·exp(^¯0+^¯1)t^·¡1
dHR=^¸1(t)=^¸0(t)=exp(^¯1)
=exp0
B@d1=t1·
d0=t0·1
CA
383
WeibullRegression:
MeansandMedians
MeanSurvivalTime
FortheWeibulldistribution, E(T)=¸(¡1=·)¡[(1=·)+1].
²ControlGroup:
T0=^¸(¡1=^·)
0¡[(1=^·)+1]
²TreatmentGroup:
T1=^¸(¡1=^·)
1¡[(1=^·)+1]
MedianSurvivalTime
FortheWeibulldistribution, M=median=·¡log(0:5)
¸¸1=·
²ControlGroup:
^M0=2
64¡log(0:5)
^¸03
751=^·
²TreatmentGroup:
^M1=2
64¡log(0:5)
^¸13
751=^·
where^¸0=exp(^¯0)and^¸1=exp(^¯0+^¯1).
384
Note:thesymbol¡isthe\gamma" function. Ifxisan
integer,then
¡(x)=(x¡1)!
Incaseswherexisnotaninteger,thisfunction hastobe
evaluatednumerically .
TheWeibullregression modelisveryeasyto¯t:
²Insas:usemodeloptiondist=weibull withinthe
proclifereg procedure
²Instata:Justspecifydist(weibull) instead
ofdist(exp) withinthestregcommand
Note:togetmoreinformation onthesemodelingprocedures,
usetheonlinehelpfacilities. Forexample, inStata,you
cantype:
.helpstreg
385
WeibullinStata:
.streggender, dist(weibull) nohr
failure _d:fail
analysis time_t:los
Fitting constant-only model:
Iteration 0:loglikelihood =-3352.5765
Iteration 1:loglikelihood =-3074.978
Iteration 2:loglikelihood =-3066.1526
Iteration 3:loglikelihood =-3066.143
Iteration 4:loglikelihood =-3066.143
Fitting fullmodel:
Iteration 0:loglikelihood =-3066.143
Iteration 1:loglikelihood =-3045.8152
Iteration 2:loglikelihood =-3045.2772
Iteration 3:loglikelihood =-3045.2768
Iteration 4:loglikelihood =-3045.2768
Weibull regression --logrelative-hazard form
No.ofsubjects = 1591 Numberofobs=1591
No.offailures = 1269
Timeatrisk = 386211
LRchi2(1) =41.73
Loglikelihood =-3045.2768 Prob>chi2 =0.0000
------------------------------------------------------------------- -----
_t|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----
gender|.4138082 .0621021 6.663 0.000.2920903 .5355261
_cons|-3.536982 .0891809 -39.661 0.000-3.711773 -3.362191
---------+--------------------------------------------------------- -----
/ln_p|-.4870456 .0232089 -20.985 0.00-.5325343 -.4415569
------------------------------------------------------------------- -----
p|.614439 .0142605 .5871152 .6430345
1/p|1.627501 .0377726 1.555127 1.703243
------------------------------------------------------------------- -----
386
WeibullinSAS
proclifereg data=morris2 covoutoutest=survres;
modellos*censor(1)=gender /dist=weibull;
run;
DataSet =WORK.MORRIS2
Dependent Variable=Log(LOS)
Censoring Variable=CENS
Censoring Value(s)= 1
Noncensored Values= 1269 RightCensored Values= 322
LeftCensored Values= 0 Interval Censored Values= 0
LogLikelihood forWEIBULL -3045.276811
Lifereg Procedure
Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value
INTERCPT 15.75644118 0.0542 11280.04 0.0001Intercept
GENDER 1-0.6734732 0.101067 44.40415 0.0001
SCALE 11.62750085 0.037773 Extreme valuescale
387
InSAS,boththeexponentialandWeibullarespecialcases
ofthegeneralclassofacceleratedlifemodelsandthe
parameter interpretations followfromthisapproach.
Totranslate theoutputofSAS(orStatausingtheereg
command) forWeibullregression, wehavetotakethenega-
tiveofthenumbersintheoutput,dividedbythe\scale"
parameter (¾,or1=·).
²^¯0=¡intercpt=scale
²^¯1=¡covariate=scale
Thenwecalculate theestimated HRasexp(^¯1).
TheMLE'sare:
²^¯0=¡intercpt=scale=¡5:756=1:627=¡3:537
²^¯1=¡covariate=scale=0:6735=1:625=0:414
andtheestimated HRisdHR=exp(^¯1)=exp(0:414)=
1:513.
388
WeibullRegression:
VarianceEstimatesandTestStatistics
Itisnotsoeasytogetvarianceestimates fromtheoutput
ofproclifereg inSASorweibullinstata,atleastfor
theparameters we'reinterested in.
Thevariances dependoninvertinga(3£3)matrixcorre-
spondingtotheparameters ¯0,¯1,and·.TheMLEfor^·
hastobeobtained numerically (i.e.,noclosedform),sothe
standard errorsalsohavetobeobtained bycomputer.
Mainobjective:toobtains:e:(^¯1),sothatwecanform
testsandcon¯dence intervalsforthehazardratio.
Theoutputgivesus^¯¤
1ands:e:(^¯¤
1),where^¯1=¡^¯¤
1=^¾.If
¾wasaconstant,thenwecouldjustcompute
var(^¯1)=1
^¾2var(^¯¤
1)
but¾isalsoarandomvariable! Instead, youneedtouse
anapproximation forthevarianceofaratiooftworandom
variables:
var(^¯1)=1
^¾4·
^¾2var(^¯¤
1)+(^¯¤
1)2var(^¾)¡2^¯¤
1^¾cov(^¯¤
1;^¾)¸
whereyougetvar(^¯¤
1)andvar(^¾)bysquaring thestandard
errorsofthecovariatetermandscaleterm,respec-
tively,fromtheproclifereg orweibulloutput.
389
ComparisonofExponentialwithKaplan-Meier
WecanseehowwelltheExponentialmodel¯tsbycompar-
ingthesurvivalestimates formalesandfemalesunderthe
exponentialmodel,i.e.,P(T¸t)=e(¡^¸zt),totheKaplan-
Meiersurvivalestimates:
S
u
r
v
i
v
a
l
0.00.10.20.30.40.50.60.70.80.91.0
Length of Stay (days)010020030040050060070080090010001100
390
ComparisonofWeibullwithKaplan-Meier
WecanseehowwelltheWeibullmodel¯tsbycomparing
thesurvivalestimates,P(T¸t)=e(¡^¸zt^·),totheKaplan-
Meiersurvivalestimates.
S
u
r
v
i
v
a
l
0.00.10.20.30.40.50.60.70.80.91.0
Length of Stay (days)010020030040050060070080090010001100
Whichdoyouthink¯tsbest?
391
Otherusefulplotsforevaluating¯ttoexponen-
tialandWeibullmodels
²¡log(^S(t))vst
²log[¡log(^S(t))]vslog(t)
Whyaretheseuseful?
IfTisexponential,thenS(t)=exp(¡¸t))
solog(S(t))=¡¸t
and ¤(t)=¸t
astraightlineintwithslope¸andintercept=0
IfTisWeibull,thenS(t)=exp(¡(¸t)·)
solog(S(t))=¡¸t·
then ¤(t)=¸t·
and log(¡log(S(t)))=log(¸)+·¤log(t)
astraightlineinlog(t)withslope·andinterceptlog(¸).
392
Sowecancalculate ourestimated ¤(t)andplotitversust,
andifitseemstoformastraightline,thentheexponential
distribution isprobably appropriate forourdataset.
Plotsfornursinghomedata:^¤(t)vstNegative Log SDF
0.00.20.40.60.81.01.21.41.61.82.0
LOS0100200300400500600700800900100011001200
393
Orwecanplotlog^¤(t)versuslog(t),andifitseemsto
formastraightline,thentheWeibulldistribution isprobably
appropriate forourdataset.
Plotsfornursinghomedata:log[¡log(^S(t))]vslog(t)Log Negative Log SDF
-0.5-0.4-0.3-0.2-0.10.00.10.20.30.40.50.60.7
Log of LOS4.504.755.005.255.505.756.006.256.506.757.007.25
394
ComparisonofMethods
fortheTwo-sampleproblem:
Data:
ZiSubjectsEventsFollow-up
Group0:Zi=0n0d0t0=Pn0i=1Xi
Group1:Zi=1n1d1t1=Pn1i=1Xi
InGeneral:
¸z(t)=¸(t;Z=z)forz=0or1:
ThehazardratedependsonthevalueofthecovariateZ.
Inthiscase,weareassuming thatweonlyhaveasingle
covariate,anditisbinary(Z=1orZ=0)
395
MODELS
ExponentialRegression:
¸z(t)=exp(¯0+¯1Z)
)¸0=exp(¯0)
¸1=exp(¯0+¯1)
HR=exp(¯1)
WeibullRegression:
¸z(t)=·exp(¯0+¯1Z)t·¡1
)¸0=·exp(¯0)t·¡1
¸1=·exp(¯0+¯1)t·¡1
HR=exp(¯1)
ProportionalHazardsModel:
¸z(t)=¸0(t)exp(¯1)
)¸0=¸0(t)
¸1=¸0(t)exp(¯1)
HR=exp(¯1)
396
Remarks
²ExponentialmodelisaspecialcaseoftheWeibullmodel
with·=1(note:Collettuses°insteadof·)
²ExponentialandWeibullmodelsarebothspecialcases
oftheCoxPHmodel.
Howcanyoushowthis?
²IfeithertheexponentialmodelortheWeibullmodelis
valid,thenthesemodelswilltendtobemoree±cient
thanPH(smaller s.e.'sofestimates). Thisisbecause
theyassumeaparticular formfor¸0(t),ratherthanes-
timating itateverydeathtime.
397
FortheExponentialmodel,thehazards areconstantover
time,giventhevalueofthecovariateZi:
Zi=0)^¸0=exp(^¯0)
Zi=1)^¸0=exp(^¯0+^¯1)
FortheWeibullmodel,wehavetoestimate thehazardasa
function oftime,giventheestimates of¯0;¯1and·:
Zi=0)^¸0(t)=^·exp(^¯0)t^·¡1
Zi=1)^¸1(t)=^·exp(^¯0+^¯1)t^·¡1
However,theratioofthehazardsisstilljustexp(^¯1),since
theothertermscancelout.
398
Here'swhattheestimatedhazardslooklikefor
thenursinghomedata:
Exponential Hazard: Female
Exponential Hazard: Male
Weibull Hazard: Female
Weibull Hazard: MaleH
a
z
a
r
d
R
a
t
e
0.0000.0050.0100.0150.0200.0250.030
Length of stay (days)01002003004005006007008009001000
399
ComparisonwithProportionalHazardsModel
.stcoxgender, nohr
failure _d:fail
analysis time_t:los
Iteration 0:loglikelihood =-8556.5713
Iteration 1:loglikelihood =-8537.8013
Iteration 2:loglikelihood =-8537.5605
Iteration 3:loglikelihood =-8537.5604
Refining estimates:
Iteration 0:loglikelihood =-8537.5604
Coxregression --Breslow methodforties
No.ofsubjects = 1591 Numberofobs=1591
No.offailures = 1269
Timeatrisk = 386211
LRchi2(1) =38.02
Loglikelihood =-8537.5604 Prob>chi2 =0.0000
------------------------------------------------------------------- ----
_t|
_d|Coef. Std.Err. zP>|z|[95%Conf.Interval]
---------+--------------------------------------------------------- ----
gender|.3943588 .0621004 6.350 0.000.2726441 .5160734
------------------------------------------------------------------- ----
ForthePHmodel,^¯1=0:394anddHR=e0:394=1:483.
400
ComparisonwiththeLogrankandWilcoxonTests
.ststestgender
failure _d:fail
analysis time_t:los
Log-rank testforequality ofsurvivor functions
------------------------------------------------
|Events
gender|observed expected
-------+-------------------------
0| 902 995.40
1| 367 273.60
-------+-------------------------
Total|1269 1269.00
chi2(1) =41.08
Pr>chi2 =0.0000
.ststestgender, wilcoxon
failure _d:fail
analysis time_t:los
Wilcoxon (Breslow) testforequality ofsurvivor functions
----------------------------------------------------------
|Events Sumof
gender|observed expected ranks
-------+--------------------------------------
0| 902 995.40 -99257
1| 367 273.60 99257
-------+--------------------------------------
Total|1269 1269.00 0
chi2(1) =41.47
Pr>chi2 =0.0000
401
ComparisonofHazardRatiosandTestStatistics
fore®ectofGender
Wald
Model/Metho d¸0¸1HRlog(HR) se(logHR)Statistic
Exponential0.00290.00491.6760.5162 0.0619 69.507
Weibull
t=50 0.00400.00601.5130.4138 0.0636 42.381
t=100 0.00300.00461.513
t=500 0.00160.00251.513
Logrank 41.085
Wilco xon 41.468
CoxPH
Ties=Breslo w 1.4830.3944 0.0621 40.327
Ties=Discrete 1.4870.3969 0.0623 40.565
Ties=Efron 1.4860.3958 0.0621 40.616
Ties=Exact 1.4860.3958 0.0621 40.617
Score(Discrete) 41.085
402
ComparisonofMeanandMedianSurvival
TimesbyGender
MeanSurvivalMedianSurvival
Model/MethodFemaleMaleFemaleMale
Exponential 344.5205.6238.8142.5
Weibull 461.6235.4174.288.8
Kaplan-Meier 318.6200.714470
CoxPH 13172
(Kalb°eisch/Prentice)
403
TheAcceleratedFailureTimeModel
Thegeneralformofanaccelerated failuretime(AFT)model
is:
log(Ti)=¯AFTZi+¾²
where
²log(Ti)isthelogofasurvivaltime
²¯AFTisthevectorofAFTmodelparameters corre-
spondingtothecovariatevectorZi
²²isarandom\error"term
²¾isascalefactor
Inotherwords,wecanmodelthelog-survival
timesasalinearfunctionofthecovariates.
proclifereg inSASandthestregcommand instata
(without theexponentialorweibulloption)allusethis\log-
linear"modelformulationfor¯ttingparametric models.
404
Bychoosingdi®erentdistributions for²,wecanobtaindif-
ferentparametric distributions:
²Exponential
²Weibull
²Gamma
²Log-logistic
²Normal
²Lognormal
Wecancompare thepredicted survivalunderanyofthese
parametric distributions totheKMestimated survivaltosee
whichoneseemsto¯tbest.
Oncewedecideonacertainclassofmodel(say,Gamma),
wecanevaluatethecontributions ofcovariatesby¯nding
theMLE's,andconstructing Wald,Score,orLRtestsofthe
covariatee®ects.
405
WecanmotivatetheAFTmodelby¯rstdemonstrating the
followingtworelationships:
²1.FortheExponentialModel:
IfthefailuretimesTi=T(Zi)followanexponential
distribution, i.e.,Si(t)=e¡¸itwith¸i=exp(¯Zi),
then
log(Ti)=¡¯Zi+²
where²followsanextremevaluedistribution (whichjust
meansthate²followsaunitexponentialdistribution).
²2.FortheWeibullModel:
IfthefailuretimesTi=T(Zi)followaWeibulldistri-
bution,i.e.,Si(t)=e¸it·with¸i=exp(¯Zi),then
log(Ti)=¡¾¯Zi+¾²
where²againfollowsanextreme valuedistribution, and
¾=1=·.
Inotherwords,boththeExponentialandWeibullmodelcan
bewrittenintheformofalog-linear modelforthesurvival
times,ifwechoosetherightdistribution for².
406
Thelog-linear formfortheexponentialcanbederivedby:
(1)Creating anewvariableT0=TZ£exp(¯Zi)
(2)TakingthelogofTZ,yieldinglog(TZ)=logÃ
T0
exp(¯Zi)!
Step(1):Foranexponentialmodel,recallthat:
Si(t)=Pr(TZ¸t)=e¡¸t;with¸=exp(¯Zi)
ItfollowsthatT0»exp(1):
S0(t)=Pr(T0¸t)=Pr(TZ¢exp(¯Z)¸t)
=Pr(TZ¸texp(¡¯Z))
=exp[¡¸texp(¡¯Z)]
=exp[¡exp(¯Z)texp(¡¯Z)]
=exp(¡t)
Step(2):Nowtakethelogofthesurvivaltime:
log(TZ)=log0
B@T0
exp(¯Zi)1
CA
=log(T0)¡log(exp(¯Zi))
=¡¯Zi+log(T0)
=¡¯Zi+²
where²=log(T0)followstheextremevaluedistribution.
407
RelationshipbetweenExponentialandWeibull
IfTZhasaWeibulldistribution, i.e.,S(t)=e¡¸t·
with¸=exp(¯Zi),thenyoucanshowthatthenewvariable
T¤
Z=T·
Z
followsanexponentialdistribution withparameter exp(¯Zi).
Basedontheprevious page,wecantherefore write:
log(T¤)=¡¯Z+²
(where²hasanextreme valuedistribution.)
Butsincelog(T¤)=log(T·)=·£log(T),wecanwrite:
log(T)=log(T¤)=·
=(1=·)(¡¯Zi+²)
=¡¾¯Zi+¾²
where¾=1=·.
408
Thismotivatesthefollowinggeneralde¯nition ofthe
AcceleratedFailureTimeModelby:
log(Ti)=¯AFTZi+¾²
where²isarandom\error"term,¾isascalefactor,Yis
thelogofasurvivalrandomvariable,and
¯AFT=¡¾¯e
where¯ecamefromthehazard¸=exp(¯Z).
Thede¯ningfeatureofanAFTmodelis:
S(t;Z)=Si(t)=S0(Át)
Thatis,thee®ectofcovariatesistoaccelerate
(stretch)ordecelerate (shrink)thetime-scale.
E®ectofAFTonhazard:
¸i(t)=Á¸0(Át)
409
OnewaytointerprettheAFTmodelisviaitse®ecton
mediansurvivaltimes.IfSi(t)=0:5,thenS0(Át)=0:5.
Thismeans:
Mi=ÁM0
Interpretation:
²ForÁ<1,thereisanacceleration oftheendpoint
(ifM0=2yrsincontrolandÁ=0:5,thenMi=1yr.
²ForÁ>1,thereisastretchingordelayinendpoint
²Ingeneral, thelifetimeofindividualiisÁtimeswhat
theywouldhaveexperienced inthereference group
SinceÁmustbepositiveandafunction ofthecovariates,we
modelÁ=exp(¯Zi).
410
WhendoesProportionalhazards=AFT?
According totheproportionalhazardsmodel:
S(t)=S0(t)exp(¯Zi)
andaccording totheaccelerated failuretimemodel:
S(t)=S0(texp(¯Zi))
SayTi»Weibull(¸;·).Then¸(t)=¸·t(·¡1)
UndertheAFTmodel:
¸i(t)=Á¸0(Át)
=e¯Zi¸0(e¯Zit)
=e¯Zi¸0·Ã
e¯Zit!(·¡1)
=Ã
e¯Zi!·
¸0·t(·¡1)
=Ã
e¯Zi!·
¸0(t)
ButthislooksjustlikethePHmodel:
¸i(t)=exp(¯¤Zi)¸0(t)
ItturnsoutthattheWeibulldistribution (andexponential,
sincethisisjustaspecialcaseofaWeibullwith·=1)
istheonlyoneforwhichtheaccelerated failuretimeand
proportionalhazardsmodelscoincide.
411
SpecialcasesofAFTmodels
²Exponentialregression: ¾=1,²followingtheextreme
valuedistribution.
²Weibullregression:¾arbitrary ,²followingtheextreme
valuedistribution.
²Lognormal regression: ¾arbitrary ,²followingthenor-
maldistribution.
Examplesinstata:Usingthestregcommand, one
hasthefollowingoptionsofdistributions forthelog-surviv al
times:
.stregtrt,dist(lognormal)
²exponential
²weibull
²gompertz
²lognormal
²loglogistic
²gamma
412
.streggender, dist(exponential) nohr
------------------------------------------------------------------- -----------
_t|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
gender|.516186 .0619148 8.337 0.000 .3948352 .6375368
------------------------------------------------------------------- -----------
.streggender, dist(weibull) nohr
------------------------------------------------------------------- -----------
_t|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
gender|.4138082 .0621021 6.663 0.000 .2920903 .5355261
1/p|1.627501 .0377726 1.555127 1.703243
------------------------------------------------------------------- -----------
.streggender, dist(lognormal)
------------------------------------------------------------------- -----------
_t|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
gender|-.6743434 .1127352 -5.982 0.000 -.8953002 -.4533866
_cons|4.957636 .0588939 84.179 0.000 4.842206 5.073066
sigma|1.94718 .040584 1.86924 2.028371
------------------------------------------------------------------- -----------
.streggender, dist(gamma)
------------------------------------------------------------------- -----------
_t|Coef. Std.Err. zP>|z| [95%Conf.Interval]
---------+--------------------------------------------------------- -----------
gender|-.6508469 .1147116 -5.674 0.000 -.8756774 -.4260163
_cons|4.788114 .1020906 46.901 0.000 4.58802 4.988208
sigma|1.97998 .0429379 1.897586 2.065951
------------------------------------------------------------------- -----------
413
Thisgivesagoodideaofthesensitivit yofthetestofgender
tothechoiceofmodel.Itisalsoeasytogetpredicted sur-
vivalcurvesunderanyoftheparametric modelsusingthe
following:
.streggender, dist(gamma)
.stcurv, survival
Theoptions hazard andcumhaz canalsobesubstituted
forsurvivalabovetoobtainplots.
414
AFTmodelsinSAS
proclifereg data=pop covout outest=survres;
modellos*censor(1)=gender /dist=exponential;
modellos*censor(1)=gender /dist=weibull;
modellos*censor(1)=gender /dist=gamma;
modellos*censor(1)=gender /dist=normal;
Otheroptionsarelognormal, logistic,andlog-logistic. The
defaultistomodellogofresponse.Canspecify"NOLOG"
fornolog-transformation. Inthiscase,"normal" isthesame
as"lognormal."
415
Lifereg Procedure
LogLikelihood forEXPONENT -3320.476626
Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value
INTERCPT 15.84213388 0.033296 307860.0001Intercept
GENDER 1-0.5161878 0.061915 69.50734 0.0001
SCALE 0 1 0 Extreme valuescale
Lagrange Multiplier ChiSquare forScale337.5998 Pr>Chiis0.0001.
LogLikelihood forWEIBULL -3045.276811
Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value
INTERCPT 15.75644118 0.0542 11280.04 0.0001Intercept
GENDER 1-0.6734732 0.101067 44.40415 0.0001
SCALE 11.62750085 0.037773 Extreme valuescale
LogLikelihood forGAMMA-2970.388508
Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value
INTERCPT 14.78811071 0.104333 2106.114 0.0001Intercept
GENDER 1-0.6508468 0.114748 32.17096 0.0001
SCALE 11.97998063 0.043107 Gammascaleparameter
SHAPE 1-0.1906006 0.094752 Gammashapeparameter
LogLikelihood forNORMAL-9593.512838
Variable DFEstimate StdErrChiSquare Pr>ChiLabel/Value
INTERCPT 1303.824624 9.919629 938.1129 0.0001Intercept
GENDER 1-107.09585 18.97784 31.84577 0.0001
SCALE 1330.093584 6.918237 Normalscaleparameter
416
Designing aSurvivalStudy
Wewillfocusonthepoweroftestsbasedontheexponential
distribution andthelogranktest.
²Asinstandard designs,thepowerdependson
{TheTypeIerror(signi¯cance level)
{Thedi®erence ofinterest,¢,underHa.
²Anotabledi®erence fromtheusualscenarioisthatpower
dependsonthenumberoffailures ,notthetotal
samplesize.
²Inpractice, designing asurvivalstudyinvolvesdeciding
howmanypatientsorindividuals toenter,aswellashow
longtheyshouldbefollowed.
²Designsmaybe¯xedsamplesizeorsequential
(Moreonthislater!)
References:
Collett Chapter 12
PocockChapter 9ofClinicalTrials
Williams Chapter 10ofAIDSClinicalTrials
(eds.FinkelsteinandSchoenfeld)
417
Reviewofpowercalculationsfor2-samplenormal
Supposewehavethefollowingdata:
Group1:(Y11;:::Y1n1)
Group0:(Y01;:::Y0n0)
andmakethefollowingassumptions:
Y1j»N(¹1;¾2)Y0j»N(¹0;¾2)
Ourobjectiveistotest:
H0:¹1=¹0)H0:4=0where4=¹1¡¹0
Thestandard testisbasedontheZstatistic:
Z=Y1¡Y0r
s2(1
n1+1
n0)
wheres2isthepooledsamplevariance(weareassuming
equalvariances here).Thisteststatistic followsaN(0;1)
distribution underH0.
Ifthesamplesizesareequalinthetwoarms,n0=n1=n=2,
(whichwillmaximize thepower),thenwehavethesimpler
form:
Z=Y1¡Y0s
s2(1
n=2+1
n=2)=Y1¡Y0
2s=pn
418
Thestepstofollowincalculating thesamplesizeare:
(1)Determine thecriticalvalue,c,forrejecting thenull
whenitistrue.
(2)Calculate theprobabilit yofrejecting thenullwhenthe
alternativ eistrue,substituting cfromabove.
(3)Rewritetheexpression intermsofthesamplesizefora
givenpower.
Step(1):
Setthesigni¯cance level,®,equaltotheprobabilit yofre-
jectingthenullhypothesiswhenitistrue:
®=Pr(jY1¡Y0j>cjH0)
=Pr0
B@jY1¡Y0j
2s=pn>c
2s=pnjH01
CA
=Pr0
B@jZj>c
2s=pn1
CA=2¢©0
B@c
2s=pn1
CA
soz1¡®=2=c
2s=pn
orc=z1¡®=22spn
Notethatz°isthevaluesuchthat©(z°)=Pr(Z<z°)=
°.
419
Step(2):
Calculate theprobabilit yofrejecting thenullwhenHais
true.Startoutbywritingdowntheprobabilit yofaTypeII
error:
¯=Pr(acceptH0jHa)
so1¡¯=Pr(rejectH0jHa)
=Pr(jY1¡Y0j>cjHa)
=Pr0
B@jY1¡Y0j¡¢
2s=pn>c¡¢
2s=pnjHa1
CA
=Pr0
B@Z>c¡¢
2s=pn1
CA
sowegetz¯=¡z1¡¯=c¡¢
2s=pn
NowwesubstitutecfromStep(1):
¡z1¡¯=z1¡®=22s=pn¡¢
2s=pn
=z1¡®=2¡¢
2s=pn
420
Step(3):
Nowrewritetheequation intermsofsamplesizeforagiven
power,1¡¯,andsigni¯cance level,®:
z1¡®=2+z1¡¯=¢
2s=pn
=¢pn
2s
=)n=(z1¡®=2+z1¡¯)24s2
¢2
Notes:
Thepowerisanincreasing function ofthestandardized dif-
ference:
¹T(4)=4
2s=pn
Thisisjustthenumberofstandard errorsbetweenthetwo
means,undertheassumption ofequalvariances.
1.Asnincreases, thepowerincreases.
2.For¯xedn,thepowerincreases with4.
3.For¯xednand4,thepowerdecreases withs.
4.Assigning equalnumbersofpatientstothetwogroups
(n1=n0=n=2)isbestintermsofmaximizing power.
421
AnExample:
n=µ
z1¡®
2+z1¡¯¶24s2
42
Saywewanttoderivethetotalsamplesizerequired toyield
90%powerfordetecting adi®erence of0.5standard devia-
tionsbetweenmeans,basedonatwo-sided0.05leveltest.
®=0.05
z1¡®
2=1.96
¯=0.10
z1¡¯=z0:90=1.28
n=(1:96+1:28)24s2
42¼42s2
42
Fora0.5standard deviation di®erence, ¢=s=0:5,so
n¼42
(0:5)2=168
Ifyouendupwithn<30,thenyoushouldbeusingthe
t-distribution ratherthanthenormaltocalculate critical
values,andthentheprocessisiterative.
422
SurvivalStudies:ComparingProportionsofEvents
Insomecases,thesamplesizeforasurvivaltrialisbased
onacrudecomparison oftheproportionofeventsatsome
¯xedpointintime.
Inthiscase,wecanapplytheresultsjustshowntogetsample
sizes,basedonthenormalapproximation tothebinomial:
De¯ne:
Pcprobabilit yofeventincontrolarmbytimet
Peprobabilit yofeventin\experimental"armbytimet
Thenumberofpatientsrequired pertreatmen tarmbased
onachi-square testcomparing binomial proportionsis:
N=fz1¡®
2q
2P(1¡P)+z1¡¯q
Pe(1¡Pe)+Pc(1¡Pc)g2
(Pc¡Pe)2
whereP=(Pe+Pc)=2
(Thislooksslightlydi®erentbecausethevarianceisnotthe
sameunderHoandHa,aswasthecaseinthenormalpre-
viousexample.)
423
Notesoncomparingproportionsoffailures:
²Useofchi-square testisbestwhen0:2<Pe;Pc<0:8
²Shouldhave¸15patientsineachcellofthe(2x2)table
²Forsmallersamplesizes,useFisher'sexacttesttomo-
tivatepowercalculations
²E±ciency vslogranktestisnear100%forstudieswith
shortdurations relativetothemedianeventtime
Whatdoesthismeanintermsoftheevent
rates?Highorlow?
²Calculation ofsamplesizeforcomparing proportionsof-
tenprovidesanupperboundtothosebasedoncompar-
isonofsurvivaldistributions
424
Samplesizebasedonthelogranktest
Recap: Consider atwogroupsurvivalproblem, withequal
numbersofindividuals inthetwogroups(sayn0ingroup0
andn1ingroup1).Let¿1;:::;¿KrepresenttheKordered,
distinctfailuretimes,andatthej-theventtime:
Die/Fail
Group Yes NoTotal
0d0jr0j¡d0jr0j
1d1jr1j¡d1jr1j
Totaldjrj¡djrj
whered0jandd1jarethenumberofdeaths(events)ingroup
0and1,respectively,atthej-theventtime,andr0jandr1j
arethecorrespondingnumbersatrisk.
Thelogranktestis:(z-statistic version)
ZLR=PK
j=1(d1j¡ej)
s
PKj=1vj
withej=djr1j=rj
vj=r1jr0jdj(rj¡dj)=[r2
j(rj¡1)]
425
Distributionofthelogrankstatistic
Supposethatthehazardratesinthetwogroupsare¸0(t)
and¸1(t),withhazardratio
µ=e¯=¸1(t)
¸0(t)
andsupposeweareinterestedintestingHo:¯=ln(µ)=0
(whichisequivalenttotestingHo:µ=1.)
[Note:wewilluseln(µ)ratherthan¯inthefollowing,sothatthere
isnoconfusionwiththeTypeIIerrorrate]
Itispossibletoshowthat
²iftherearenoties,and
²weare\near"H0:
then:
²E(d1j¡ejjd1j;d0j;r1j;r0j)¼ln(µ)=4
²vj¼1=4
So,atapointln(µ)inthealternativ e,weget:
ZLR¼PK
j=1ln(µ)=4
rPKj=11=4=dln(µ)=4
r
d=4=p
dln(µ)
2
andZLR»N(ln(µ)p
d=2;1)
426
HeuristicProof:
E(d1jjd1j;d0j;r1j;r0j)=Pr(d1j=1jdj=1;r1j;r0j)
=r1j¸0µ
r1j¸0µ+r0j¸0
=r1jµ
r1jµ+r0j
=r1j
r1j+r0j+ln(µ)2
64r1jr0j
(r1j+r0j)23
75
Butej=r1j=(r1j+r0j),so:
E(d1jjd1j;d0j;r1j;r0j)¡ej=ln(µ)2
64r1jr0j
(r1j+r0j)23
75
Ifn0=n1,thennearH0:,r1j¼r0j,hence,
E(d1jjd1j;d0j;r1j;r0j)¡ej=ln(µ)=4
Similarly ,withnoties,wehave
vj=r1jr0j=r2
j¼1=4
427
Thiscanalsobederivedviathepartiallikelihood:
Wecanwritethepartiallikelihoodas:
l(¯)=log2
664nY
j=10
B@e¯Zj
P
`2R(¿j)e¯Z`1
CA±j3
775
=nX
j=1±j2
64¯Zj¡log0
B@X
`2R(¿j)e¯Z`1
CA3
75
andthenthe\score"(partialderivativeoflog-likelihood)becomes:
U(¯)=@
@¯`(¯)
=nX
j=1±j2
64Zj¡P
`2R(¿j)Z`e¯Z`
P
`2R(¿j)e¯Z`3
75
Wecanwritethe\information" (minussecondpartialderivative
ofthelog-likelihood)as:
¡@2
@¯2`(¯)=nX
j=1±j2
64P
`2R(¿j)e¯Z`P
`2R(¿j)Z`e¯Z`¡(P
`2R(¿j)Z`e¯Z`)2
P
`2R(¿j)e¯Z`3
75
Thelogrankstatistic(withnoties)isequivalenttothescorestatistic
fortesting¯=0:
ZLRU(0)p
I(0)
ByaTaylorseriesexpansion:
U(0)»=U(¯)¡¯@U
@¯(0)
E[U(0)]»=¯d=4andI(0)»=d=4
428
PoweroftheLogrankTest
Usingasimilarargumen ttobefore,thepowerofthelogrank
test(basedonatwo-sided®leveltest)isapproximately:
Power(µ)¼1¡©·
z1¡®
2¡ln(µ)p
d=2¸
Note:Powerdependsonlyondandµ!
Wecaneasilysolvefortherequired numberofeventsto
achieveacertainpowerataspeci¯edvalueofµ:
Toyieldpower(µ)=1¡¯,wewantdsothat
1¡¯=1¡©µ
z1¡®
2¡ln(µ)p
d=2¶
)z¯=z1¡®
2¡ln(µ)p
d=2
)d=4µ
z1¡®
2¡z¯¶2
[ln(µ)]2
ord=4µ
z1¡®
2+z1¡¯¶2
[ln(µ)]2
429
Example:
Saywewereplanning a2-armstudy,andwantedtobeable
todetectahazardratioof1.5with90%powerata2-sided
signi¯cance levelof®=0:05.
Required numberofevents:
d=4µ
z1¡®
2+z1¡¯¶2
[ln(µ)]2
=4(1:96+1:282)2
[ln(1:5)]2
¼42
0:1644=256
#EventsrequiredforvariousHazardRatios
HazardPower
Ratio 80% 90%
1.5 191 256
2.0 66 88
2.5 38 50
3.0 26 35
Moststudiesaredesigned todetectahazardratioof1.5-2.0.
430
PracticalConsiderations
²Howdowedecideonµ?
²Howdowetranslate numbersoffailurestonumbersof
patients?
Hazardratiosfortheexponentialdistribution
Thehazardratiofromtwoexponentialdistributions canbe
easilytranslated intomoreintuitivelyinterpretable quanti-
ties:
Median:
IfTi»exp(¸i),then
Median(Ti)=¡ln(0:5)=¸i
Itfollowsthat
Median(T1)
Median(T0)=¸0
¸1=e¡¯=1
µ
Hence,doubling themediansurvivalofatreatedcompared
toacontrolgroupwillcorrespondtohalvingthehazard.
431
R-yearsurvivalrates
SupposetheR-yearsurvivalrateingroup1isS1(R)andin
group0isS0(R).Undertheexponentialmodel:
Si(R)=exp(¡¸iR)
Hence,
ln(S1(R))
ln(S0(R))=¡¸1R
¡¸0R=¸1
¸0=e¯=µ
Hence,doubling thehazardratefromgroup1togroup0will
correspondtodoubling thelogoftheR-yearsurvivalrate.
NotethatthisresultdoesnotdependonR!.
Example: Supposethe5-yearsurvivalrateontreatmen tA
is20%andwewant90%powertodetectanimprovementof
thatrateto30%.Thecorrespondinghazardratiooftreated
tocontrolis:
ln(0:3)
ln(0:2)=¡1:204
¡1:609=0:748
Fromourprevious formula,thenumberofevents(deaths)
neededtodetectthisimprovementwith90%power,based
ona2-sided5%leveltestis:
d=4(1:96+1:282)2
[ln(0:748)]2=499
432
TranslatingtoNumberofEnrolledPatients
First,supposethatwewillenterNpatientsintoourstudy
attime0,andwillthencontinuethestudyforFunitsof
time.
UnderH0,theprobabilit ythatanindividual willfailduring
thestudyis:
Pr(fail)=ZF
0¸0e¡¸0tdt
=1¡e¡¸0F
Hence,ifourcalculations sayweneeddfailures, thento
decidehowmanypatientstoenter,wesimplysolve
d=(N=2)(1¡e¡¸0F)+(N=2)(1¡e¡¸1F)
Tosolvetheaboveequation forN,weneedtosupplyvalues
ofFandd.Inotherwords,herewearealreadydeciding
whatHRwewanttodetect(withwhatpower,etc),andfor
howlongwearegoingtofollowpatients.Whatwegetis
thetotalnumberofpatientsweneedtoenrollinorderto
observethedesirednumberofeventsinFunitsoffollow-up
time.
433
Example: Supposewewanttodetecta50%improvement
inthemediansurvivalfrom12monthsto18monthswith
80%powerat®=0:05,andweplanonfollowingpatients
for3years(36months).
Wecanusethetwomedians tocalculate boththeparameters
¸0and¸1andthehazardratio,µ:
Median(Ti)=¡ln(0:5)=¸i
so¸1=¡ln(0:5)
M1=0:6931
18=0:0385
¸0=¡ln(0:5)
M0=0:6931
12=0:0578
µ=¸1
¸0=0:0385
0:0578=12
18=0:667
andfromourprevious table,#eventsrequired isd=191
(sameforµ=1:5asitisfor1/1.5=0.667).
Soweneedtosolve:
191=(N=2)(1¡e¡0:0578¤36)+(N=2)(1¡e¡0:0385¤36)
=(N=2)(0:875)+(N=2)(0:7500)=(N=2)(1:625)
)N=235
(forpractical reasons, wewouldprobably roundupto236
andrandomize 118patientstoeachtreatmen tarm)
434
Amorerealisticaccrualpattern
Inreality,noteveryonewillenterthestudyonthesameday.
Instead, theaccrualwilloccurina\staggered" mannerover
aperiodoftime.
Thestandardassumption:
Supposeindividuals enterthestudyuniformly overanac-
crualperiodlastingAunitsoftime,andthataftertheac-
crualperiod,follow-upwillcontinueforanotherFunitsof
time.
TotranslatedtoN,weneedtocalculate theprobabilit ythat
apatientfailsunderthisaccrualandfollow-upscenario.
Pr(fail)=ZA
0Pr(failjenterata)f(a)da
=1¡RA
0S(a+F)da
A(2)
Thensolve:d=(N=2)Pr(fail;¸0)+(N=2)Pr(fail;¸1)
=(N=2)Pc+(N=2)Pe
=(N=2)(Pc+Pe)
IfwenowsolveforN(substituting informulaford),weget:
N=2d
(Pc+Pe)
N=8µ
z1¡®
2+z1¡¯¶2
[ln(µ)]21
(Pc+Pe)
435
HowcanwegetPcandPefrom(2)?
Ifweassumethattheexponentialdistribution holds,then
wecansolve(2)toobtain:
Pi=1¡exp(¡¸iF)(1¡exp(¡¸iA))
¸iA(3)
(fori=c;e)
Freedman suggested anapproximation forPcandPe,by
computing theprobabilit yofaneventatthemedianduration
offollow-up,(A=2+F):
Pi=Pr(fail;¸i)=1¡exp[¡¸i(A=2+F)](4)
Heshowedthatthisapproximation worksprettywellforthe
exponentialdistribution (i.e.,itgivesvaluescloseto(3)).
436
Analternativeformulation
Rubenstein,Gail,andSantner(1981)suggestthefollowing
approachforcalculating thetotalsamplesizethatmustbe
enrolled:
N=2µ
z1¡®
2+z1¡¯¶2
[ln(µ)]22
41
Pc+1
Pe3
5
wherePcandPearetheexpectedproportionofpatientsor
individuals whowillfail(haveanevent)onthecontroland
treatmen tarms.
Howdowecalculate (estimate)PcandPe?
²usingthegeneralformulaforadistribution Sgivenin
(2)
²usingtheexactformulaforanexponentialdistribution
givenin(3)
²usingtheapproximation givenby(4)
Note:alloftheseformulascanbemodi¯edforunequal
assignmen ttotreatmen t(orexposure)groupsbychanging
(N=2)intheformulasonp.17-19to(qc¤N)and(qe¤N),
whereqcandqearetheproportionsassigned tothecontrol
andexposedgroups,respectively.
437
Freedman'sApproach(1982)
Freedman's approachisbasedonthelogrankstatisticunder
theassumption ofproportionalhazards, butdoesnotrequire
theassumption ofexponentialsurvivaldistributions.
Totalnumberofevents:
d=µ
z1¡®
2+z1¡¯¶20
@µ+1
µ¡11
A2
Totalsamplesize:
N=2µ
z1¡®
2+z1¡¯¶2
Pe+Pc0
@µ+1
µ¡11
A2
wherePeandPcareestimated using(4).
Thisapproximation dependsontheassumption ofacon-
stantratiobetweenthenumberofpatientsatriskinthetwo
treatmen tgroupspriortoeacheventtime=)r0j¼r1j(as
showninthe\heuristic proof").Whenthisassumption is
notsatis¯ed, therequired samplesizestendtobeoveresti-
mated.
Q.Whenwouldthisassumptionnotbesatis¯ed?
A.Whenthesmallestdetectabledi®erenceislarge.
438
Someexamplesofstudydesign
ExampleI:
Aclinicaltrialinesophageal cancerwillrandomize patients
toradiotherap yalone(RxA)versusradiotherap ypluschemother-
apy(RxB).Thegoalofthestudyistocompare thetwo
treatmen tswithrespecttosurvival,andweplantousethe
logranktest.Fromhistorical data,weknowthatthemedian
survivalonRXAforthisdiseaseisaround9months.We
want90%powertodetectanimprovementinthismedian
to18months.Paststudieshavebeenabletoaccrueap-
proximately 50patientsperyear.Chooseasuitable study
design.
439
ExampleII:
Aclinicaltrialinearlystagebreastcancerwillrandomize
patientsaftertheirsurgerytoTamoxifen(ARMA)versus
observationonly(ARMB).Thegoalofthestudyistocom-
parethetwotreatmen tswithrespecttotimetorelapse,and
thelogranktestwillbeusedintheanalysis. Fromhistorical
data,weknowthatafter¯veyears,65%ofthepatientswill
stillbediseasefree.Wewouldliketohave90%powerto
detectanimprovementinthisdiseasefreerateto75%.Past
studieshavebeenabletoaccrueapproximately 200patients
peryear.Chooseasuitablestudydesign.
440
ExampleIII:
Someinvestigators intheenvironmen talhealthdepartmen t
wanttoconduct astudytoassessthee®ectsofexposureto
tolueneontimetopregnancy .Theywillconduct acohort
studyinvolvingwomenwhoworkinachemical factoryin
China. Itisestimated that20%ofthewomenwillhave
workplace exposuretotoluene. Furthermore, itisknown
thatamongunexposedwomen,80%willbecomepregnant
withinayear.Theinvestigators willbeabletoenroll200
womenperyearintothestudy,andplananadditional year
offollow-upattheendofaccrual. Assuming theyhave2
yearsaccrual, whatreduction inthe1-yearpregnancy rate
forexposedwomenwilltheybeabletodetectwith85%
power?Whatiftheyhave3yearsofaccrual?
441
Otherimportantissues:
Theapproachesjustdescribedaddressthebasicquestion of
calculating asamplesizeforstudywithasurvivalendpoint.
Theseapproachesoftenneedtobemodi¯edslightlytoad-
dressthefollowingcomplications:
²Losstofollow-up
²Non-compliance (orcross-overs)
²Strati¯cation
²Sequentialmonitoring
²Equivalencehypotheses
Nextwesummarize someofthemainpoints.
442
Losstofollow-up
Ifsomepatientsarelosttofollowup(asopposedtocensored
attheendofthetrialwithouttheevent),thepowerwillbe
decreased.
Therearetwomainapproachesfordealingwiththis:
²Simplein°ationmethod-If`*100%ofpatients
areanticipated tobelosttofollowup,calculate target
samplesizetobe
N¤=0
@1
1¡`1
A¢N
Example: SayyoucalculateN=200,andanticipate
lossesof20%.Thesimplein°ation methodwouldgive
youatargetsamplesizeofN¤=(1=0:8)¤200=250.
Warning: peopleoftenmakethemistakeofjustin-
°atingtheoriginalsamplesizeby`*100%,whichwould
havegivenN¤=240fortheexample above.
²Exponentiallossassumption -theaboveapproach
assumes thatlossescontributeNOinformation. Butwe
actuallyhaveinformation onthemupuntilthetimethat
theyarelost.Incorporatethisbyassuming thattimeto
lossalsofollowsanexponentialdistribution, andmodify
PeandPc.
443
Noncompliance
Ifsomepatientsdon'ttaketheirassigned treatmen ts,the
powerwillbedecreased. Thisissuehastwosides:
²Drop-outs(de)-patientswhocannottolerate the
medication stoptakingit;theirhazardratewouldbe-
comethesameastheplacebogroup(ifincluded instudy)
atthatpoint.
²Drop-ins(dc)-patientsassigned tolesse®ectivether-
apymaynotgetrelieffromsymptoms andseekother
therapy,orrequesttocross-over.
Aconservativeremedy{adjustPeandPcasfollows:
P¤
e=Pe(1¡de)+Pcde
P¤
c=Pc(1¡dc)+Pedc
444
DesignStrategy:
1.Decideon
²TypeIerror(signi¯cance level)
²clinically importantdi®erence (intermsofHR)
²desiredpower
2.Determine thenumberoffailuresneeded
3.Basedonpastexperience
²decideonareasonable distribution forthecontrols
(usually exponential)
²estimate anticipated accrualperunittime
²estimate expectedrateoflosstofollowup
VarythevaluesofAandFuntilyougetsomething prac-
ticallyfeasiblethatgivestherightnumberoffailures.
4.Consider noncompliance, sequentialmonitoring, andother
issuesimpacting samplesize
445
Included onthenextseveralpagesisaSASprogram tocal-
culatesamplesizesforsurvivalstudies.Itusesseveralofthe
approacheswe'vediscussed, including:
²Rubenstein,GailandSantner(RGS,1981)
²Freedman (1982)
²LachinandFoulkes(1986)
Acopyofthisprogram isshownonthenextseveralpages.
Theprogram requiresentryof:
²Signi¯cance level(alpha)
²Power
²Sides(1forone-sided test,2fortwo-sidedtest)
²Accrualperiod
²Followupperiod
²Yearlyrateoflosstofollow-up
²Proportionrandomized toexperimentaltreatmen tarm
²Oneofthefollowing:
{Yearlyeventrateoncontrolandexperimentaltreat-
mentarms
{Yearlyeventrateoncontrolarm,andthehazard
ratio
{Mediantimetoeventoncontrolandexperimental
treatmen tarms
446
TheSASprogram rgsnew.sas
datargs;
******************************************************************* ****;
***enterthefollowing information inthisblock;
alpha=0.05; /*significance level*/
sides=2; /*one-sided ortwo-sided test*/
power=0.90; /*Desired power*/
accrual =2; /*Accrual periodinyears*/
fu=1.5; /*Followupafterlastpatient isaccrued */
loss=0.0; /*yearlyrateofloss*/
qe=0.5; /*proportion randomized toexperimental arm*/
***eitherenterthemediantimetoeventinyearsoncontrol;
***ortheyearlyeventrate-leavetheothervaluemissing;
medianc =0.75; /*mediantimetoeventoncontrol arm*/
probc=.; /*yearlyeventrateincontrol arm*/
***eitherentertheyearlyeventrateintheexperimental arm;
***orthehazardratioforcontrol vsexperimental ;
***orthemediantimetoeventonexperimental arminyears ;
***leavetheothervaluesmissing (.);
mediane =1.5; /*mediantimetoeventonexperimental */
probe=.; /*yearlyeventrateinexperimental arm*/
rr=.; /*hazardratio*/
******************************************************************* ****;
beta=1-power;
qc=1-qe;
zalpha=probit(1-alpha/sides);
zbeta=probit(1-beta);
***calculate yearlyeventrateinbotharmsusingmedians, ifsupplied;
ifmedianc^=. thendo;
hazc=-log(0.5)/medianc;
probc=1-exp(-hazc);
end;
ifmediane^=. thendo;
haze=-log(0.5)/mediane;
probe=1-exp(-haze);
end;
hazc=-log(1-probc);
447
***calculate hazardinexperimental group,usingyearlyeventrate;
***orhazardratio;
ifprobe^=. thenhaze=-log(1-probe);
ifrr^=.thenhaze=hazc/rr;
ifprobe^=. andhaze^=. thendo;
put"**************************************************************";
put"WARNING: bothyearlyeventrateandhazardratio(HR)have";
put" beenspecified. Calculations willusetheHR";
put"**************************************************************" /;
end;
***calculate mediansurvival timesifnotsupplied;
medianc=-log(0.5)/hazc;
mediane=-log(0.5)/haze;
hazl=-log(1-loss);
avghaz=qc*hazc +qe*haze;
rr=hazc/haze;
log_rr=log(hazc/haze);
totloss=(accrual*0.5 +fu)*loss;
***compute expected probability ofdeath(event) duringtrial;
***givenstaggered accrual butNOloss;
pc0loss =1-((exp(-hazc*fu)-exp(-hazc*(accrual+fu)))/(hazc*accrual) );
pe0loss =1-((exp(-haze*fu)-exp(-haze*(accrual+fu)))/(haze*accrual) );
***compute expected probability ofeventduringtrial;
***givenstaggered accrual ANDloss;
pc=(1-(exp(-(hazc+hazl)*fu)-exp(-(hazc+hazl)*(fu+accrual)))
/((hazc+hazl)*accrual))*(hazc/(hazc+hazl));
pe=(1-(exp(-(haze+hazl)*fu)-exp(-(haze+hazl)*(fu+accrual)))
/((haze+hazl)*accrual))*(haze/(haze+hazl));
pbar=(1-(exp(-(avghaz+hazl)*fu)-exp(-(avghaz+hazl)*(fu+accrual)) )
/((avghaz+hazl)*accrual))*(avghaz/(avghaz+hazl));
***compute totalsamplesizeassuming loss;
N=int(((zalpha+zbeta)**2)/(log_rr**2)*(1/(qc*pc)+1/(qe*pe) ))+1;
***Compute samplesizeusingmethodofFreedman (1982);
N_FRD=int((2*(((rr+1)/(rr-1))**2)*(zalpha+zbeta)**2)/(2*(q e*pe+qc*pc)))+1;
448
***Compute samplesizeusingmethodofLachinandFoulkes (1986);
***withratesunderH0givenbypooledhazard;
N_LF=int((zalpha*sqrt((avghaz**2)*(1/pbar)*(1/qc +1/qe))+
zbeta*sqrt((hazc**2)*(1/(qc*pc)) +(haze**2)*(1/(qe*pe))))**2/
((hazc-haze)**2)) +1;
***compute totalsamplesizeassuming noloss;
n_0loss =int(((zalpha+zbeta)**2)/(log_rr**2)*
(1/(qc*pc0loss)+1/(qe*pe0loss))) +1;
***compute samplesizeusingsimpleinflation methodforloss;
naive=int(n_0loss/(1-totloss)) +1;
***verifythatactualpowerissameasdesired power;
newpower =probnorm(sqrt(((N)*(log_rr**2))/(1/(qc*pc) +1/(qe*pe)))
-zalpha);
ifabs(newpower-power)>0.001 thendo;
put'***WARNING: actualpowerisnotequaltodesired power';
put'Desired power:'power'Actualpower:'newpower;
end;
******************************************************************* ****;
***compute numberofeventsexpected duringtrial;
******************************************************************* ****;
***Compute expected numberundernull;
n_evth0 =int(n*pbar) +1;
r=qc/qe;
***Rubinstein, GailandSantner (1981)method-simpleapproximation;
n_evtrgs =int((((r+1)**2)/r)*((zalpha +zbeta)**2)/(log_rr**2)) +1;
***Freedman (1982);
n_evtfrd =int((((rr+1)/(rr-1))**2) *(zalpha+zbeta)**2)+1;
***Usingbacktracking methodofLachinandFoulkes (1986);
n_evt_c =int(N*qc*pc) +1;
n_evt_e =int(N*qe*pe) +1;
n_evtlf =n_evt_c +n_evt_e;
449
labelsides='Sides'
alpha='Alpha'
power='Power'
beta='Beta'
zalpha='Z(alpha)'
zbeta='Z(beta)'
accrual='Accrual (yrs)'
fu='Follow-up (yrs)'
loss='Yearly Loss'
totloss='Total Loss'
probc='Yearly eventrate:control'
probe='Yearly eventrate:active'
rr='Hazard ratio'
log_rr='Log(HR)'
N='Total Samplesize(RGS)'
N_FRD='Total Samplesize(Freedman)'
N_LF='Total Samplesize(L&F)'
n_0loss='Sample size(noloss)'
pc='Pr(event), control'
pe='Pr(event), active'
pc0loss='Pr(event| noloss),control'
pe0loss='Pr(event| noloss),active'
n_evth0='# events(Ho-pooled)'
n_evtrgs='# events(RGS)'
n_evtfrd='# events(Freedman)'
n_evtlf='# events(L&F)'
medianc='Median survival, control'
mediane='Median survival, active'
naive='Sample size(naiveloss)';
procprintdata=rgs labelnoobs;
title'Sample size&expected eventsforcomparing twosurvival distributions';
title2'UsingmethodofRubinstein, GailandSanter(RGS,1981)';
title3'Freedman (1982), orLachinandFoulkes (L&F,1986)';
varsidesalphapoweraccrual fulosstotloss
probcprobemedianc mediane pcpepc0loss pe0loss rrlog_rr
n_evth0 n_evtrgs n_evtfrd n_evtlf NN_FRDN_LFn_0loss naive;
formatpowerlosstotloss f4.2medianc mediane f5.3
probcproberrlog_rrpcpepc0loss pe0loss f6.4;
450
BacktoExampleI:
Aclinicaltrialinesophageal cancerwillrandomize patients
toradiotherap yalone(RxA)versusradiotherap ypluschemother-
apy(RxB).Thegoalofthestudyistocompare thetwo
treatmen tswithrespecttosurvival,andweplantousethe
logranktest.Fromhistorical data,weknowthatthemedian
survivalonRxAforthisdiseaseisaround9months.We
want90%powertodetectanimprovementinthismedian
to18months.Paststudieshavebeenabletoaccrueap-
proximately 50patientsperyear.Chooseasuitable study
design.
First,let'swritedownwhatweknow:
²desiredsigni¯cance levelnotstated,souse®=0:05
(assume atwo-sidedtest)
²assumeequalrandomization totreatmen tarms
(unlessotherwise stated)
²desiredpoweris90%
²mediansurvivaloncontrolis9months)M0=9
²wanttodetectimprovementto18monthsonRxB)
M1=18
²Maximumaccrualperyearis50patients
451
Wehavealloftheinformation weneedtoruntheprogram,
excepttheaccrualandfollowuptimes.Weneedtousetrial
anderrortogetthese.
NumberofTotalTotal
Accrual Follow-up EventsSample Study
PeriodPeriodRequired SizeDuration
1 2.5 88 106 3.5
2 1.5 88 115 3.5
2.5 1 88 122 3.5
3 0.5 88 133 3.5
3 1 88 117 4
Shownonthenextpageistheoutputfromrgsnew.sas us-
ingAccrual=2, Follow-up=1.5. I'vegiventheRGSnumbers
above.
Whichoftheabovearefeasibledesigns?
452
Samplesize&expected eventsforcomparing twosurvival distributions
UsingmethodofRubinstein, GailandSanter(RGS,1981)
Freedman (1982), orLachinandFoulkes (L&F,1986)
Yearly
event
Accrual Follow-up Yearly Total rate:
Sides Alpha Power (yrs) (yrs) Loss Loss control
20.050.90 2 1.5 0.00 0.00 0.6031
Yearly
event Median Median Pr(event|
rate: survival, survival, Pr(event), Pr(event), noloss),
active control active control active control
0.3700 0.750 1.500 0.8860 0.6737 0.8860
Pr(event|
noloss), Hazard #events #events #events
active ratio Log(HR) (Ho-pooled) (RGS) (Freedman)
0.6737 2.0000 0.6931 94 88 95
Total Sample
Total Sample Total Sample size
#events Sample size Sample size(no(naive
(L&F) size(RGS) (Freedman) size(L&F) loss) loss)
90 115 122 121 115 115
453
Howdowepickfromthefeasibledesigns?
The¯rst4designsallhave31/2yearstotalduration, since
thefollow-upperiodstartsafterthelastpatienthasbeen
accrued. Theshorterthefollow-upperiodgiventhis¯xed
studyduration, themorepatientswehavetoenroll.
Insomecases,itwillbemuchmorecost-e®ectiv etoenroll
fewerpatientsandfollowthemforlonger.Thiscorresponds
tocaseswheretheinitialcostperpatientisveryhigh.
Inothercases(wheretheinitialcostperpatientislower),it
willbebettertoenrollmorepatients.Themedianfollow-up
forthe¯rst4designsare3,2.5,2.25,and2years,respec-
tively.Thetotalcostoftreatmen tcouldbeestimated by
multiplying thenumberofpatientsbythemedianfollow-up
time.
Someprefertokeeptheaccrualperiodasshortaspossi-
ble,givenhowmanypatientscanfeasiblybeenrolled. This
willtendtogivethesmallest numberofpatientsamongthe
feasibledesigns. Whichdesignwouldthiscorrespondto?
Another issuetothinkaboutiswhetherthebackground con-
ditionsofthediseasearechangingrapidly(likeAIDS)orare
fairlystable(likemanytypesofcancer). Fortheformersitu-
ation,itwouldbebesttohaveastudywithashortduration
sotheresultswillhavemoreinterpretation.
454
Usingtheinformation given,therearealotofotherquanti-
tieswecancalculate:
²Thehazardratioofcontroltotreated is:
median(Rx B)
median(Rx A)=18
9=2
²Thehazardratesforthetwotreatmen tarmsare:
forRxA:¸0=¡log(0:5)
median(Rx A)=¡log(0:5)
9=0:0770
forRxB:¸1=¡log(0:5)
median(Rx B)=¡log(0:5)
18=0:0385
²Theyearlyprobabilityofaneventis:
forRxA:Pr(T<1j¸0)=1¡e(¡¸0¤t)
=1¡e(¡0:0770¤12)=0:603
forRxB:Pr(T<1j¸1)=1¡e(¡¸1¤t)
=1¡e(¡0:0385¤12)=0:370
Whatwouldhappenaboveifweusedtimet
inyears(i.e.,t=1)insteadofmonths?
Whatwouldhappenifwecalculatedboththe
hazardrateandyearlyeventprobabilityusing
timeinyears?
455
Basedonadesignwith2.5yearsaccrualand
1yearfollow-up:
²Themedianfollow-uptime
medianFU=A=2+F
=30=2+12=27months
²Theprobabilit yofaneventduringtheentirestudyis:
(usingtheapproximation innotes)
forRxA:Pc=1¡exp(¡¸0¤[A=2+F])
=1¡exp(¡0:0770¤27)=0:875
forRxB:Pe=1¡exp(¡¸0¤[A=2+F])
=1¡exp(¡0:0385¤27)=0:646
(theabovenumbersdi®erfromwhatyou'dgetinthe
printoutfromtheprogram, sinceitcalculates theexact
probabilit yundertheexponentialdistribution, instead
ofusingtheapproximation)
Inthecalculations above,allofthe\time"periodswerein
termsofmonths.Youhavetoremembertokeepthescale
thesamethroughout.
Tousetheprogram, youneedtotranslate thetimescalein
termsofyears.Soamedianof18monthssurvivalwouldbe
enteredasmedian=1.5.
456
Whathappensifweaddlosstofollow-up?
RequiredsamplesizeforA=2.5,FU=1year
YearlyNumberofTotalTotal
LosstoEventsSample Study
Follow-upRequired SizeDuration
0 88 122 3.5
5% 88 128 3.5
10% 88 133 3.5
20% 88 147 3.5
457
SequentialDesignandAnalysisofsurvivalstud-
ies
Inclinicaltrialsandotherstudies,itisoftendesirable to
conduct interimanalyses ofastudywhileitisstillongoing.
Rationale:
²ethical: ifonetreatmen tissubstantiallyworsethan
another, thenitiswrongtocontinuetogivetheinferior
treatmen ttopatients.
²timelyreporting: ifthehypothesisofinteresthas
beenclearlyestablished halfwaythroughthestudy,then
scienceandthepublicmaybene¯tfromearlyreporting.
WARNING!!
Unplanned interimanalyses canseriously in°atethetrue
typeIerrorofatrial.Ifinterimanalyses aretobeperformed,
itisESSENTIAL tocarefully plantheseinadvance,andto
adjustalltestsappropriately sothethetypeIerrorisofthe
desiredsize.
458
HowdoesthetypeIerrorbecomein°ated?
Consider atwogroupstudycomparing treatmen tsAandB.
Supposethedataarenormally distributed (sayXi»N(¹A;¾2)
ingroupA,andsimilarly forgroupB),sothatthenullhy-
pothesisofinterestis
H0:¹A=¹B
Itisnottoohardto¯gureouthowthetypeIerrorcanget
in°atedifanaiveapproachisused.
SupposeweplantodoKinterimanalyses, andthatexactly
mindividuals willentereachtreatmen tbetweeneachanal-
ysis.Theteststatistic atthekthanalysiswillbe
Zk=Pk
i=1Pm
j=1(XAij¡XBij)=km
r
2¾=km=Pk
i=1di=k
r
2¾=km
wherediisthedi®erence betweenthetwogroupmeansat
theithanalysis,
di=XAi¡XBi
andXAiandXBiarethemeansingroupsAandBofthe
mindividuals whoenteredinthei-thtimeperiod.
459
(naive)Interimmonitoringprocedure:
²Allowmpatientstoenteroneachtreatmen tarm
(totalof2madditional patients)
²CalculateZkbasedonthecurrentdata
²RejectthenullhypothesisifjZkj>z1¡®=2,where®is
thedesiredtypeIerror.
TheoveralltypeIerrorrateforthestudyis:
Pr(jZ1j>z1¡®=2orjZ2j>z1¡®=2...orjZKj>z1¡®=2)
Ifthetestateachinterimanalysis isperformed atlevel®,
thenclearlythisprobabilit ywillexceed®.Thetablebelow
showstheTypeIerrorrateifeachtestisdoneat®=0:05
forvariousvaluesofK:
Numberofinterimanalyses(K)
123451025
5%8.3%10.7%12.6%14.2%19.3%26.6%
(fromLee,StatisticalMethodsforSurvivalData,Table
12.9)
460
Forsurvivaldata,thecalculations becomeMUCHmorecom-
plicated sincethedatacollected withineachtimeinterval
continuestochangeastimegoeson!
WhatcanwedotoprotectagainstthistypeIerrorin°ation?
PocockApproach:
Pickasmallersigni¯cance level(say®0)touseateachinterim
analysissothattheoveralltypeIerrorstaysatlevel®.
Aproblem withthePocockmethodisthateventhevery
lastanalysis hastobeperformed atlevel®0.Thistendsto
beveryconservativeatthe¯nalanalysis.
O'BrienandFlemingApproach:
Apreferable approachwouldbetovarythealphalevelsused
foreachoftheKinterimanalyses, andtrytokeepthevery
lastone\close"tothedesiredoverallsigni¯cance level.The
O'Brien-Fleming approachdoesthat.
461
Commentsandnotes:
²Thereareseveralotherapproachesavailableforsequen-
tialdesignandanalysis. TheO'BrienandFleming
approachisprobably themostpopularinpractice.
²Therearemanyvariations onthethemeofsequential
design.Thetypewehavediscussed hereiscalledGroup
sequentialanalysis .
{Thereareotherapproachesthatrequirecontinuous
analysisaftereachnewindividual entersthestudy!
{Therearealsoapproacheswheretherandomization
itselfismodi¯edasthetrialproceeds.E.g.Ze-
len's\Playthewinnerrule"(NewEngland Journalof
Medicine 300,1979,page1242)andWare's\ECMO"
study(Statistical Science, 4,1989,page298)
²Somedesignsallowforearlystopping intheabsenceofa
su±cienttreatmen te®ectasthetrialprogresses. These
proceduresarereferredtoas\stochasticcurtailmen t"or
\conditional power"calculations.
462
²Designing agroupsequentialtrialforsurvivaldatare-
quiressophisticated andhighlyspecialized software.EaSt,
apackagefromCYTEL SOFTWAREthatdoesstan-
dard(¯xed)survivaldesigns, aswellassequentialde-
signs.
²Many\non-statistical" issuesenterdecisions aboutwhether
ornottostopatrialearly
²P-valuesbasedonanalyses ofstudieswithsequential
designsaredi±culttointerpret.
²Onceyoudo5interimanalyses, thenaddingmoremakes
littledi®erence. Someclinicaltrialsgroups(HSPHAIDS
group)havelargerandomized PhaseIIIstudiesmon-
itoredatleastonceperyear(forsafetyreasons), and
moststudieshave1-3interimlooks.
²Goingfroma¯xedtoagroupsequentialdesignadds
onlyabout3-4%totherequired maximumsamplesize.
Thisisagoodruleofthumbtouseincalculating the
samplesizewhenyouplanondoinginterimmonitoring.
463
CompetingRisksandMultipleFailureTimes
Sofar,we'vebeenactingasiftherewasonlyoneendpoint
ofinterest,andthatcensoring duetodeath(orsomeother
event)wasindependentoftheeventofinterest.
However,inmanycontextsitislikelythatthetimetocen-
soringissomehowcorrelated withthetimetotheeventof
interest.Ingeneral, weoftenhaveseveraldi®erenttypes
offailure(death,relapse,opportunistic infection, etc)which
arerelated(i.e.,dependentor\competing"risks).
Examples:
²Afterabonemarrowtransplan tation,patientsarefol-
lowedtoevaluate\leukemia-fr eesurvival",sotheend-
pointistimetoleukemiarelapseordeath.Thisendpoint
conistsoftwotypesoffailures(competingrisks):
{leukemiarelapse
{non-relapse deaths
²Incardiovascularstudies,deathsfromothercauses(such
ascancer)areconsidered competingrisks.
²Inactuarial analyses, welookattimetodeath,butwant
toprovideseparate estimates ofhazardsforeachcause
ofdeath(multipledecremen tlifetables).
464
Anotherexample: FortheMACstudy,theanalyses you
havebeendoingoftimetoMACassumethatthecensoring
timeisindependent.
Recall:
T=timetoeventofinterest(MAC)
U=timetocensoring (death,losstoFU)
X=min(T;U)
±=I(T·U)
ObservableData:(X;±)
Whatarethepossiblitieshere?
²(1)FailureTandcensoringUareindependent
²(2)FailureTandcensoringUaredependent
465
Case(1):Independentfailuretimes
(thisincludes thecaseofindependentcensoring)
BOTTOMLINE)NOPROBLEM
Nonparametric estimation:
Inthiscase,wecanusetheKaplan-Meier estimator toesti-
mateST(t)=P(T>t).
Parametricestimation:
Ifweknowthejointdistribution of(T;U)hasacertain
parametric form(exponential,Weibull,log-logistic), thenwe
canusethelikelihoodfor(X;±)togetparameter estimates
ofthemarginal distribution ofST(t).
Semi-parametric estimation:
WecanapplytheCoxregression modeltoassessthee®ects
ofcovariatesonthemarginal hazard.
466
Case(2):Dependentfailuretimes
BOTTOMLINE)BIGPROBLEM
Tsiatis(1975)showedthatST(t)=P(T¸t)(i.e.,thesur-
vivalfunction fortheeventTofinterest)cannotbe\identi-
¯ed"fromdataoftheform(X;±)foreachsubject.
Infact,observing (X;±)doesnotprovideenoughinforma-
tiontoestimate thejointdistribution of(T;U)sothatwe
canevencheckwhether theassumption ofindependence is
valid.
Whenisitreasonabletoassumeindependentrisks?
²whencensoring occursbecausethestudyends,orbe-
causethesubjectmovestoadi®erentstate
²andthereisnotrendovertimeinhealthstatusofen-
rollingpatients
InthecaseofourMACstudy,thefactthatsomeone dies
mayre°ectthattheywouldhavebeenatgreaterriskofMAC
iftheyhadnotdiedthansomeone elsewhoremained alive
atthatpoint.
Theassumption ofindependence meansthatthehazardfor
someone whoiscensored attimetisexactlythesameasthat
forsomeone withthesamecovariateswhoisalsoatriskat
timet.
467
Whatistheimpactofdependentcompetingrisks?
SludandByar(1988)showthatdependentcausesofdeath
canpotentiallymakeriskfactorsappearprotectiv e:
Ifwehave
T=deathfromcauseofinterest
andU=censoring, fromdeathduetoothercause
andasinglebinarycovariateZ
Z=8
><
>:1ifriskfactorispresent
0otherwise
andwecalculate theKaplan-Meier survivalestimates ^S1(t)
forZ=1and^S0(t)forZ=0assuming independentcen-
soring,thenwecould(intheirhypothetical example) endup
reversingthesignofthesurvivalfunctions:
Trueorderingbetweensurvivaldistributions:
S1(t)<S0(t)forallt
KaplanMeierestimatesofsurvivaldistributions:
^S1(t)>^S0(t)forallt
468
Whatcanwedoifwesuspectdependentrisks?
Alotofpeoplehavetriedtotacklethisproblem!
References
AlyEAA,KocharSC,andMcKeagueIW(1994).Sometestsfor
comparingcumulativeincidencefunctionsandcause-speci¯chazard
rates. JASA89,994-999.
BenichouJandGailMH(1990).Estimatesofabsolutecausespeci¯c
riskincohortstudies. Biometrics 46,813-826.
(*)BoothALandSatchellSE(1995).ThehazardsofdoingaPhD:an
analysisofcompletionandwithdrawalratesofBritishPhDstudents
inthe1980's. JRSS-A,297-318.
FarewellVT(1979).AnapplicationofCox'sproportionalhazard
modeltomultipleinfectiondata.Applie dStatistics28,136-143.
Gail,M(1982).Competingrisks. Encyclop ediaofStatistic alScienc es
2,75-81.
LinDY,RobinsJM,andWeiLJ(1996).Comparingtwofailuretime
distributionsinthepresenceofdependentcensoring. Biometrika
83,381-393.
LunnMandMcNeilD(1995).ApplyingCoxregressiontocompeting
risks. Biometrics 51,524-532
MoeschbergerMLandKleinJP(1988).Boundsonnetsurvivalprob-
abilitiesfordependentcompetingrisks. Biometrics 44,529-538.
(*)MoeschbergerMLandKleinJP(1995).Statisticalmethodsforde-
pendentcompetingrisks. Lifetime DataAnalysis1,193-204.
469
PepeMS(1991).Inferenceforeventswithdependentrisksinmultiple
endpointstudies. JASA86,770-778.
(*)PepeMSandMoriM(1993).Kaplan-Meier, marginalorcondi-
tionalprobabilitycurvesinsummarizingcompetingrisksfailure
timedata? Statistics inMedicine12,737-751.
PrenticeRL,Kalb°eischJD,PetersonAV,FlournoyN,FarewellVT,
andBreslowNE(1978).Theanalysisoffailuretimesinthepresence
ofcompetingrisks. Biometrics 34,541-554.
SludEV,ByarDP,andSchatzkinA(1988).Dependentcompeting
risksandthelatentfailuremodel.Biometrics 44,1203-1205.
SludEandByarD(1988).Howdependentcausesofdeathcanmake
riskfactorsappearprotective.Biometrics 44,265-269.
Tsiatis,A.(1975).Anonidenti¯abilityaspectoftheproblemofcom-
petingrisks. ProceedingsoftheNational Academy ofScienc es72,
20-22.
470
Therehasbeenalivelydebateintheliterature
aboutthebestwaytoattackthisproblem.The
twosidesarebasicallydividedaboutwhichtype
ofmodeltouse:
²basedoncause-speci¯chazard functions (observ-
ables)
²basedonlatentvariable models(unobserv ables)
The¯rstapproachfocusesonwhattheobservedsurvivalis
duetoacertaincauseoffailure,acknowledingthatthereare
othertypesoffailuresoperatingatthesametime.
Thesecondapproachattempts toestimate whatthesurvival
associatedwithacertainfailuretypewouldhavebeen,ifthe
othertypesoffailureshadbeenremoved.
471
GeneralCaseofMultipleFailureTypes
Ingeneral, saywehavemdi®erenttypesoffailure(say,
causesofdeath),andtherespectivetimestofailureare:
T1;T2;T3;¢¢¢;Tm
andweobserveT=min(T1;T2;:::;Tm)
Wecanwritethecause-speci¯chazardfunction forthej-th
failuretypeas:
¸j(t)=lim
¢t!01
¢tPr(t·T<t+¢t;J=jjT¸t)
Theoverallhazardofdeathisthesumoverthefailuretypes:
¸(t)=mX
j=1¸j(t)
where¸(t)=lim
¢t!01
¢tPr(t·T<t+¢tjT¸t)
Q.Canweestimatethesequantities?...evenif
therisksaredependent?
A.Yes,Prentice(1978)showsthatprobabilities
thatcanbeexpressedasafunctionofthecause-
speci¯chazardscanbeestimated.
472
Forexample, estimable quantitiesinclude:
(a)Theoverallsurvivalprobability3:
ST(t)=P(T¸t)=exp"
¡Zt
0¸(u)du#
=exp2
64¡Zt
0X
j¸j(u)du3
75
(b)Conditionalprobabilityoffailingfromcause
jintheinterval(¿i¡1;¿i]
Q(i;j)=[ST(¿i¡1)]¡1Z¿
¿¡1¸j(u)ST(u)du
(c)Conditionalprobabilityofsurviving ithinter-
val
½i=1¡mX
j=1Q(i;j)
3Note:previously Isaidyoucouldn'testimateST(t),butthatwaswhenTwas
thetimetoeventofinterest(possiblyunobservable)andUwasthepossiblycor-
relatedcensoring time.Here,ST(t)isthesurvivaldistribution fortheminimum
ofallfailures,whichcanalwaysbeobserved
473
Estimators:
(a)TheMLEofQ(i;j)issimply
^Q(i;j)=dij
ri
i.e.,thenumberoffailures(deaths) duetocausejduring
thei-thintervalamongtherisubjectsatriskoffailure
atthebeginning oftheinterval.
(b)TheMLEof½iis::
^½i=ri¡Pm
j=1dij
ri=1¡Pm
j=1dij
ri
(c)TheMLEofST(t)isbasedon½i:
^ST(¿i)=iY
k=1½k
474
Sowhatcan'tweestimate?
Compare thecause-speci¯chazardfunction :
¸j(t)=lim
¢t!01
¢tPr(t·T<t+¢t;J=jjT¸t)
withthemarginalhazardfunction :
¸j(t)=lim
¢t!01
¢tPr(t·Tj<t+¢tjTj¸t)
Wecangetestimates ofthecause-speci¯chazardfunction,
sincewecanestimateST(t)=P(T¸t)evenifthefail-
uretimesaredependent.(Inotherwords,wecanobserve
whether eachpatientisstillaliveornot)
Butunfortunately ,wecan'testimate themarginal hazard
function whentherisksaredependent,sincewecan'testi-
mateSj(t)=P(Tj¸t).(wecan'ttellwhentheywould
havehadeventTjiftheyhaveadi®erentevent¯rst)
Thisisthemaintrickyissueofcompetingrisksanalyses.
475
Backtooriginalquestion...
Whatcanwedoifwesuspectdependentrisks?
Ex.Saywehavetwotypesoffailures,T1andT2,andwe
thinktheyaredependent.However,weareinterestedinthe
¯rsttypeoffailureT1,andviewthecompetingriskT2as
censoring (likeinbonemarrowtransplan texample).
Methodsofsummarizing datawithcompetingrisks:
(1)Summarize thecause-speci¯chazard rateovertime
(2)UsetheKaplan-Meier estimate anyway,^ST1(t)
(3)ReportthecomplementtotheKM,1¡^ST1(t)
(4)Usecumulativeincidence curves(crudeincidence
curve)
(5)Usetheconditionalprobabilityfunction
(6)Giveupperandlowerboundsforthetruemarginal
survivalfunction, intheabsenceofthecompetingrisk
PepeandMorireviewthe¯rst5oftheseoptions, andrec-
ommendagainstoption(2),butnotethatthisisoftenwhat
peopleendupdoing.
476
Tomaketheexamplemoreconcrete:
Sayweareinterested intimetoMACordeath,whichever
occurs¯rst.Wede¯ne:
T=timetoMACordeath
andU=censoring (assumed independent)
andthetypeoffailureisdenoted byj
j=8
><
>:MifeventisMAC
Difeventisdeathfromothercauses
Inthealternativ e\latentvariable"framework,wewouldde-
¯ne
TM=timetoMAC
andTD=timetoDeath
although wemightnotbeabletoobserveTMifTDoccurred
¯rst.
477
Methodsforcompetingrisks:
(1)Summarizingthecause-speci¯chazardover
time
Asmentionedabove,thisisoneofthequantitiesthatwe
canestimate. Ourbasicestimator duringtimeintervaliis
^Q(i;j)=dij
ri.
Sowecanplot^¸j(t)overtime,andgetsomeinsightasto
biological phenomona involved.
However,ifyouremembersomeoftheplotsIshowedyouof
hazardsovertime,theytendedtobehighlyvariable.Several
contributions havebeenmadetowards\smoothing"outthe
inherentvariabilityintheestimates ofcause-speci¯chazards.
²Efron(1988)
²Ramlau-Hansen (1983)
²TannerandWong(1983)
Drawback:thehazardfunctions alonedonotgiveoverall
e®ectofacovariateonsurvival.
Example: Ifthehazardfunctions for¸j(t)fortwotreat-
mentscross,wecan'tsaywhichtreatmen tleadstolower
overalleventrate.
478
FigurefromPepeandMoriforLeukemiaData
KernelEstimatesofCause-Speci¯cHazards
479
Methodsforcompetingrisks:
(2)ApplyingKaplan-Meier tocause-speci¯c¸j's
SayweevaluatetheMACsurvivaldistribution bytreating
²allMACcasesas\events"
²anydeathswithoutMACas\censorings"
andconstruct theKaplan-Meier survivalcurve.
Whatareweestimating?
S¤
M(t)=exp"
¡Zt
0¸M(u)du#
where¸M(t)isthecause-speci¯chazardforMAC:
¸M(t)=lim
¢t!01
¢tPr(t·T<t+¢t;j=MjT¸t)
i.e.,theconditional probabilit ythatMACoccursinashort
periodoftime,giventhatthesubjectisaliveandMAC-free.
Theinterpretation oftheKaplan-Meier curveisasthe
\exponentialofthenegativecumulativecause-
speci¯chazard".
Clinicians (andothers!) havedi±cultyunderstanding this
function, sinceithasnodirectclinicalinterpretation.
480
FigurefromPepeandMoriforLeukemiaData
KaplanMeierwithCauseSpeci¯cHazards
Theythoughtthiswassuchabadidea,thatthey
didnotincludeanyplotofthis!
481
Methodsforcompetingrisks:
(3)UsingthecomplementoftheKaplan-Meier
Another function usedfairlyofteninthecompetingrisks
areaissometimes referred toasthepureprobability
function :
1¡S¤
j(t)=1¡exp"
¡Zt
0¸j(u)du#
InourMACexample, 1¡S¤
M(t)couldbeinterprested asthe
predictiv eprobabilit yofMACbytimetiftheriskofdeath
couldberemoved.
Ifweweredesigning anewstudyforamiracledrugthat
seemedsopowerfulthatitwouldnotonlyreduceMACbut
preventalldeathfromothercausesinHIV-infected patients,
thenwecoulduseestimates 1¡^S¤
M(t)tohelpdesignour
newstudy.
Thiswouldbeprettyoptimistic, andtherehasbeenalotof
workonthestrict(anduntestable )assumptions required to
interprettheKMcurveinthismanner.
PepeandMoricontendthatthisfunction isirrelevantfor
summarizing datafromacompetingrisksstudy.
482
FigurefromPepeandMoriforLeukemiaData
ComplementKaplan-Meier Functions
483
Methodsforcompetingrisks:
(4)UsingCumulativeIncidenceCurves
Thishasalsobeentermedthe\crudeincidence curve",and
estimates themarginalprobabilityofaneventinthe
settingwhereothercompetingrisksareacknowledged toex-
ist.
DescriptionofMethod:themethodisdescribedin
moredetailinKalb°eisc handPrentice(p.169).
Testsforcovariates: Testsforcomparing cumulative
incidence curvesamongtreatmen tgroups(orsomeotherco-
variate)havebeendevelopedbyBobGray(1988). They
aresimilartologranktestsinthattheyare\linearrank"
statistics.
Ifwewereabletofollowupallsubjectstotimet,thenthe
cumulativeincidence curveswouldrelfectwhatproportion
ofthetotalstudypopulation havehadtheparticular event
(i.e.,MAC)bytimet.
484
FigurefromPepeandMoriforLeukemiaData
CumulativeIncidenceCurves
485
Methodsforcompetingrisks:
(5)ConditionalProbabilityCurves
Thishasthesame°avorasthecomplemen tKM,butamore
naturalinterpretation. PepeandMoride¯netheconditional
probabilit yfunction as:
dCPM(t)=P(TM·tjTD¸t)
=^PM(t)
1¡^PD(t)
where^PM(t)=Zt
0^ST(u)dNM(u)
Y(u)
and^PD(t)=Zt
0^ST(u)dND(u)
Y(u)
Intheabove,^ST(u)istheKMestimate oftheoverallmac-
freesurvivaldistribution, andthetermsNM(u),ND(u),and
Y(u)re°ectthe\countingprocess"forthenumberofsub-
jectswithMAC,death,andatriskattimeu,respectively.
Intheabsenceofcensoring, theinterpretation isthepropor-
tionofpatientswhodevelopMACamongthosewhodonot
dieofothercauses.
Tests:PepeandMorialsopresenttests,whicharesums
overtimeofweighteddi®erences betweentheconditional
(ormarginal) probabilities fortwogroups.
486
FigurefromPepeandMoriforLeukemiaData
ConditionalProbabilityandMarginalCurves
487
Methodsforcompetingrisks:
(6)BoundsonNetSurvivalCurves
Asnotedpreviously ,wecannotestimateSj(t)=P(Tj¸t)
ifthefailuretimesaredependent(eg,wecan'testimate the
survivalfunction forMACiftimetoMACiscorrelated with
timetodeathwithoutMAC).
However,wemaybeabletosaysomething abouttherange
ofSj(t)by¯ndingupperandlowerboundsthatcontain
Sj(t).
²Peterson (1976)obtained generalboundsbasedonthe
minimal andmaximal dependencestructure for(TM;TD).
Theboundsallowanypossibledependence structure,
butcanbeverywide.
²SludandRubenstein (1983)obtained tighterbounds
onSj(t)byusingadditional information, butrequirethe
usertospecifyreasonable boundsonafunction½.Once
½issupplied, themarginal distribution Sj;½(t)canbe
obtained.
²KleinandMoeshberger(1988)usetheframework
ofClaytonandOakesforbivariatesurvivaltoobtain
tighterboundsthanthoseofPeterson. Again,theuser
hastosupplyboundsonafunctionµ,andoncethisis
given^Sj;µ(t)canbeobtained.
488
FigurefromKleinandMoeschberger(1988)
BoundsonNetSurvivalCurves
489
Onelastexample:PromotionofFacultyatHSPH
Iwasaskedtoanalyzetheschool'sdatafrom1980-1995 on
promotion ofFacultyfromAssistantProfessor toAssociate
Professor toassesswhether thereweredi®erences between
malesandfemalesandamongacademic areas(social,labo-
ratory,quantitative).
Problem:
Wouldyouthinkthat\censoring" (someone leavingtheir
tenuretrackpositionpriortogettingpromoted) isindepen-
dentoftheprobabilit yofpromotion?
Iconsidered 3approachesforaccountingforcensoring:
MethodI:assumes thosewhodeparted wouldNOThave
beenpromoted
MethodII:assumes thosewhodeparted wouldhavebeen
promoted atthesamerateasthosewhostayed
MethodIII:assumes 50%ofthosewhodeparted wouldnot
havebeenpromoted, andtheother50%would
havebeenpromoted atthesamerateasthose
whostayed
WhichMethodcorrespondsto\non-informativ e"
(independent)censoring?
490
Results:Cumulativeprobabilitiesofpromotion
E®ectofGenderonPromotion
OverallMalesFemales
MethodI:0.6310.719 0.451
MethodII:0.9331.000 0.674
MethodIII:0.7360.825 0.531
E®ectofAcademicAreaonPromotion
OverallQuantitativeSocialLaboratory
MethodI:0.631 0.703 0.238 0.701
MethodII:0.933 0.950 0.389 1.000
MethodIII:0.736 0.803 0.287 0.801
491