Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Probability and Statistics / Survival Analysis

p244-26.pdf survival primer

PDF · 9 pages · 204.9 KB
Open PDF file

A SAS conference paper (Paper 244-26) by Tyler Smith and Besa Smith of the Naval Health Research Center, San Diego, kept in the archive as a survival analysis primer. It covers the history of survival analysis, the pdf, cdf, survivor and hazard functions, Kaplan-Meier estimation, and left and right censoring. It then looks at time-dependent covariates, the extended Cox model, and using the PROC PHREG baseline option to estimate survival functions for hospitalization risk.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
Paper 244 -26 Survival Analysis And The Application Of Cox's Proportional Hazards Modeling Using SAS Tyler Smith, and Besa Smith, Department of Defense Center for Deployment Health Research, Naval Health Research Center, San Diego, CA Abstract In recent papers published in the American Journalof Epidemiology, the authors used Cox'sproportional hazards regression modeling to modelthe time until an event of interest and compare thecumulative probability of hospitalization over timefor two or more cohorts while adjusting for otherinfluential covariates. In this presentation thesestatistical procedures will be looked at more closelyby using SAS. The usefulness of the baselineoption in PROC PHREG will be demonstrated withthe creation and output of survival functionestimates, which are a function of the cumulativeprobability estimates over time.This real data example using Cox modeling willshow what increased risk for hospitalization for anevent of interest might look like graphically and in arisk of event type ratio. Further analyses using thesurvivor function estimates will look for graphicalrepresentation of a temporal bias during theobservation period. The SAS system's PROCPHREG with baseline option was instrumental insearching for possible temporal biases and dealingwith attrition of subjects over our study period. Introduction to Survival Analysis The term "survival analysis" pertains to a statisticalapproach designed to take into account the amountof time an experimental unit contributes to a study.That is, it is the study of time between entry intoobservation and a subsequent event. Originally,the event of interest was death hence the term,"survival analysis." The analysis consisted offollowing the subject until death. The uses in thesurvival analysis of today vary quite a bit.Applications now include time until onset ofdisease, time until stockmarket crash, time untilequipment failure, time until earthquake, and so on.The best way to define such events is simply torealize that these events are a transition from onediscrete state to another at an instantaneousmoment in time. Of course, the term"instantaneous", which may be years, months,days, minutes, or seconds, is relative and has onlythe boundaries set by the researcher. The History of Survival Analysis The origin of survival analysis goes back tomortality tables from centuries ago. However, it wasnot until World War II that a new era of survivalanalysis emerged. This new era was stimulated byinterest in reliability (or failure time) of militaryequipment. At the end of the war these newlydeveloped statistical methods emerging from strictmortality data research to failure time research,quickly spread through private industry ascustomers became more demanding of safer, morereliable products. As the uses of survival analysisgrew, parametric models gave way tononparametric and semiparametric approaches fortheir appeal in dealing with the ever-growing field ofclinical trials in medical research. Survival analysiswas well suited for such work because medicalintervention follow-up studies could start without allexperimental units enrolled at start of observationtime and could end before all experimental unitshad experienced an event. This is extremelyimportant because even in the best-developedstudies, there will be subjects who choose to quitparticipating, who move too far away to follow, orwho will die from some unrelated event. Theresearcher was no longer forced to withdraw theexperimental unit and all associating data from thestudy, instead techniques called censoring enabledresearchers to analyze incomplete data due todelayed entry or withdrawal from the study. Thiswas important in allowing each experimental unit tocontribute all of the information possible to themodel for the amount of time the researcher wasable to observe the unit.The last great strides in the application of survivalanalysis techniques has been a direct result of the Statistics, Data Analysis, and Data Mining availability of software packages and highperformance computers which are now able to runthese difficult and computationally intensivealgorithms relatively efficiently. Some Tools Used in Survival Analysis First, recall that time is continuous, which results inthe probability of an event at a single point of acontinuous distribution being zero. We arechallenged to define the probability of these eventsover distribution. This is best described by graphing the distribution of event times. To ensurethe readers will start with the same fundamentaltools of survival analysis, a brief descriptive sectionof these important concepts will follow. A moredetailed description of the probability densityfunction, the cumulative distribution function, thehazard function, and the survivor function, can befound in any intermediate level statistical textbook.So that the reader will be able to look for certainrelationships while reading, it is important to notebefore the brief descriptions the one-to-onerelationship that these four functions possess. Thepdf can be obtained by taking the derivative of thecdf and likewise, the cdf can be obtained by takingthe integral of the pdf. The survivor function issimply 1 minus the cdf. Which leaves the hazardfunction as simply being the pdf over the survivorfunction. It will be these relationships later that willallow us to calculate the cdf from the survivorfunction estimates that the SAS procedure PROCPHREG will output. The Cumulative Distribution Function The cumulative distribution function (cdf) is veryuseful in describing the continuous probabilitydistribution of a random variable, such as time, in asurvival analysis. The cdf of a random variable T,denoted F T (t), is defined by FT (t) = PT (T /G23 t). This is interpreted as a function that will give theprobability that the variable T will be less than orequal to any value t that we choose. Severalproperties of a distribution function F(t) can be listedas a consequence of the knowledge of probabilities.Because F(t) has the probability 0 /G23 F(t) /G23 1, then F(t) is a nondecreasing function of t, and as tapproaches /G34, F(t) approaches 1.The Probability Density Function The probability density function (pdf) is also veryuseful in describing the continuous probabilitydistribution of a random variable. The pdf of arandom variable T, denoted f T(t), is defined by fT(t) = d FT (t) / dt. That is, the pdf is the derivative or slope of the cdf. Every continuous random variablehas its own density function, the probability P(a /G23 T /G23 b) is the area under the curve between times a and b. The Survival Function Let T /G24 0 have a pdf f(t) and cdf F(t). Then the survival function takes on the following form:S(t) = P{T > t} = 1 - F(t)That is, the survival function gives the probability ofsurviving or being event-free beyond time t.Because S(t) is a probability, it is positive andranges from 0 to 1. It is defined as S(0) = 1 and ast approaches /G34, S(t) approaches 0. The Kaplan- Meier estimator, or product limit estimator, is theestimator used by most software packages becauseof the simplistic step idea. The Kaplan-Meierestimator incorporates information from all of theobservations available, both censored anduncensored, by considering any point in time as aseries of steps defined by the observed survival andcensored times. The survival curve describes therelationship between the probability of survival andtime. The Hazard Function The hazard function h(t) is given by the following:h(t) = P{ t < T < (t + /G29) | T >t} = f(t) / (1 - F(t)) = f(t) / S(t)The hazard function describes the concept of therisk of an outcome (e.g., death, failure,hospitalization) in an interval after time t, conditionalon the subject having survived to time t. It is theprobability that an individual dies somewherebetween t and t + /G29, divided by the probability that the individual survived beyond time t. The hazardfunction seems to be more intuitive to use in Statistics, Data Analysis, and Data Mining survival analysis than the pdf because it attempts toquantify the instantaneous risk that an event willtake place at time t given that the subject survivedto time t. Incomplete Data Observation time has two components that must becarefully defined in the beginning of any survivalanalysis. There is a beginning point of the studywhere time=0 and a reason or cause for theobservation of time to end. For example, in acomplete observation cancer study, observation ofsurvival time may begin on the day a subject isdiagnosed with the cancer and end when thatsubject dies as a result of the cancer. This subjectis what is called an uncensored subject, resultingfrom the event occurring within the time period ofobservation. Complete observation time data likethis example are desired but not realistic in moststudies. There is always a good possibility that thepatient might recover completely or the patientmight die due to an entirely unrelated cause. Inother words, the study cannot go on indefinitely,waiting for an event from a participant andunforeseen things happen to study participants thatmake them unavailable for observation. Thecensoring of study participants therefore deals withthe problems of incomplete observations of timedue to assumed random factors not related to the study design. Note: This differs from truncation where observations of time are incomplete due to aselection process inherent to the study design. Left and Right Censoring The most common form of incomplete data is rightcensoring. This occurs when there is a definedtime (t=0) where the observation of time is startedfor all subjects involved in the study. A rightcensored subject's time terminates before theoutcome of interest is observed. For example, asubject could move out of town, die of anunexpected cause, or could simply choose not toparticipate in the study any longer. Right censoringtechniques allow subjects to contribute to the modeluntil they are no longer able to contribute (end ofthe study, or withdrawal), or they have an event.Conversely, an observation is left censored if theevent of interest has already occurred whenobservation of time begins. For the purposes of thisstudy we focused on right censoring.The following graph shows a simple study designwhere the observation times start at a consistentpoint in time (t=0). The X's represent events andthe O's represent censored observations. Noticethat all observations are classified with an event, orthey are censored at time of separation or at theend of the study period. Some subjects haveevents early in the study period and others haveevents at the end of the study period. Likewisesome subjects leave early, but most do not have anevent during the entire study and are simply rightcensored at the end. There is no need for leftcensoring or truncation techniques in this simpleexample. Time Dependencies In some situations the researcher may find that thedynamical nature of a variable causes changes invalue over the observation time. In other instancesthe researcher may find that certain trends effectthe probability of the event of interest over time.There are easy ways to test and account for thesetemporal biases within PROC PHREG but becareful if you have a large number of observationsas the computation of the subsequent partiallikelihood is very taxing and time consuming. Aneasier way to see if there are existing temporalbiases is to look at the plots of the cumulativedistributions of the probability of event. If there is asteep increase or decrease in the cumulativeprobability, it may suggest more investigation isneeded. It is important to note here that when atime dependent variable is introduced into themodel, the ratios of the hazards will not remainsteady. This only effects the model structure. Wewill still be doing a Cox regression but instead themodel used is called the extended Cox model. Statistics, Data Analysis, and Data Mining The Studies The underlining purpose of these studies was toinvestigate the effect of a specific exposure on anoutcome of interest. Then we sought to identify 2 ormore cohorts who might have had this exposure orsome degree of exposure and compare thehospitalization experiences for certain outcomes ofinterest to another similar cohort without thatparticular exposure. Demographic data Demographic data available for analysis includedsocial security numbers (for linking purposes only),gender, date of birth, race, ethnicity, home ofrecord, marital status, military occupational status,military pay grade, length of military service,deployment status, salary, date of separation frommilitary service, military service branch, andexposure status. Hospitalization Data Data describing hospitalization experiences werecaptured from all United States Department ofDefense military treatment facilities for the period ofOctober 1, 1988, through December 31, 1999. Theactual observation period varied by study. Removalof personnel with diagnoses of interest prior to thestart of the study follow-up period was completed.These data included date of admission in a hospitaland up to eight discharge diagnoses associatedwith the admission to the hospital. Additionally, apreexposure period covariate (coded as yes or no)was used to reflect a hospital admission during the12 months prior to the start of the exposure period.Note: the exposure period was the year fromAugust 1, 1990, to August 1, 1991. Diagnoseswere coded according to the International Classification of Diseases, Ninth Revision (ICD-9). For these analyses, we scanned for the specific 3-,4-, or 5-digit component of the ICD-9 diagnoses. Observation Time The focus of each study was to see if a certainexposure or lack of exposure had any influence onthe targeted disease outcomes. For each subject,hospitalizations (if any) were scanned inchronological order and diagnostic fields werescanned in numerical order for the ICD-9 codes ofinterest. Only the first hospitalization meeting theoutcome criteria was counted for each subject.Subjects were classified as having an event if theywere hospitalized in any Department of Defensehospital facility worldwide with the targeteddiagnoses, and as censored otherwise.Observation time varied with the dates we chose tostart and end observation but was calculated fromthe start of follow-up until event, separation frommilitary service, or the end of the study period,whichever occurred first. Subjects were allowed toleave the study and assumed a random earlydeparture distribution. Delayed entry and eventsoccurring before the start date of the study were nota concern, therefore only right censoring wasneeded to allow for the random early departure ofsubjects (see previous graph). Cox's Proportional Hazards Regression There are several reasons Cox's proportionalhazards modeling was chosen to explain the effectof covariates on time until event. They arediscussed below and include: the relative risk, noparametric assumptions, the use of the partiallikelihood function, and the creation of survivorfunction estimates. Relative Risk The simple interpretation given by the Cox modelas "relative risk" type ratio is very desirable inexplaining the risk of event for a certain covariate.For example, when we have a two-level covariatewith a value of 0 or 1, the hazard ratio becomes e /G24. If the value of the coefficient is /G24 = ln(3) then it is simply saying that the subjects labeled with a 1 arethree times more likely to have an event than thesubjects labeled with a 0. In this way we had ameasure of difference between our exposurecohorts instead of simply knowing whether theywere different. No Parametric Assumptions Another attractive feature of Cox regression is nothaving to choose the density function of aparametric distribution. This means that Cox'ssemiparametric modeling allows for no assumptionsto be made about the parametric distribution of thesurvival times, making the method considerablymore robust. Instead, the researcher must onlyvalidate the assumption that the hazards areproportional over time. The proportional hazardsassumption refers to the fact that the hazardfunctions are multiplicatively related. That is, theirratio is assumed constant over survival time, Statistics, Data Analysis, and Data Mining thereby not allowing a temporal bias to become aninfluential player on the endpoint. Use of the Partial Likelihood Function The Cox model has the flexibility to introduce time-dependent explanatory variables and handlecensoring of survival times due to its use of thepartial likelihood function. This was important to ourstudy in that any temporal biases due to differencesin hospitalization practices for different strata of thesignificant covariates over the years of studyneeded to be handled correctly. This ensured thatany differences in hospitalization experiencesbetween the exposed and nonexposed would notbe coming from these temporal differences. Survivor Function Estimates With the SAS option BASELINE, a SAS datasetcontaining survival function estimates can becreated and output. These estimates correspond tothe means of the explanatory variables for eachstratum. Analysis Univariate Analyses Using PROC FREQ, and PROC UNIVARIATE, aninitial univariate analysis of the demographicvariables crossed with hospitalization experiencewas carried out to determine possible significantexplanatory variables to be included in the modelruns. All variables with a chi-square value or tstatistic of .15 or less were considered possiblysignificant and were therefore retained for themodel analysis. Additionally the distributions ofattrition were checked to see if the cohortsseparated from active duty military service equally. Modeling Approach Using PROC PHREG, a saturated Cox model wasrun after creating dummy variables, necessary forthe output of hazard ratios for the categoricalexplanatory variables. A manual backwardstepwise analysis was carried out to create a modelwith statistically significant effects of explanatoryvariables on survival times.Programming PROC PHREG DATA=ANALYDAT; MODEL INHOSP*CENSOR(0)= expose1 pwhspstatus1 sex1 age1-age3 ms1 paygr1-paygr2oc_cat1-oc_cat9 ccep /RL TIES=EFRON ;TITLE1 'Cox Regression With Exposure Status Inthe Model ';RUN;The options used in this survival analysis procedureare described below:DATA=ANALYDAT names the input data set for the survival analysis.RL requests for each explanatory variable, the 95% (the default alpha level because the ALPHA= optionis not invoked) confidence limits for the hazardratios.TIES=EFRON gives the researcher the approximations to the EXACT method without usingthe tremendous CPU it takes to run the EXACTmethod. Both the EFRON and the BRESLOWmethods do reasonably well at approximating theEXACT when there are not a lot of ties. If there area lot of ties, then the BRESLOW approximation ofthe EXACT will be very poor. If the time scale is notcontinuous and is therefore discrete, the optionTIES=DISCRETE should be used. Stratification By Exposure Status These data were then stratified by exposure andthe models were run with the exposure flagcovariate withdrawn from the model. This allowedfor inspection of interaction between exposurestatus and covariates. Running these separatemodels also allowed for the computation of survivalfunction estimates using the BASELINE function inPROC PHREG. The survival curves (which arereally step functions for such numerous events thatthey appear continuous) were now available tocompute the cumulative distribution function for theseparate cohorts. Time Dependent Covariates After the final model of significant explanatoryvariables was created, it was necessary to validatethe proportional hazards assumption. If theresearcher believes that there may be a timedependency from a certain variable then simply add Statistics, Data Analysis, and Data Mining x1time to the list of independent variables and thefollowing below the model statement.x1time=x1*(t)Where t is the time variable and x1 is the suspectedtime dependent variable.If the interaction term is found to be insignificant wecan conclude that the proportional hazardsassumption holds. This is necessary to ensure thatthere was no adverse effect from time-dependentcovariates creating different rates for differentsubjects, thus making the ratios of their hazardsnonconstant. Survivor Function Estimates By Exposure The following is the code used after the ANALYDATwas stratified into exposed or nonexposed. Thisproduces the survivor function estimates byexposure while simultaneously checking to see ifthere were any interactions between the covariatesand the exposure status.PROC PHREG DATA=EXPOSE1; MODEL INHOSP*CENSOR(0)=pwhsp status1sex1 age1-age3 ms1 paygr1-paygr2 oc_cat1-oc_cat9 ccep /RL TIES=EFRON ; BASELINE OUT=SURVS SURVIVAL=S;RUN;The new options used in this survival analysisprocedure are described below:BASELINE without the COVARIATES= option produces the survivor function estimatescorresponding to the means of the explanatoryvariables for each stratum.OUT=SURVS names the data set output by the BASELINE option.SURVIVAL=S tells SAS to produce the survivor function estimates in the output data set.A simple calculation of 1-Survivor functionestimates in SURVS, obtained from running theBASELINE option, produced the cumulativedistribution functions. We could now see thecumulative probability estimates of hospitalizationover time. We were then able to visually scan fordifferences in hospitalization experiences ofbetween the cohorts and look for insight as towhether or not the proportional hazards assumptionhad been violated. The Plots Figure 1 is what the cumulative distribution functionwould look like if there were a violation of theproportional hazards assumption. Note the sharpincrease in probability of hospitalization beginningright before the third year and lasting forapproximately 1 year. After this one year period thetop curve then levels off and becomes parallelagain with the bottom curve.Figure 1. Figure 2 shows what the cumulative distributionfunction would look like if there were no violation ofthe assumption of proportional hazards but theredid happen to be an observed significant differencein the disease experience between the two cohortsover the length of the study period.Figure 2. Statistics, Data Analysis, and Data Mining Figure 3 shows what the cumulative distributionfunction would look like if there were no problemwith the assumption of proportional hazards. Thefigure also shows what the curves would look like itthere were not a significant difference observedbetween the diagnosis experience of the twocohorts.Figure 3. Years from March 10, 19915 4 3 2 1 0Probability of Hospitalization.4 .3 .2 .1 0.0 Computing the Generalized R2 Recently I was asked whether SAS computed theR 2 value and what it was for that particular model. If the researcher desires, the R2 value can be computed easily from the output of your regression,although it is not an option of PROC PHREG.Simply compute R2 = 1 - exp(LR2/n) Where LR is the Likelihood-ratio chi-square statisticfor testing the null hypothesis that all variablesincluded in the model have coefficients of 0, and nis the number of observations. The researcherneeds to take extreme caution when comparing theR 2 values of Cox regression models. Remember from linear regression analysis, R2 can be artificially increased by simply adding explanatory variables tothe regression model (ie; more variables does notequal a better model necessarily). Also, the abovecomputation does not give you the proportion ofvariance of the dependent variable explained by theindependent variables as it would in linearregression, but does give you a measure of howassociated the independent variables are with thedependent variable.Residual Analysis A residual analysis is very important especially ifthe sample size is relatively small. Add thefollowing after your model statement to output themartingale and deviance residuals:BASELINE OUT=SURVS SURVIVAL=SXBETA=XBET RESMART=MARTINGRESDEV=RDEV;Then a simple plot of the residuals against thelinear predictor scores will give the researcher anidea of the fit or lack of fit of the model to individualobservations.PROC GPLOT DATA=SURVS; PLOT (MARTING RDEV) * XBET / VREF=0; SYMBOL1 VALUE=CIRCLE; Results Using the initial univariate comparisons for eventsoccurring during the study, the following variableswere selected for the subsequent model analyses:gender, age group, marital status, race/ethnicity,military occupational category, military pay grade,salary, service branch, pre-exposure periodhospitalization, and exposure status. Home ofrecord was not shown to be significantly affectingthe endpoints in any either of the studies' modelsand was dropped. Salary and length of servicewere dropped from analyses due to colinearity withage.The subjects in cohort 1 had similar risks for two ofthe three diseases during the August 1, 1991, toJuly 31, 1997 study period compared with subjectsin cohort 2. The corresponding cumulativeprobability plots (above) were nearly parallel for thefollow-up period. However, the Cox model didreveal some consistently better predictors ofhospitalization with the two diseases, whichincluded female gender, preexposure periodhospitalization, enlisted pay grade, and US Reserveservice type.The subjects in cohort 1 had significantly differentrisks for the third disease during the August 1,1991, to July 31, 1997 study period compared withsubjects in cohort 2. The corresponding cumulativeprobability plots (above) were nearly parallel for thefirst three years of follow-up, then there was adrastic increase in hospitalization for a period ofabout 1 year and then once again the curvesbecame nearly parallel again. Time-dependent Statistics, Data Analysis, and Data Mining variables included in the modeling also confirmedthis result. The hazards ratio was not significantlygreater than 1 for the first 3 years of follow-up andthe last 3 years as well. However, during the 1 yearin question, though, the risk of hospitalization withthat particular disease was almost 3 times that ofcohort 1. Further investigation found that certaintreatment facilities had adopted an approach ofadministratively hospitalizing these subjects forextensive clinical evaluations. This approach waslater dropped almost 1 year to the date of the start. Conclusions The Cox proportional hazard model's robust natureallows us to closely approximate the results for thecorrect parametric model when the parametric isunknown or in question. Using the SAS® systemprocedure PROC PHREG, Cox's proportionalhazards modeling was used to compare thehospitalization experiences of two or more cohorts.The two studies produced models suggesting noincrease in risk among the exposed which werelater confirmed by producing the cumulativedistribution function. The study also found at leastone model which initially suggested an increase inrisk among the exposed. Further analysis revealedthat this sharp increase in risk for approximately 1year was likely due to an outside factor affecting theprocess of hospitalizing personnel for this particulardisease event.The SAS® system's PROC PHREG with censoringand the baseline option is a powerful tool forhandling early departure of subjects during thestudy period. It is also useful for producing datasets, including survival function estimates, whichcan be used in a simple equation to produceestimates of probability of events. When graphed,these show cumulative probability of event curvesas a function of time. If it were not for the graphs ofthe cumulative distribution functions, which showeda sharp temporal bias, results may have beenreported and interpreted with much different results. References Hosmer JR. DW, Lemeshow S. Applied Survival Analysis; Regression Modeling of Time to EventData. New York: John Wiley & Sons; 1999 Kleinbaum DG, Survival Analysis: A self-Learning Text. New York: Springer-Verlag; 1996SAS Institute Inc., SAS/STAT® User's Guide, Version 6, Fourth Edition, Volume 1, Cary, NC: SAS Institute Inc., 1989. 943 pp.SAS Institute Inc., SAS/STAT® User's Guide, Version 6, Fourth Edition, Volume 2, Cary, NC: SAS Institute Inc., 1989. 846 pp.SAS Institute Inc. SAS/STAT® Software: Changes and Enhancements through Release 6.11 . Cary, NC: SAS Institute Inc., 1996. 1104 pp.Allison, Paul D., Survival Analysis Using the SAS® system: A Practical Guide , Cary, NC: SAS Institute Inc., 1995. 292 pp.Smith TC, Gray GC, Knoke JD. Is systemic lupus erythematosus, amyotrophic lateral sclerosis, orfibromyalgia associated with Persian Gulf Warservice? An examination of Department of Defensehospitalization. Amer J of Epidemiol; June 1, 2000, Vol 151.Gray GC, Smith TC, Knoke JD, Heller JM. The postwar hospitalization experience among Gulf Warveterans exposed to chemical munitions destructionat Khamisiyah, Iraq. Amer J of Epidemiol; September 1, 1999 - Volume 150 No 5. Acknowledgments Thank you to Navy CAPT Greg Gray, Director ofthe Department of Defense Center for DeploymentHealth Research at the Naval Health ResearchCenter, San Diego. CAPT Gray's support andencouragement of learning new concepts anddevising better ways to tackle statistical problems inpursuit of the best possible research answers.Approved for public release: distribution unlimited.This research was supported by the Department ofDefense, Health Affairs, under work unit no. 60002.SAS software is a registered trademark of SASInstitute, Inc. in the USA and other countries. About The Authors Besa Smith has used SAS for 4 years includingwork as an undergraduate in Biology and Chemistryat CSU, Chico, a graduate at the SDSU GraduateSchool of Public Health in Biostatistics, andcurrently as a biostatistician with the DoD Center forDeployment Health Research at NHRC. Statistics, Data Analysis, and Data Mining Responsibilities include management of largemilitary and demographic data, mathematicalmodeling and statistical analysis. She has been aninvited presenter for WUSS 2000, and the SanDiego Area's SAS users Group fall 2000 meeting.Besa Smith, MPHBiostatistician, Henry Jackson FoundationDepartment of Defense Center for DeploymentHealth Research, at the Naval Health ResearchCenter, San Diego(619) [email protected] Smith has used SAS for 9 years, includingwork as a student in Math and Statistics, as agraduate student at the University of KentuckyDepartment of Statistics, and currently as a seniorstatistician and data analyst with the Naval HealthResearch Center. His responsibilities includemathematical modeling, analysis, management,and documentation of large hospitalization anddemographic data sets. He has been invited tospeak at the International Biometrics Societymeetings, WUSS 99, SUGI 2000, WUSS 2000,and the San Diego SAS users group 1999 and2000 fall meetings.Tyler C. Smith, MSSenior Statistician, Henry Jackson FoundationDepartment of Defense Center for DeploymentHealth Research, at the Naval Health ResearchCenter, San Diego(619) [email protected] Statistics, Data Analysis, and Data Mining