Abstract:
Routinely collected data that are commonly used for research may be prone to errors. However, it is infeasible to validate all routinely collected data in large cohorts; data validation can be performed in subsamples of records. Restricting the analysis to the validation subset is inefficient, whereas using only error-prone data may yield biased estimates and misleading conclusions. We apply and compare novel methods for combining validation data with the original data to obtain estimates of regression coefficients. We describe the approaches and practical considerations that arose in an analysis of a large, multiwave validation study of Kaposi sarcoma (KS) among people with HIV. Investigators validated 939 of 257,429 eligible patient records across multiple international HIV cohorts in East Africa and Latin America. We used inverse probability weighting, generalized raking, and multiple imputation to analyze the original data and validation data together. We compare the approaches while investigating factors associated with three outcomes: prevalence of KS at enrollment into HIV care, incidence of KS over time in care, and time to death after KS diagnosis. Lastly, we discuss the advantages and disadvantages of each approach in the context of our study.