Modelling, Inference, and Prediction for Positive and Unlabelled Data
- 2026-07-16 (Thu.), 10:30 AM
- 統計所B1演講廳;茶 會:上午10:10。
- 實體與線上視訊同步進行。
- Prof. Chi-Kuang Yeh (葉啓光 助理教授)
- Department of Mathematics and Statistics, Georgia State University
Abstract
Case-control is a study design widely used in biomedical research to investigate the causes of diseases. However, data contamination is a common issue in case-control studies due to, for instance, some medical conditions may go unrecognized in many patients, and they are misclassified as healthy one. This situation may be characterized as positive and unlabeled (PU) data. We introduce new approach to addressing through the double exponential tilting model (DETM). Traditional methods often fall short because they only apply to selected completely at random PU data, where the labeled positive and unlabeled positive data are assumed to be from the same distribution. In contrast, our DETM's dual structure effectively accommodates the more complex and underexplored selected at random PU data, where the labeled and unlabeled positive data can be from different distributions. Through theoretical insights and practical applications, this study highlights DETM as a comprehensive framework for addressing the challenges of PU data.
線上視訊請點選連結
