GEMSS-Driven Subsampling for Information Extraction and Redundancy Elimination
- 2026-09-02 (Wed.), 14:00 PM
- 統計所B1演講廳;茶 會:下午13:40。
- 實體與線上視訊同步進行。
- Dr. Sheng-Zhan Hua (花聖展 先生)
- Department of Statistics and Data Science, UCLA
Abstract
Subsampling provides an effective strategy for addressing the computational and methodological challenges of applying statistical methods to large datasets. In particular, the training of Gaussian process models, which is notoriously difficult with large-scale data, can benefit substantially from subsampling in big-data contexts. In this study, we propose a subsampling methodology aimed at improving the predictive accuracy of Gaussian process models in unexplored regions of the input space. The proposed approach, termed Generalization Error Minimization in SubSampling (GEMSS), identifies informative subsets of data while discarding redundant points that may cause numerical instability. The development of GEMSS leverages an equivalence between Gaussian process models and linear models. Theoretical results establish its validity, and numerical studies across diverse scenarios demonstrate its effectiveness. An accompanying R package, GEMSS, has also been developed to facilitate implementation.
線上視訊請點選連結