MCMC Convergence and Tessellation Budget Sensitivity in the AddiVortes Model
Project Overview
The Additive Voronoi Tessellations (AddiVortes) model is a multivariate regression model that uses Voronoi tessellations to partition the covariate space in an additive ensemble model. Unlike other partition methods, such as decision trees, this has the benefit of allowing the boundaries of the partitions to be non-orthogonal
and nonparallel to the covariate axes. The AddiVortes model uses a similar sum-of-tessellations approach and a Bayesian backfitting MCMC algorithm to the BART
model. We use regularization priors to limit the strength of individual tessellations and accepts new models based on a likelihood.
Two lesser understood aspects of the AddiVortes model are MCMC convergence and budget sensitivity. We want to be able to better understand MCMC convergence behaviour as the dimensionality of the problem (the number of input
covariates), and the sample size grow. We also want to understand the sensitivity of the predictive accuracy of the model to changes in budget parameters, which can dramatically affect the compute time it takes to fit the model.
In this 8-week project, funded by N8 CIR through their undergraduate internship programme, I investigated these properties of the model, making use of the Hamilton HPC Service of Durham University. I worked directly with the existing codebases for the Python and R implementations of the algorithm, collaborating with my project supervisor, Prof. John Paul Gosling, to make optimisations which have had a direct impact on the speed of model fitting, as well as introducing features to aide the exploration of the key project aims, specifically convergence metrics.
What were the key results of your research project?
- In summary, throughout this project, we have been able to further investigate the parameters and convergence of the AddiVortes model, and have made the following progress.
- The Python implementation of the algorithm is preferred to the R implementation in terms of model strength and compute time, and following extensive analysis of the package code, drastic improvements have been made to both code bases, meaning the algorithm now runs significantly faster than it did before this project took place.
- Convergence diagnostics tools now exist within the package, allowing users to analyse the validity of the AddiVortes fits they create, and can be used as a tool for tuning mcmcIter and mcmcBurnin (two model parameters).
- The cleaned Earthquakes dataset in its current form is ready for wider use, and can be packaged with future AddiVortes versions, to act as a more robust model fitting test.
- We now have a better understanding of how the size of the dataset we are fitting affects key metrics such as fit time, and out-of-sample RMSE, in particular, the diminishing effect increasing n has on out-of-sample RMSE. This understanding can be applied to situations such as sequential data collection, where the practitioner can choose to stop data collection after some amount of data points n, where there may be some cost associated with data collection.
- We now have a much better understanding of how changing model parameters affects key metrics, and we have been able to make good recommendations of how the default parameters can be changed to improve these metrics
Additionally, the code written to carry out parameter screening will be used to aide the creation of an R package which contains an implementation of this method, with no current implementation existing.
GitHub Repository: https://github.com/snowy-the-leopard/AddiVortes
A presentation of this research will be shared here and on our YouTube site when available.
How do you feel you have benefitted from completing this internship and has it made you consider future career paths?
This internship has helped me develop my research skills and has made me consider further study. I am starting my master's degree this year and am starting to apply for PhD opportunities to explore further study.
Download presentation slides