This project investigates interpretable machine learning methods that explain model predictions using examples. The focus will be on prototypes, which represent typical examples of a class or decision, and counterfactuals, which show how an input would need to change to receive a different prediction. Students will explore and compare state-of-the-art (SOTA) approaches. Interpretability in prediction is something frontier AI labs care deeply about.
This book gives a brief and general overview of the topic – https://christophm.github.io/interpretable-ml-book/
This project would be ideal for BA ICS/MCS, BA CS (JH), BA CSLL and integrated masters. Experience with python programming and an avid interest in machine learning is desirable. Experience with pytorch and a strong track record of projects on Github is a plus.