This project evaluates guardrail or safety models designed to detect, prevent, or filter unsafe and undesirable outputs from large language models (LLMs). The focus is on developing a systematic evaluation framework and comparing the effectiveness of different guardrail approaches.
This project would be ideal for BA ICS/MCS, BA CS (JH), BA CSLL and integrated masters. Experience with python programming and an avid interest in machine learning is desirable. Experience with pytorch and a strong track record of projects on Github is a plus.