Evaluation and Explanation of Guardrail Models in LLMs

This project evaluates guardrail or safety models designed to detect, prevent, or filter unsafe and undesirable outputs from large language models (LLMs). The focus is on developing a systematic evaluation framework and comparing the effectiveness of different guardrail approaches.

This project would be ideal for BA ICS/MCS, BA CS (JH), BA CSLL and integrated masters. Experience with python programming and an avid interest in machine learning is desirable. Experience with pytorch and a strong track record of projects on Github is a plus.