LLM evaluation pipeline development

While LLMs are claimed to be useful for an ever growing range of tasks, systematic evaluation is required to produce evidence of their actual performance. A core component of an automated LLM evaluation is the data used to prompt the model, i.e. the user’s “side” of the transcript. These inputs can be designed in a … Read more

AI companionship evaluation on Roblox

Experiences centred on AI-powered chatbot “characters” – including some offering forms of companionship – are proliferating on the popular gaming platform Roblox. In parallel, the European Commission has designated Roblox as a Very Large Online Platform (VLOP) under the Digital Service Act as the platform surpasses 45 million average monthly users in the EU, and … Read more