Large language models can be used as autonomous agents that interact with applications through a browser, potentially allowing tests to be expressed as high-level goals rather than detailed scripts. This project’s vision is to investigate whether such agents can effectively explore and test web applications, and which factors affect their performance. The idea could be to implement or adapt an agent-based testing prototype and evaluate it on open source web applications, measuring properties such as successful task completion, coverage, failures discovered, reproducibility and execution cost. A particular question could be whether autonomous agents offer advantages over conventional scripted or automated exploration approaches.
References:
[1] Anna Arnaudo, Riccardo Coppola, Maurizio Morisio, Flavio Giobergia, Van-Thanh Nguyen, Enrico Chen, Minh-Thai Mai, Xiaoquan Ji, Xiaoning Ma, “Automated Black-Box Testing: A Comparative Study of LLM Agent Architectures and Prompt Engineering”, ICST 2026
[2] Antoine Chevrot, Alexandre Vernotte, Jean-Rémy Falleri (LaBRI), Xavier Blanc (LaBRI), Bruno Legeard, “Are Autonomous Web Agents Good Testers?”, ISSTA 2025