Recent work in my research group has shown that AI feedback can help balance competing objectives, like vehicle delay, emissions, and fairness across approaches, in traffic signal control, using preferences over individual decisions to guide reinforcement learning. However, decision-making in traffic control in sequential and long-horizon; a signal-timing choice that is suboptimal in the short-term might set up a much better outcome three or four steps later (e.g., clearing a queue before a predicted surge). This project asks whether learning from trajectory-level preferences (eg comparing whole sequences of decisions rather than single steps) produces better, more strategically sound traffic control policies than step-by-step feedback alone. You’ll extend an existing RLAIF-based traffic control pipeline to generate and compare trajectory segments, explore how to elicit or synthesize meaningful multi-step preferences, and evaluate whether trajectory-level optimization improves long-horizon traffic outcomes over the current single-step approach.
Note: Suitable for MSc-level student and requires prior RL experience. Please outline your experience in the email.
Reference:
Chenyang Zhao, Vinny Cahill, and Ivana Dusparic. “Balancing Multiple Objectives in Urban Traffic Control with Reinforcement Learning from AI Feedback.” 21st International Conference on Software Engineering for Adaptive and Self-Managing Systems (SEAMS 2026).