Intelligent & Learning Agents
Bandit algorithms, MDP solvers, and a Minimax chess agent.
- Period
- August–November 2025
- Kind
- course
- Stack
- Python, Linear Programming
- Guide
- Prof. Shivaram Kalyanakrishnan, IIT Bombay
Three connected pieces of work on sequential decision-making.
- Regret minimisation — implemented and benchmarked UCB, KL-UCB and Thompson Sampling, including a computationally optimised KL-UCB variant for stochastic environments.
- Planning — modelled a strategic card game as a Markov Decision Process and solved it with Howard’s Policy Iteration and Linear Programming to derive optimal winning policies.
- Search — built a Microchess agent using Minimax with custom heuristics, tuning search depth against execution time.