Lecture Notes in Computer Science, 2008, Volume 5323/2008, 41-54, DOI: 10.1007/978-3-540-89722-4_4

Efficient Reinforcement Learning in Parameterized Models: Discrete Parameter Case

Kirill Dyagilev, Shie Mannor and Nahum Shimkin

View Related Documents

Abstract

We consider reinforcement learning in the parameterized setup, where the model is known to belong to a finite set of Markov Decision Processes (MDPs) under the discounted return criteria. We propose an on-line algorithm for learning in such parameterized models, the Parameter Elimination (PEL) algorithm, and analyze its performance in terms of the total mistake bound criterion. The algorithm relies on Wald’s sequential probability ratio test to eliminate unlikely parameters, and uses an optimistic policy for effective exploration. We establish that, with high probability, the total mistake bound for the algorithm is linear (up to a logarithmic term) in the size of the parameter space, independently of the cardinality of the state and action spaces. We further demonstrate that much better dependence on is possible, depending on the specific information structure of the problem.

Fulltext Preview

Image of the first page of the fulltext document