Conference Proceedings

Width-based Lookaheads with Learnt Base Policies and Heuristics Over the Atari-2600 Benchmark

Stefan O'Toole, Nir Lipovetzky, Miquel Ramirez, Adrian Pearce, Marc'Aurelio Ranzato (ed.), Alina Beygelzimer (ed.), Khai Nguyen (ed.), Percy Liang (ed.), Jenn Wortman Vaughan (ed.), Yann Dauphin (ed.)

Advances in Neural Information Processing Systems 34 pre-proceedings (NeurIPS 2021) | NeurIPS | Published : 2021

Abstract

We propose new width-based planning and learning algorithms applied over the Atari-2600 benchmark. The algorithms presented are inspired from a careful analysis of the design decisions made by previous width-based planners. We benchmark our new algorithms over the Atari-2600 games and show that our best performing algorithm, RIW C +CPV, outperforms previously introduced width-based planning and learning algorithms π -IW(1), π -IW(1)+ and π -HIW(n, 1). Furthermore, we present a taxonomy of the set of Atari-2600 games according to some of their defining characteristics. This analysis of the games provides further insight into the behaviour and performance of the width-based algorithms intro..

View full abstract

Grants

Funding Acknowledgements

Funding in direct support of this work: Australian Government Research Training Program Scholarship provided by the Australian Commonwealth Government and the University of Melbourne; the Defence Science Institute, an initiative of the State Government of Victoria; and computing resources provided by the University of Melbourne through the Melbourne Research Cloud.