Exploration-Exploitation in MDPs with Options

被引:0
|
作者
Fruit, Ronan [1 ]
Lazaric, Alessandro [1 ]
机构
[1] Inria Lille, SequeL Team, Villeneuve Dascq, France
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
While a large body of empirical results show that temporally-extended actions and options may significantly affect the learning performance of an agent, the theoretical understanding of how and when options can be beneficial in online reinforcement learning is relatively limited. In this paper, we derive an upper and lower bound on the regret of a variant of UCRL using options. While we first analyze the algorithm in the general case of semi-Markov decision processes (SMDPs), we show how these results can be translated to the specific case of MDPs with options and we illustrate simple scenarios in which the regret of learning with options can be provably much smaller than the regret suffered when learning with primitive actions.
引用
收藏
页码:576 / 584
页数:9
相关论文
共 50 条
  • [1] The exploration-exploitation dilemma in pain
    Krypotos, Angelos
    [J]. PSYCHOSOMATIC MEDICINE, 2020, 82 (06) : A166 - A166
  • [2] Explaining Exploration-Exploitation in Humans
    Candelieri, Antonio
    Ponti, Andrea
    Archetti, Francesco
    [J]. BIG DATA AND COGNITIVE COMPUTING, 2022, 6 (04)
  • [3] Social Learning and the Exploration-Exploitation Tradeoff
    Mintz, Brian
    Fu, Feng
    [J]. COMPUTATION, 2023, 11 (05)
  • [4] The Exploration-Exploitation Dilemma: A Multidisciplinary Framework
    Berger-Tal, Oded
    Nathan, Jonathan
    Meron, Ehud
    Saltz, David
    [J]. PLOS ONE, 2014, 9 (04):
  • [5] Unpacking the exploration-exploitation tradeoff on Snapchat: The relationships between users' exploration-exploitation interests and server log data
    Gomez-Zara, Diego
    Liu, Yozen
    Neves, Leonardo
    Shah, Neil
    Bos, Maarten W.
    [J]. COMPUTERS IN HUMAN BEHAVIOR, 2024, 150
  • [6] SPACE - EXPLORATION-EXPLOITATION AND THE ROLE OF MAN
    LOFTUS, JP
    [J]. AVIATION SPACE AND ENVIRONMENTAL MEDICINE, 1986, 57 (10): : A69 - A77
  • [7] The two facets of the exploration-exploitation dilemma
    Zhang, Kaifu
    Pan, Wei
    [J]. 2006 IEEE/WIC/ACM INTERNATIONAL CONFERENCE ON INTELLIGENT AGENT TECHNOLOGY, PROCEEDINGS, 2006, : 371 - +
  • [8] Approximate information for efficient exploration-exploitation strategies
    Barbier-Chebbah, Alex
    Vestergaard, Christian L.
    Masson, Jean-Baptiste
    [J]. PHYSICAL REVIEW E, 2024, 109 (05)
  • [9] Exploration-exploitation and acquisition likelihood in new ventures
    Keyhani, Mohammad
    Deutsch, Yuval
    Madhok, Anoop
    Levesque, Moren
    [J]. SMALL BUSINESS ECONOMICS, 2022, 58 (03) : 1475 - 1496
  • [10] Adaptive Exploration-Exploitation Tradeoff for Opportunistic Bandits
    Wu, Huasen
    Guo, Xueying
    Liu, Xin
    [J]. INTERNATIONAL CONFERENCE ON MACHINE LEARNING, VOL 80, 2018, 80