>
Fa   |   Ar   |   En
   employing chaos theory for exploration–exploitation balance in deep reinforcement learning  
   
نویسنده khodadadi habib ,derhami vali
منبع مهندسي برق دانشگاه تبريز - 2025 - دوره : 55 - شماره : 1 - صفحه:113 -121
چکیده    Deep reinforcement learning is widely used in machine learning problems and the use of methods to improve its performance is important. balance between exploration and exploitation is one of the important issues in reinforcement learning and for this purpose, action selection methods that involve exploration such as ɛ-greedy and soft-max are used. in these methods, by generating random numbers and evaluating the action-value, an action is selected that can maintain this balance. over time, with appropriate exploration, it can be expected that the environment becomes better understood and more valuable actions are identified. chaos, with features such as high sensitivity to initial conditions, non-periodicity, unpredictability, exploration of all possible search space states, and pseudo-random behavior, has many applications. in this article, numbers generated by chaotic systems are used for the ɛ-greedy action selection method in deep reinforcement learning to improve the balance between exploration and exploitation; in addition, the impact of using chaos in replay buffer will also be investigated. experiments conducted in the lunar lander environment demonstrate a significant increase in learning speed and higher rewards in this environment
کلیدواژه action selection ,chaos theory ,deep reinforcement learning ,exploration and exploitation
آدرس yazd university, department of computer engineering, iran, yazd university, department of computer engineering, iran
پست الکترونیکی vderhami@yazd.ac.ir
 
   employing chaos theory for exploration–exploitation balance in deep reinforcement learning  
   
Authors khodadadi habib ,derhami vali
Abstract    deep reinforcement learning is widely used in machine learning problems and the use of methods to improve its performance is important. balance between exploration and exploitation is one of the important issues in reinforcement learning and for this purpose, action selection methods that involve exploration such as ɛ-greedy and soft-max are used. in these methods, by generating random numbers and evaluating the action-value, an action is selected that can maintain this balance. over time, with appropriate exploration, it can be expected that the environment becomes better understood and more valuable actions are identified. chaos, with features such as high sensitivity to initial conditions, non-periodicity, unpredictability, exploration of all possible search space states, and pseudo-random behavior, has many applications. in this article, numbers generated by chaotic systems are used for the ɛ-greedy action selection method in deep reinforcement learning to improve the balance between exploration and exploitation; in addition, the impact of using chaos in replay buffer will also be investigated. experiments conducted in the lunar lander environment demonstrate a significant increase in learning speed and higher rewards in this environment
 
 

Copyright 2023
Islamic World Science Citation Center
All Rights Reserved