Multiuser Resource Control With Deep Reinforcement Learning in IoT Edge Computing

By leveraging the concept of mobile edge computing (MEC), massive amount of data generated by a large number of Internet of Things (IoT) devices could be offloaded to MEC server at the edge of wireless network for further computational intensive processing. However, due to the resource constraint of...

Full description

Saved in:

Bibliographic Details
Published in	IEEE internet of things journal Vol. 6; no. 6; pp. 10119 - 10133
Main Authors	Lei, Lei, Xu, Huijuan, Xiong, Xiong, Zheng, Kan, Xiang, Wei, Wang, Xianbin
Format	Journal Article
Language	English
Published	Piscataway IEEE 01.12.2019 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subjects	Algorithms Approximation algorithms Computation offloading Computational modeling Computer simulation Deep learning Deep reinforcement learning (DRL) Edge computing Electronic devices Internet of Things Internet of Things (IoT) Machine learning Markov analysis Markov processes Mobile computing mobile edge computing (MEC) Neural networks Optimization Power consumption Servers Stochastic processes Task analysis Traffic delay Wireless communications Wireless networks
Online Access	Get full text
ISSN	2327-4662 2327-4662
DOI	10.1109/JIOT.2019.2935543

Cover

More Information
Summary:	By leveraging the concept of mobile edge computing (MEC), massive amount of data generated by a large number of Internet of Things (IoT) devices could be offloaded to MEC server at the edge of wireless network for further computational intensive processing. However, due to the resource constraint of IoT devices and wireless network, both communications and computation resources need to be allocated and scheduled efficiently for better system performance. In this article, we propose a joint computation off-loading and multiuser scheduling algorithm for IoT edge computing system to minimize the long-term average weighted sum of delay and power consumption under stochastic traffic arrival. We formulate the dynamic optimization problem as an infinite-horizon average-reward continuous-time Markov decision process (CTMDP) model. One critical challenge in solving this MDP problem for the multiuser resource control is the curse-of-dimensionality problem, where the state space of the MDP model and the computation complexity increase exponentially with the growing number of users or IoT devices. In order to overcome this challenge, we use the deep reinforcement learning (RL) techniques and propose a neural network architecture to approximate the value functions for the post-decision system states. The designed algorithm to solve the CTMDP problem supports semi distributed auction-based implementation, where the IoT devices submit bids to the BS to make the resource control decisions centrally. The simulation results show that the proposed algorithm provides significant performance improvement over the baseline algorithms, and also outperforms the RL algorithms based on other neural network architectures.
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14
ISSN:	2327-4662 2327-4662
DOI:	10.1109/JIOT.2019.2935543