>
Fa   |   Ar   |   En
   یک رویکرد مبتنی بر نمونه‌برداری برای خوشه‌بندی کلان‌داده‌ها با استفاده از k-means  
   
نویسنده طلایی کاظم ,راحتی امین
منبع اولين كنفرانس بين المللي هوش مصنوعي و فناوري هاي مرتبط - 1404 - دوره : 1 - اولین کنفرانس بین المللی هوش مصنوعی و فناوری های مرتبط - کد همایش: 04250-48654 - صفحه:0 -0
چکیده    امروزه کلان‌داده‌ها در بسیاری از زمینه‌های تحقیقاتی موثر بر دانش بشری مانند پزشکی، مهندسی و علوم کاربرد دارند. خوشه‌بندی از ابزارهای کلیدی در تحلیل انواع مختلف داده‌، به‌ویژه کلان‌داده‌ها شناخته می‌شود. در میان الگوریتم‌های خوشه‌بندی، k-means به‌دلیل سادگی و کارایی بالا از محبوبیت ویژه‌ای برخوردار است. با این حال، عملکرد این الگوریتم به اندازه‌ی مجموعه داده وابسته بوده و در مواجهه با کلان‌داده‌ها، همگرایی آن کند شده و عملکردش نیز ضعیف می‌شود. علاوه بر این، حساسیت بالای k-means نسبت به مقداردهی اولیه‌ی مراکز خوشه می‌تواند منجر به‌ دام افتادن آن در بهینه‌های محلی شود. ازاین‌رو در این مقاله، روش خوشه‌بندی جدیدی برای کلان‌‌داده‌ها ارائه می‌شود که از روش‌ پایه‌ی k-means استفاده می‌کند. ویژگی اصلی روش پیشنهادی، استفاده از تکنیک نمونه‌برداری است که به‌جای استفاده از تمام داده‌ها با زیرمجموعه‌ای از آنها سروکار دارد. بر این اساس، روش پیشنهادی به‌طور یکنواخت زیرمجموعه‌ای از داده‌ها با اندازه‌ی مشخص را انتخاب کرده و k-means را روی آن اعمال می‌کند. طی این فرآیند تکراری، هر موقع بهبودی در مقادیر تابع هدف مشاهده شود، مراکز حاصل از k-means به‌عنوان مراکز خوشه‌ِ‌ی آغازین زیرمجموعه‌ی بعدی انتخاب می‌شوند. برای ارزیابی عملکرد روش پیشنهادی از 4 کلان‌داده دنیای واقعی استفاده شده و نتایج آن با الگوریتم‌های k-means، forgy k-means و k-means++ مقایسه شده است. نتایج حاصل نشان می‌دهند روش پیشنهادی برحسب معیارهای rand index و زمان محاسباتی، روی 75 درصد داده‌ها عملکرد بهتری نسبت به سایر الگوریتم‌ها ارائه کرده است.
کلیدواژه خوشه بندی، کلانداده ها، means-k، نمونه برداری
آدرس , iran, , iran
پست الکترونیکی a.rahati@basu.ac.ir
 
   a sampling-based approach for big data clustering using k-means  
   
Authors
Abstract    nowadays, big data has increasingly become predominant in many research fields affecting human knowledge, including medicine, engineering, and science. clustering is widely recognized as one of the most effective processes to deal with various types of data, especially big data. k-means is one of the most popular clustering algorithms due to its simplicity and high efficiency. nevertheless, in the context of big data, its convergence significantly slows down and its effectiveness is reduced. in addition, the high sensitivity of k-means to the initialization of cluster centers can lead to entrapment in local optima. therefore, this paper presents a new clustering method for big data based on the basic k-means algorithm. the main feature of the proposed method is employing a sampling technique that selects a subset of the data and applies clustering. more precisely, the proposed method uniformly selects a subset of the data with a certain size and applies k-means to it. during this iterative process, whenever an improvement in the objective function values ​​is observed, the centers obtained from k-means are selected as the initial cluster centers of the next subset. to evaluate the performance of the proposed method, four real-world big datasets are used and its results are compared with the k-means, forgy k-means, and k-means++ algorithms. the results show that the proposed method outperforms the other algorithms on 75% of the datasets in terms of the rand index and computational time criteria
Keywords clustering ,big data ,k-means ,sampling
 
 

Copyright 2023
Islamic World Science Citation Center
All Rights Reserved