|
|
|
|
یک رویکرد مبتنی بر نمونهبرداری برای خوشهبندی کلاندادهها با استفاده از k-means
|
|
|
|
|
|
|
|
نویسنده
|
طلایی کاظم ,راحتی امین
|
|
منبع
|
اولين كنفرانس بين المللي هوش مصنوعي و فناوري هاي مرتبط - 1404 - دوره : 1 - اولین کنفرانس بین المللی هوش مصنوعی و فناوری های مرتبط - کد همایش: 04250-48654 - صفحه:0 -0
|
|
چکیده
|
امروزه کلاندادهها در بسیاری از زمینههای تحقیقاتی موثر بر دانش بشری مانند پزشکی، مهندسی و علوم کاربرد دارند. خوشهبندی از ابزارهای کلیدی در تحلیل انواع مختلف داده، بهویژه کلاندادهها شناخته میشود. در میان الگوریتمهای خوشهبندی، k-means بهدلیل سادگی و کارایی بالا از محبوبیت ویژهای برخوردار است. با این حال، عملکرد این الگوریتم به اندازهی مجموعه داده وابسته بوده و در مواجهه با کلاندادهها، همگرایی آن کند شده و عملکردش نیز ضعیف میشود. علاوه بر این، حساسیت بالای k-means نسبت به مقداردهی اولیهی مراکز خوشه میتواند منجر به دام افتادن آن در بهینههای محلی شود. ازاینرو در این مقاله، روش خوشهبندی جدیدی برای کلاندادهها ارائه میشود که از روش پایهی k-means استفاده میکند. ویژگی اصلی روش پیشنهادی، استفاده از تکنیک نمونهبرداری است که بهجای استفاده از تمام دادهها با زیرمجموعهای از آنها سروکار دارد. بر این اساس، روش پیشنهادی بهطور یکنواخت زیرمجموعهای از دادهها با اندازهی مشخص را انتخاب کرده و k-means را روی آن اعمال میکند. طی این فرآیند تکراری، هر موقع بهبودی در مقادیر تابع هدف مشاهده شود، مراکز حاصل از k-means بهعنوان مراکز خوشهِی آغازین زیرمجموعهی بعدی انتخاب میشوند. برای ارزیابی عملکرد روش پیشنهادی از 4 کلانداده دنیای واقعی استفاده شده و نتایج آن با الگوریتمهای k-means، forgy k-means و k-means++ مقایسه شده است. نتایج حاصل نشان میدهند روش پیشنهادی برحسب معیارهای rand index و زمان محاسباتی، روی 75 درصد دادهها عملکرد بهتری نسبت به سایر الگوریتمها ارائه کرده است.
|
|
کلیدواژه
|
خوشه بندی، کلانداده ها، means-k، نمونه برداری
|
|
آدرس
|
, iran, , iran
|
|
پست الکترونیکی
|
a.rahati@basu.ac.ir
|
|
|
|
|
|
|
|
|
|
|
|
|
a sampling-based approach for big data clustering using k-means
|
|
|
|
|
Authors
|
|
|
Abstract
|
nowadays, big data has increasingly become predominant in many research fields affecting human knowledge, including medicine, engineering, and science. clustering is widely recognized as one of the most effective processes to deal with various types of data, especially big data. k-means is one of the most popular clustering algorithms due to its simplicity and high efficiency. nevertheless, in the context of big data, its convergence significantly slows down and its effectiveness is reduced. in addition, the high sensitivity of k-means to the initialization of cluster centers can lead to entrapment in local optima. therefore, this paper presents a new clustering method for big data based on the basic k-means algorithm. the main feature of the proposed method is employing a sampling technique that selects a subset of the data and applies clustering. more precisely, the proposed method uniformly selects a subset of the data with a certain size and applies k-means to it. during this iterative process, whenever an improvement in the objective function values is observed, the centers obtained from k-means are selected as the initial cluster centers of the next subset. to evaluate the performance of the proposed method, four real-world big datasets are used and its results are compared with the k-means, forgy k-means, and k-means++ algorithms. the results show that the proposed method outperforms the other algorithms on 75% of the datasets in terms of the rand index and computational time criteria
|
|
Keywords
|
clustering ,big data ,k-means ,sampling
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|