|
|
|
|
تکنولوژی خوشه بندی خودکار اسناد علمی بر مبنای الگوریتم چرخه آب
|
|
|
|
|
|
|
|
نویسنده
|
عبدالرزاقنژاد مجید ,هاشمزاده بهاره ,قاسمی عفت
|
|
منبع
|
تصميم گيري و تحقيق در عمليات - 1403 - دوره : 9 - شماره : 4 - صفحه:1064 -1086
|
|
چکیده
|
هدف: خوشهبندی متون با سازماندهی پیکرههای بزرگ متنی، نقش کلیدی در پیمایش مرور آسان انبوهی از متون دارد. یکی از قابلیتهای خوشهبندی متون در کنفرانسهای علمی، برای دستهبندی مقالات با موضوعات مشترک میباشد که کاربردهای زیادی در جستوجو و انتخاب مقالات دارد. هدف این تحقیق، بهبود کیفیت و سرعت خوشهبندی متون علمی بهویژه مقالات پژوهشی با تاکید بر تشخیص خودکار تعداد خوشهها و کاهش نیاز به تنظیمات دستی پارامترها است.روششناسی پژوهش: در این مقاله، یک روش خوشهبندی خودکار اسناد علمی جدید بر اساس الگوریتم چرخه آب (wca) ارایه میشود. ایده پیشنهادی متشکل از مراحل مختلف پیشپردازش، نمایش اسناد علمی بر اساس tf-idf سازگار شده برای اسناد علمی، تعریف مکانیزم فعال و غیرفعال شدن مراکز خوشهها از تعداد معینی مرکز خوشه بهمنظور ایجاد انعطاف در تعداد خوشههای اسناد علمی و الگوریتم چرخه آب بهمنظور بهینهیابی تعداد مراکز خوشه و مختصات آنها میباشد.یافتهها: در این مقاله از دو مجموعه داده استاندارد nips 2015 و aaai 2013 که حاوی اطلاعات مقالات ارایهشده به دو کنفرانس در حوزه یادگیری ماشین و هوش مصنوعی هستند، استفاده شده است. همچنین خوشهبندی خودکار بر اساس چهار الگوریتم فرا ابتکاری تکامل تفاضلی، ژنتیک، زنبورعسل و بهینهسازی ازدحام ذرات نیز بر روی دادههای استاندارد یادشده پیادهسازی شدهاند. از شاخص دیویس بودلین (db) و شاخص چو و سو (cs) جهت ارزیابی کیفیت نتایج بهدستآمده استفاده شده است. نتایج حاصل نشان میدهد که روش پیشنهادی در مقایسه با سایر روشهای فرا ابتکاری، کیفیت و کارایی بهتری در خوشهبندی اسناد علمی داشته و قادر به غلبه بر چالشهای خوشهبندی دادههای متنی نامتوازن و بزرگ مقیاس است.اصالت/ارزشافزوده علمی: در روش خوشهبندی خودکار پیشنهادی برای اولین بار از الگوریتم چرخه آب که توانایی سازگاری با دادههای ناهمگن و نامتوازن را دارد استفاده شده است. با توجه به اینکه مقالات علمی هم زمینه در یک مجله یا کنفرانس ارایهشده و در خوشهبندی این مستندات تحلیل آماری در شناسایی سریع کلمات کلیدی جایگاه ویژهای دارد، ترکیب tf-idf و مکانیزم فعال و غیرفعال شدن مراکز خوشه در فرآیند خوشهبندی اسناد علمی ارایه شده است.
|
|
کلیدواژه
|
متن کاوی، خوشهبندی خودکار متون علمی، tf-idf، الگوریتمهای فرا ابتکاری، الگوریتم چرخه آب
|
|
آدرس
|
دانشگاه صنعتی بیرجند, دانشکده مهندسی کامپیوتر و صنایع, گروه علوم کامپیوتر, ایسلند, دانشگاه الزهرا (س), گروه مهندسی کامپیوتر, ایران, دانشگاه آزاد اسلامی واحد بیرجند, گروه مهندسی کامپیوتر, ایران
|
|
پست الکترونیکی
|
ef.ghasemi@gmail.com
|
|
|
|
|
|
|
|
|
|
|
|
|
automatic scientific documents clustering technology based on water cycle algorithm
|
|
|
|
|
Authors
|
abdolrazzagh-nezhad majid ,hashemzadeh bahareh ,ghasemi efat
|
|
Abstract
|
purpose: the clustering of large-scale textual data plays a key role in the easy browsing and scrolling of huge documents by organizing their structures. one of its applications is scientific document clustering، which was presented or published in conferences and journals to categorize articles with common topics. the research's purpose is to improve the quality and speed of scientific document clustering and reduce the need for manual parameter settings.methodology: the paper proposes a novel automatic scientific documents clustering method based on the water cycle algorithm (wca). the proposed method consists of different stages of pre-processing، scientific document representation based on tf-idf adapted for scientific documents، defining the mechanism of activating and deactivating cluster centers from a certain number of cluster centers in order to create flexibility in the number of scientific document clusters and the wca to optimize the number of cluster centers and their coordinates.findings: in this paper، two benchmark datasets، nips 2015 and aaai 2013، are used، which contain information on articles presented at two conferences. also، the automatic clustering has been implemented based on four meta-heuristic algorithms: differential evolution algorithm، genetic algorithm، artificial bee colony algorithm and particle swarm optimization. davis-bodelin index and chu and su index were utilized to evaluate the quality of the obtained results. the comparison of the obtained results shows that the proposed method has better quality and efficiency in scientific document clustering and is able to overcome the challenges of unbalanced and large textual data clustering.originality/value: in the proposed automatic clustering method، the wca has been used for the first time. with regard to scientific articles in the same field are presented in a journal or conference، and statistical analysis has a special place to identify keywords of these documents in their clustering quickly، tf-idf and the mechanism of activation and deactivation of cluster centers have been combined in the proposed scientific documents clustering.
|
|
Keywords
|
text mining ,automatic scientific documents clustering ,tf-idf ,meta-heuristics ,water cycle algorithm
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|