>
Fa   |   Ar   |   En
   تکنولوژی خوشه بندی خودکار اسناد علمی بر مبنای الگوریتم چرخه آب  
   
نویسنده عبدالرزاق‌نژاد مجید ,هاشم‌زاده بهاره ,قاسمی عفت
منبع تصميم گيري و تحقيق در عمليات - 1403 - دوره : 9 - شماره : 4 - صفحه:1064 -1086
چکیده    هدف: خوشه‌بندی متون با سازماندهی پیکره‌های بزرگ متنی، نقش کلیدی در پیمایش مرور آسان انبوهی از متون دارد. یکی از قابلیت‌های خوشه‌بندی متون در کنفرانس‌های علمی، برای دسته‌بندی مقالات با موضوعات مشترک می‌باشد که کاربردهای زیادی در جست‌وجو و انتخاب مقالات دارد. هدف این تحقیق، بهبود کیفیت و سرعت خوشه‌بندی متون علمی به‌ویژه مقالات پژوهشی با تاکید بر تشخیص خودکار تعداد خوشه‌ها و کاهش نیاز به تنظیمات دستی پارامترها است.روش‌شناسی پژوهش: در این مقاله، یک روش خوشه‌بندی خودکار اسناد علمی جدید بر اساس الگوریتم چرخه آب (wca) ارایه می‌شود. ایده پیشنهادی متشکل از مراحل مختلف پیش‌پردازش، نمایش اسناد علمی بر اساس tf-idf سازگار شده برای اسناد علمی، تعریف مکانیزم فعال و غیرفعال شدن مراکز خوشه‌ها از تعداد معینی مرکز خوشه به‌منظور ایجاد انعطاف در تعداد خوشه‌های اسناد علمی و الگوریتم چرخه آب به‌منظور بهینه‌یابی تعداد مراکز خوشه و مختصات آن‌ها می‌باشد.یافته‌ها: در این مقاله از دو مجموعه داده استاندارد nips 2015 و aaai 2013 که حاوی اطلاعات مقالات ارایه‌شده به دو کنفرانس در حوزه یادگیری ماشین و هوش مصنوعی هستند، استفاده شده است. همچنین خوشه‌بندی خودکار بر اساس چهار الگوریتم فرا ابتکاری تکامل تفاضلی، ژنتیک، زنبورعسل و بهینه‌سازی ازدحام ذرات نیز بر روی داده‌های استاندارد یادشده پیاده‌سازی شده‌اند. از شاخص دیویس بودلین (db) و شاخص چو و سو (cs) جهت ارزیابی کیفیت نتایج به‌دست‌آمده استفاده شده است. نتایج حاصل نشان می‌دهد که روش پیشنهادی در مقایسه با سایر روش‌های فرا ابتکاری، کیفیت و کارایی بهتری در خوشه‌بندی اسناد علمی داشته و قادر به غلبه بر چالش‌های خوشه‌بندی داده‌های متنی نامتوازن و بزرگ مقیاس است.اصالت/ارزش‌افزوده علمی: در روش خوشه‌بندی خودکار پیشنهادی برای اولین بار از الگوریتم چرخه آب که توانایی سازگاری با داده‌های ناهمگن و نامتوازن را دارد استفاده شده است. با توجه به اینکه مقالات علمی هم زمینه در یک مجله یا کنفرانس ارایه‌شده و در خوشه‌بندی این مستندات تحلیل آماری در شناسایی سریع کلمات کلیدی جایگاه ویژه‌ای دارد، ترکیب tf-idf و مکانیزم فعال و غیرفعال شدن مراکز خوشه در فرآیند خوشه‌بندی اسناد علمی ارایه شده است.
کلیدواژه متن کاوی، خوشه‌بندی خودکار متون علمی، tf-idf، الگوریتم‌های فرا ابتکاری، الگوریتم چرخه آب
آدرس دانشگاه صنعتی بیرجند, دانشکده مهندسی کامپیوتر و صنایع, گروه علوم کامپیوتر, ایسلند, دانشگاه الزهرا (س), گروه مهندسی کامپیوتر, ایران, دانشگاه آزاد اسلامی واحد بیرجند, گروه مهندسی کامپیوتر, ایران
پست الکترونیکی ef.ghasemi@gmail.com
 
   automatic scientific documents clustering technology based on water cycle algorithm  
   
Authors abdolrazzagh-nezhad majid ,hashemzadeh bahareh ,ghasemi efat
Abstract    purpose: the clustering of large-scale textual data plays a key role in the easy browsing and scrolling of huge documents by organizing their structures. one of its applications is scientific document clustering، which was presented or published in conferences and journals to categorize articles with common topics. the research's purpose is to improve the quality and speed of scientific document clustering and reduce the need for manual parameter settings.methodology: the paper proposes a novel automatic scientific documents clustering method based on the water cycle algorithm (wca). the proposed method consists of different stages of pre-processing، scientific document representation based on tf-idf adapted for scientific documents، defining the mechanism of activating and deactivating cluster centers from a certain number of cluster centers in order to create flexibility in the number of scientific document clusters and the wca to optimize the number of cluster centers and their coordinates.findings: in this paper، two benchmark datasets، nips 2015 and aaai 2013، are used، which contain information on articles presented at two conferences. also، the automatic clustering has been implemented based on four meta-heuristic algorithms: differential evolution algorithm، genetic algorithm، artificial bee colony algorithm and particle swarm optimization. davis-bodelin index and chu and su index were utilized to evaluate the quality of the obtained results. the comparison of the obtained results shows that the proposed method has better quality and efficiency in scientific document clustering and is able to overcome the challenges of unbalanced and large textual data clustering.originality/value: in the proposed automatic clustering method، the wca has been used for the first time. with regard to scientific articles in the same field are presented in a journal or conference، and statistical analysis has a special place to identify keywords of these documents in their clustering quickly، tf-idf and the mechanism of activation and deactivation of cluster centers have been combined in the proposed scientific documents clustering.
Keywords text mining ,automatic scientific documents clustering ,tf-idf ,meta-heuristics ,water cycle algorithm
 
 

Copyright 2023
Islamic World Science Citation Center
All Rights Reserved