>
Fa   |   Ar   |   En
   automatic construction of persian ict wordnet using princeton wordnet  
   
نویسنده ahmadi tameh a. ,nassiri m. ,mansoorizadeh m.
منبع journal of ai and data mining - 2019 - دوره : 7 - شماره : 1 - صفحه:109 -119
چکیده    Wordnet is a large lexical database of the english language in which nouns, verbs, adjectives, and adverbs are grouped into sets of cognitive synonyms (synsets). each synset expresses a distinct concept. synsets are interlinked by both semantic and lexical relations. wordnet is essentially used for word sense disambiguation, information retrieval, and text translation. in this paper, we propose several automatic methods to extract information and communication technology (ict)-related data from princeton wordnet. we then add these extracted data to our persian wordnet. the advantage of automated methods is to reduce the interference of human factors and accelerate the development of our bilingual ict wordnet. in our first proposed method, based on a small subset of ict words, we use the definition of each synset to decide whether that synset is ict. the second mechanism is to extract the synsets that are in a semantic relation with the ict synsets. we also use two similarity criteria, namely lcs and s3m, to measure the similarity between a synset definition in wordnet and definition of any word in microsoft dictionary. our last method is to verify the coordinate of ict synsets. the results obtained show that our proposed mechanisms are able to extract the ict data from princeton wordnet at a good level of accuracy.
کلیدواژه wordnet; semantic relation; synset; part of speech; information and communication technology
آدرس bu-ali sina university, faculty of engineering, computer department, iran, bu-ali sina university, faculty of engineering, computer department, iran, bu-ali sina university, faculty of engineering, computer department, iran
پست الکترونیکی mansoorm@basu.ac.ir
 
     
   
Authors
  
 
 

Copyright 2023
Islamic World Science Citation Center
All Rights Reserved