بررسی روش‌های موثر بر عملکرد تجزیه‌گر دستور مستقل از متن آماری زبان فارسی

Fa | Ar | En

بررسی روش‌های موثر بر عملکرد تجزیه‌گر دستور مستقل از متن آماری زبان فارسی


نویسنده	صادق زاده محمدباقر ,رزازی محمدرضا ,قیومی مسعود
منبع	پردازش علائم و داده ها - 1398 - شماره : 3 - صفحه:23 -36
چکیده	عدم دقّت در طراحی دستورهای مستقل از متن و استفاده از ساختارهای نامناسب مانند فرم نرمال چامسکی به خودی خود می ‌تواند عملکرد تجزیه‌ zwj; گرهای آماری مستقل از متن را تضعیف کند. در این پژوهش ساختار ترکیبات عطفی درخت‌بانک فارسی را مورد بررسی قرار دادیم. نتایج حاصل از این پژوهش نشان می‌ دهد که با اضافه‌کردن وابستگی ‌های ساختاری به دستورهای مستقل از متن و اصلاح قواعد اولیه، می‌‌توان از ترکیبات عطفی رفع ابهام کرد و صحت عملکرد تجزیه ‌گر دستور مستقل از متن آماری را افزایش داد. فرض استقلال ضعیف، یکی از مشکلات مربوط به دستورهای مستقل از متن است که سعی شده است تا با تزریق وابستگی ‌های ساختاری از طریق نشانه ‌گذاری گره‌ های والد و فرزند مرتفع شود. تاثیر ریزدانگی و درشت دانگی برچسب‌ های اجزای واژگانی کلام و همین‌طور ادغام ناپایانه ‌ها بر تجزیه‌ گر دستور مستقل از متن آماری فارسی از جمله مواردِ مورد بررسی قرار گرفته‌شده در این پژوهش است.
کلیدواژه	دستور مستقل از متن آماری، تجزیه‌گر، ترکیبات عطفی، نشانه‌گذاری قواعد، برچسب اجزای واژگانی کلام
آدرس	دانشگاه صنعتی امیرکبیر, ایران, دانشگاه صنعتی امیرکبیر, ایران, پژوهشگاه علوم انسانی و مطالعات فرهنگی, ایران

Studying impressive parameters on the performance of Persian probabilistic context free grammar parser

Authors	sadeghzadeh mohammadbagher ,razzazi mohammadreza ,ghayoomi Masood
Abstract	In linguistics, a tree bank is a parsed text corpus that annotates syntactic or semantic sentence structure. The exploitation of tree bank data has been important ever since the first largescale tree bank, The Penn Treebank, was published. However, although originating in computational linguistics, the value of tree bank is becoming more widely appreciated in linguistics research as a whole. For example, annotated tree bank data has been crucial in syntactic research to test linguistic theories of sentence structure against large quantities of naturally occurring examples.The natural language parser consists of two basic parts, POS tagger and the syntax parser. A PartOfSpeech Tagger (POS Tagger) is a piece of software that reads text in some languages and assigns parts of speech to each word (and other token), such as noun, verb, adjective, etc., although generally computational applications use more finegrained POS tags like 'nounplural'. A natural language parser is a program that works out the grammatical structure of sentences, for instance, which groups of words go together (as phrases ) and which words are the subject or object of a verb.Probabilistic parsers use knowledge of language gained from handparsed sentences to try to produce the most likely analysis of new sentences. These statistical parsers still make some mistakes, but commonly work rather well. Inaccurate design of contextfree grammars and using bad structures such as Chomsky normal form can reduce accuracy of probabilistic contextfree grammar parser. Weak independence assumption is one of the problems related to CFG. We have tried to improve this problem with parent and child annotation, which copies the label of a parent node onto the labels of its children, and it can improve the performance of a PCFG.In grammar, a conjunction (conj) is a part of speech that connects words, phrases, or clauses that are called the conjuncts of the conjunctions. In this study, we examined the conjunction phrases in the Persian tree bank. The results of this study show that adding structural dependencies to grammars and modifying the basic rules can remove conjunction ambiguity and increase accuracy of probabilistic contextfree grammar parser.When a partofspeech (PoS) tagger assigns word class labels to tokens, it has to select from a set of possible labels whose size usually ranges from fifty to several hundred labels depending on the language. In this study, we have investigated the effect of fine and coarse grain POS tags and merging nonterminals on Persian PCFG parser.
Keywords	Probabilistic context free grammar ,parser ,tree bank ,conjunction phrases ,parent annotation ,child annotation ,part of speech tags