تفکیک کور منابع گفتار دوکاناله بر اساس مکان‌یابی

Fa | Ar | En

تفکیک کور منابع گفتار دوکاناله بر اساس مکان‌یابی


نویسنده	علی‌صوفی حسن ,خادمی مرتضی ,ابراهیمی مقدم عباس
منبع	مهندسي برق و مهندسي كامپيوتر ايران - 1400 - دوره : 19 - شماره : 1 - صفحه:59 -64
چکیده	در این مقاله یک روش جدید برای تفکیک کور منابع گفتار دوکاناله، بدون نیاز به دانش قبلی در مورد منابع گفتار آمده است. در روش پیشنهادی، با وزن‌دادن به طیف سیگنال ترکیب‌شده بر اساس فاصله منابع گفتار با میکروفون، تفکیک منابع گفتار انجام می‌شود. بنابراین ابتدا با تشکیل اسپکتوگرام زاویه‌ای توسط تابع همبستگی متقابل تعمیم‌یافته، منابع گفتار موجود در سیگنال ترکیب‌شده مکان‌یابی می‌شوند. سپس با توجه به موقعیت مکانی منابع از نظر فاصله با میکروفون‌ها، اندازه طیف سیگنال ترکیب‌شده، وزن‌دهی می‌شود. با ضرب اندازه طیف وزن داده شده در مقادیر حاصل از اسپکتوگرام زاویه‌ای و مقایسه آنها با هم، برای هر منبع یک نقاب باینری ساخته می‌شود. با اعمال نقاب باینری به اندازه طیف سیگنال ترکیب‌شده، منابع گفتار موجود در آن از هم جدا می‌شوند. این روش روی داده‌های پایگاه داده sisec آزمایش و از ابزار سنجش و معیارهای موجود در این پایگاه، برای ارزیابی استفاده شده است. نتایج نشان می‌دهد که روش پیشنهادی، از جهت معیارهای موجود در پایگاه مذکور با روش‌های رقیب قابل مقایسه بوده و پیچیدگی محاسباتی کمتری دارد.
کلیدواژه	اسپکتوگرام زاویه‌ای، تابع همبستگی متقابل تعمیم‌یافته، تفکیک کور منابع گفتار
آدرس	دانشگاه فردوسی مشهد, دانشکده مهندسی, گروه برق, ایران, دانشگاه فردوسی مشهد, دانشکده مهندسی, گروه برق, ایران, دانشگاه فردوسی مشهد, دانشکده مهندسی, گروه برق, ایران
پست الکترونیکی	a.ebrahimi@um.ac.ir

Blind TwoChannel Speech Source Separation Based on Localization

Authors	Alisufi Hassan ,Khademi M. ,Ebrahimi moghadam Abbas
Abstract	This paper presents a new method for blind twochannel speech sources separation without the need for prior knowledge about speech sources. In the proposed method, by weighting the mixture signal spectrum based on the location of the speech sources in terms of distance to the microphone, the speech sources are separated. Therefore, by forming an angular spectrum by generalized crosscorrelation function, the speech sources in the mixture signal are localized. First, by creating an angular spectrogram by generalized crosscorrelation function, the speech sources in the mixture signal are localized. Then according to the location of the sources, the amplitude of the mixture signal spectrum is weighted. By multiplying the weighted spectrum by the values obtained from the angular spectrograms, a binary mask is constructed for each source. By applying the binary mask to the amplitude of the mixture signal spectrum, the speech sources are separated. This method is evaluated on SiSEC database and the measurement tools and criteria contained in this database are used for evaluation. The results show that the proposed method is comparable in terms of the criteria available in the database to the competing ones, has lower computational complexity.
Keywords