CIVILICA We Respect the Science
(ناشر تخصصی کنفرانسهای کشور / شماره مجوز انتشارات از وزارت فرهنگ و ارشاد اسلامی: ۸۹۷۱)

AWS: Automatic Webpage Segmentation

عنوان مقاله: AWS: Automatic Webpage Segmentation
شناسه ملی مقاله: IRANWEB02_032
منتشر شده در دومین کنفرانس بین المللی وب پژوهی در سال 1395
مشخصات نویسندگان مقاله:

Mohammad Mehdi Yadollahi - Social Networks Lab., Faculty of Electrical and Computer Engineering, University of Tehran, Tehran, Iran
Masoud Asadpour - Social Networks Lab., Faculty of Electrical and ComputerEngineering, University of Tehran, Tehran, Iran

خلاصه مقاله:
a webpage contains many blocks of data, which can be informative or non-informative. In content extraction methods, informative data such as page title, headlines, news article and post body are distinguished from non-informative data such as advertisement, sidebar and navigational menus. The content extraction tasks have many difficulties because of the variety structure of webpages. In this paper, we proposed a content extraction method named Automatic Webpage Segmentation, AWS, which classifies the main content of a given webpage using a feature set consisting of structural and shallow text features. We benefit DOM tree of webpages for feature extraction. The obtained results are promising due to the effectiveness of proposed method to classify individual text elements of a webpage. Besides, feature selection methods such as wrapper and filter are utilized to improve performance of AWS.

کلمات کلیدی:
Content Extraction, Web Information Extraction, Full-text Extraction, Web Document Modeling

صفحه اختصاصی مقاله و دریافت فایل کامل: https://civilica.com/doc/481676/