Distinguishing Real Web Crawlers from Fakes: Googlebot Example

机译：从假货中区分真实的Web爬虫：GoogleBot示例

获取原文

页面导航

摘要
著录项
相似文献
相关主题

摘要

Web crawlers are programs or automated scripts that scan web pages methodically to create indexes. such as Google, Bing use crawlers in order to provide web surfers with relevant information. Today there are also many crawlers that impersonate well-known web crawlers. For example, it has been observed that Google's Googlebot crawler is impersonated to a high degree. This raises ethical and security concerns as they can potentially be used for malicious purposes. In this paper, we present an effective methodology to detect fake Googlebot crawlers by analyzing web access logs. We propose using Markov chain mohels to learn profiles of real and fake Googlebots based on their patterns of web resource access sequences. We have calculated log-odds ratios for a given set of crawler sessions and our results show that the higher the log-odds score, the higher the probability that a given sequence comes from the real Googlebot. Experimental results show, at a threshold log-odds score we can distinguish the real Googlebot from the fake.

机译：Web爬网程序是程序或自动脚本，可维护Web页面以创建索引。如谷歌，Bing使用爬虫，以便为Web冲浪者提供相关信息。今天，也有许多追逐众所周知的网络爬虫。例如，已经观察到谷歌的GoogleBot履带冒充高度。这提高了道德和安全问题，因为它们可能被用于恶意目的。在本文中，我们提出了一种通过分析Web Access日志来检测假GoogleBot爬虫的有效方法。我们建议使用Markov Chain Mohels根据其Web资源访问序列模式学习真实和假Googlebots的配置文件。我们已经计算了一套给定的爬网赛会话的log-odds比率，我们的结果表明，Log-odds评分越高，给定序列来自真正的GoogleBot的概率越高。实验结果表明，在阈值对数量的分数下，我们可以将真正的Googlebot与假的真实的Googlebot区分开来。

著录项

来源
《International Moratuwa Engineering Research Conference》|2018年|611p|共6页
会议地点
作者
Nilani Algiryage;
展开▼
作者单位

展开▼
会议组织
原文格式 PDF
正文语种
中图分类输配电工程、电力网及电力系统;
关键词
Crawlers; Markov processes; Google; Web sites; Robots; Web servers; Search engines;

机译：爬虫;马尔可夫进程;谷歌;网站;机器人;Web服务器;搜索引擎;

相似文献

外文文献
中文文献
专利

1. Analyzing and distinguishing fake and real news to mitigate the problem of disinformation [J] . Vereshchaka Alina, Cosimini Seth, Dong Wen Computational & Mathematical Organization Theory . 2020,第3期

机译：分析与区分虚假的真正新闻，减轻虚假信息问题
2. Distinguishing real from fake ivory products by elemental analyses: A Bayesian hybrid classification method [J] . Buddhachat Kittisak, Brown Janine L., Thitaram Chatchote, Forensic science international . 2017,第期

机译：通过元素分析区分真实的象牙产品：贝叶斯混合分类方法
3. Semiconductor parts: Distinguishing real from fake requires a trained eye [J] . Steve Martin ECN . 2013,第13期

机译：半导体零件：辨别真伪需要训练有素的眼睛
4. Distinguishing Real Web Crawlers from Fakes: Googlebot Example [C] . Nilani Algiryage 4th International Moratuwa Engineering Research Conference . 2018

机译：区分真实的Web爬虫与假冒：Googlebot示例
5. Constructing Web Crawlers for the World Art Dynamics Technology Platform [D] . Guo, Xueyuan. 2019

机译：为世界艺术动力学技术平台构建网络爬虫
6. A user-oriented web crawler for selectively acquiring online content in e-health research [O] . Songhua Xu, Hong-Jun Yoon, Georgia Tourassi -1

机译：面向用户的网络爬虫用于在电子卫生研究中选择性地获取在线内容
7. Applying Clickstream Data Mining to Real-Time Web Crawler Detection and Containment Using ClickTips Platform [O] . Anália Lourenço, O Belo 2013

机译：使用ClickTips平台将Clickstream数据挖掘应用于实时Web爬网程序检测和遏制

Distinguishing Real Web Crawlers from Fakes: Googlebot Example

摘要

著录项

相似文献

相关主题

期刊订阅