首页> 中文期刊> 《新疆农业大学学报》 >中文农业搜索引擎字符编码识别

中文农业搜索引擎字符编码识别

         

摘要

针对农业网页中汉字编码标识混乱的情况,提出了一种综合运用编码规则和网页文本特征的字符编码识别模型。利用卡方检验算法,结合最小二乘多元线性回归方法,得到了基于网页文本特征的字符识别模型。实验结果显示,在适当的选取阈值(r =1,阈值=属于某一编码的字符数/网页总字符数)和文本特征数(≥65)的基础上,模型准确率达到100%,且结果稳定。%The character encoding identification model comprehensively using encoding rules and Web page text was put forward in accordance with the confused conditions of character encoding identification in Chinese agriculture web page.Using the chi-square test algorithm,combining with the method of least square multivariale linear regression,the model of character identification based on Web page text feature was obtained.The experimental results showed that the accuracy of the model reached 100% and the result was stable on the basis of the appropriate threshold selected (r =1,threshold=the number of characters belonging to a coding/the total number of web page)and the text feature number (≥65).

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号