首页> 外文会议>9th International conference on language resources and evaluation >Thematic Cohesion: measuring terms discriminatory power toward themes
【24h】

Thematic Cohesion: measuring terms discriminatory power toward themes

机译:主题凝聚力:测量歧视性致歧视权给主题

获取原文

摘要

We present a new measure of thematic cohesion. This measure associates each term with a weight representing its discriminatory power toward a theme, this theme being itself expressed by a list of terms (a thematic lexicon). This thematic cohesion criterion can be used in many applications, such as query expansion, computer-assisted translation, or iterative construction of domain-specific lexicons and corpora. The measure is computed in two steps. First, a set of documents related to the terms is gathered from the Web by querying a Web search engine. Then, we produce an oriented co-occurrence graph, where vertices are the terms and edges represent the fact that two terms co-occur in a document. This graph can be interpreted as a recommendation graph, where two terms occurring in a same document means that they recommend each other. This leads to using a random walk algorithm that assigns a global importance value to each vertex of the graph. After observing the impact of various parameters on those importance values, we evaluate their correlation with retrieval effectiveness.
机译:我们提出了一种新的主​​题凝聚力。该措施将每个术语与代表其歧视权给主题的权重相关联,此主题本身由术语列表(主题词典)表示。该主题凝聚力标准可用于许多应用,例如查询扩展,计算机辅助翻译或域名lexicons和corpora的迭代构造。该度量分两步计算。首先,通过查询Web搜索引擎从Web收集与术语相关的一组文档。然后,我们生成面向导向的共同发生图,其中顶点是术语,边缘代表两个术语在文档中发生的事实。该图可以解释为推荐图,其中在同一文件中发生的两个术语意味着它们互相推荐。这导致使用随机漫游算法,将全局重要性值分配给图形的每个顶点。在观察各种参数对这些重要性值的影响之后,我们评估其与检索效果的相关性。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号