首页> 外文会议>9th International conference on language resources and evaluation >Thematic Cohesion: measuring terms discriminatory power toward themes
【24h】

Thematic Cohesion: measuring terms discriminatory power toward themes

机译:主题衔接:衡量术语对主题的歧视性

获取原文

摘要

We present a new measure of thematic cohesion. This measure associates each term with a weight representing its discriminatory power toward a theme, this theme being itself expressed by a list of terms (a thematic lexicon). This thematic cohesion criterion can be used in many applications, such as query expansion, computer-assisted translation, or iterative construction of domain-specific lexicons and corpora. The measure is computed in two steps. First, a set of documents related to the terms is gathered from the Web by querying a Web search engine. Then, we produce an oriented co-occurrence graph, where vertices are the terms and edges represent the fact that two terms co-occur in a document. This graph can be interpreted as a recommendation graph, where two terms occurring in a same document means that they recommend each other. This leads to using a random walk algorithm that assigns a global importance value to each vertex of the graph. After observing the impact of various parameters on those importance values, we evaluate their correlation with retrieval effectiveness.
机译:我们提出了主题凝聚力的一种新方法。此度量将每个术语与代表其对主题的区分能力的权重相关联,该主题本身由一系列术语(主题词典)表示。该主题衔接标准可以用于许多应用程序中,例如查询扩展,计算机辅助翻译或特定领域词典和语料库的迭代构建。该度量分两步计算。首先,通过查询Web搜索引擎从Web收集与这些术语相关的一组文档。然后,我们生成一个定向的同时出现图,其中顶点是项,边表示文档中两个项同时出现的事实。该图可以解释为推荐图,在同一文档中出现的两个术语表示它们彼此推荐。这导致使用随机游走算法,该算法为图形的每个顶点分配全局重要性值。在观察了各种参数对那些重要性值的影响之后,我们评估了它们与检索有效性的相关性。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号