针对为检索服务的语义知识库存在的内容不全面和不准确的问题,提出一种基于维基百科的软件工程领域概念语义知识库的构建方法.以SWEBOK V3概念为标准,从维基百科提取概念的解释文本,并抽取其关键词表示概念的语义;通过概念在维基百科中的层次关系、概念与其他概念的解释文本关键词之间的链接关系、不同概念解释文本关键词之间的链接关系构成概念语义知识库;利用LDA主题模型分别与TF-IDF、TextRank算法相结合的两种方法抽取关键词;对构建好的概念语义知识库用随机游走算法计算概念间的语义相似度.将实验结果与人工标注结果对比后发现,本方法构建的语义知识库语义相似度准确率能够达到84%以上,充分验证了所提方法的有效性.%The problem of incomplete and inaccurate content for the retrieval of semantic knowledge base existed,this paper proposed a method of constructing the concept semantic knowledge base in the field of software engineering based on Wikipedia.First,taking the concept of SWEBOK V3 as the standard,it extracted the interpretation of the concept from Wikipedia and extracted the keywords to represent the semantic meaning of the concept.Second,through hierarchical relationships of the concept in Wikipedia,link relationships between concepts and explanatory text of other concepts and link relationships between explanatory texts of different concepts,it built concept semantic knowledge base.Then,it combined the LDA topic model with the two methods that were called TF-IDF algorithm and TextRank algorithm respectively serve the keywords extraction.Finally,it calculated the semantic similarity between concepts by the random walk algorithm for the construction of the concept semantic knowledge base.The experimental results were compared with the manual annotation results.The semantic similarity of knowledge base constructed by this method can reach more than 84%.The effectiveness of the proposed method is verified.
展开▼