Using Empirically Constructed Lexical Resources for Named Entity Recognition

机译：使用经验构造的词汇资源进行命名实体识别

代理获取

本网站仅为用户提供外文OA文献查询和代理获取服务，本网站没有原文。下单后我们将采用程序或人工为您竭诚获取高质量的原文，但由于OA文献来源多样且变更频繁，仍可能出现获取不到、文献不完整或与标题不符等情况，如果获取不到我们将提供退款服务。请知悉。

页面导航

摘要
著录项
相似文献
相关主题

摘要

Because of privacy concerns and the expense involved in creating an annotated corpus, the existing small-annotated corpora might not have sufficient examples for learning to statistically extract all the named-entities precisely. In this work, we evaluate what value may lie in automatically generated features based on distributional semantics when using machine-learning named entity recognition (NER). The features we generated and experimented with include n-nearest words, support vector machine (SVM)-regions, and term clustering, all of which are considered distributional semantic features. The addition of the n-nearest words feature resulted in a greater increase in F-score than by using a manually constructed lexicon to a baseline system. Although the need for relatively small-annotated corpora for retraining is not obviated, lexicons empirically derived from unannotated text can not only supplement manually created lexicons, but also replace them. This phenomenon is observed in extracting concepts from both biomedical literature and clinical notes.

机译：由于隐私问题和创建带注释的语料库所涉及的费用，现有的带小注释的语料库可能没有足够的示例来学习精确地统计提取所有命名实体。在这项工作中，我们使用机器学习命名实体识别（NER）时，会基于分布语义评估自动生成的要素中可能具有的价值。我们生成和试验的特征包括n个最近词，支持向量机（SVM）区域和术语聚类，所有这些都被认为是分布语义特征。与通过使用手动构建的词典到基线系统相比，增加了n个最近字词功能可导致F分数的更大提高。尽管没有消除对相对较小注释的语料库进行再培训的需求，但凭经验从未经注释的文本派生的词典不仅可以补充手动创建的词典，还可以替换它们。从生物医学文献和临床笔记中提取概念时都观察到这种现象。

著录项

期刊名称 Biomedical Informatics Insights
作者
Siddhartha Jonnalagadda; Trevor Cohen; Stephen Wu; Hongfang Liu; Graciela Gonzalez;
展开▼
作者单位

展开▼
年(卷),期 2013(6),Suppl 1
年度 2013
页码 17–27
总页数 11
原文格式 PDF
正文语种
中图分类生物学;
关键词
natural language processing distributional semantics concept extraction named entity recognition empirical lexical resources;

机译：自然语言处理;分布语义;概念提取;命名实体识别;经验词汇资源;

相似文献

外文文献
中文文献
专利

1. Using Empirically Constructed Lexical Resources for Named Entity Recognition: [J] . Siddhartha Jonnalagadda, Trevor Cohen, Stephen Wu, Biomedical Informatics Insights . 2013,第1期

机译：使用经验构造的词汇资源进行命名实体识别：
2. Leveraging Lexical Features for Chinese Named Entity Recognition via Static and Dynamic Weighting [J] . Dong Zhang, Chengying Chi, Xuegang Zhan IAENG Internaitonal journal of computer science . 2021,第1Pta2期

机译：利用静态和动态加权的中文命名实体识别的词汇特征
3. ME-Based Biomedical Named Entity Recognition Using Lexical Knowledge [J] . KYUNG-MI PARK, SEON-HO KIM, HAE-CHANG RIM, ACM transactions on Asian language information processing . 2006,第1期

机译：使用词法知识的基于ME的生物医学命名实体识别
4. Evaluating the Use of Empirically Constructed Lexical Resources for Named Entity Recognition [C] . Siddhartha Jonnalagadda, Trevor Cohen, Stephen Wu, Workshop on computational semantics in clinical text . 2013

机译：评估使用经验构造的词汇资源进行命名实体识别的使用
5. Low-resource Named Entity Recognition [D] . Mayhew, Stephen. 2019

机译：低资源名为实体识别
6. Constructing a Chinese electronic medical record corpus for named entity recognition on resident admit notes [O] . Yan Gao, Lei Gu, Yefeng Wang, 2019

机译：构建中国电子病历语料库以在居民录取通知书上命名实体
7. Chinese named entity recognition using lexicalized HMMs [O] . Fu G, Luke KK 2005

机译：使用词汇化Hmm的中文命名实体识别

Using Empirically Constructed Lexical Resources for Named Entity Recognition

摘要

著录项

相似文献

相关主题

期刊订阅