首页> 美国卫生研究院文献>AMIA Annual Symposium Proceedings >Automated Disambiguation of Acronyms and Abbreviations in Clinical Texts: Window and Training Size Considerations
【2h】

Automated Disambiguation of Acronyms and Abbreviations in Clinical Texts: Window and Training Size Considerations

机译:临床文本中首字母缩写词和缩写的自动消歧:窗口和培训规模的考虑

代理获取
本网站仅为用户提供外文OA文献查询和代理获取服务,本网站没有原文。下单后我们将采用程序或人工为您竭诚获取高质量的原文,但由于OA文献来源多样且变更频繁,仍可能出现获取不到、文献不完整或与标题不符等情况,如果获取不到我们将提供退款服务。请知悉。

摘要

Acronyms and abbreviations within electronic clinical texts are widespread and often associated with multiple senses. Automated acronym sense disambiguation (WSD), a task of assigning the context-appropriate sense to ambiguous clinical acronyms and abbreviations, represents an active problem for medical natural language processing (NLP) systems. In this paper, fifty clinical acronyms and abbreviations with 500 samples each were studied using supervised machine-learning techniques (Support Vector Machines (SVM), Naïve Bayes (NB), and Decision Trees (DT)) to optimize the window size and orientation and determine the minimum training sample size needed for optimal performance. Our analysis of window size and orientation showed best performance using a larger left-sided and smaller right-sided window. To achieve an accuracy of over 90%, the minimum required training sample size was approximately 125 samples for SVM classifiers with inverted cross-validation. These findings support future work in clinical acronym and abbreviation WSD and require validation with other clinical texts.
机译:电子临床文本中的首字母缩写词和缩写很普遍,并且经常与多种含义相关联。自动首字母缩写词义消除(WSD)是将上下文合适的意义分配给歧义的临床首字母缩写词和缩写的任务,代表了医学自然语言处理(NLP)系统的一个活跃问题。在本文中,使用有监督的机器学习技术(支持向量机(SVM),朴素贝叶斯(NB)和决策树(DT))研究了五十个临床首字母缩写词和500个样本的缩写,以优化窗口的大小和方向,以及确定获得最佳性能所需的最小训练样本量。我们对窗口大小和方向的分析表明,使用较大的左侧窗口和较小的右侧窗口时,效果最佳。为了达到90%以上的精度,具有反向交叉验证的SVM分类器所需的最小训练样本量约为125个样本。这些发现支持将来在临床首字母缩写词和缩写WSD中的工作,并且需要与其他临床文献一起进行验证。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
代理获取

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号