首页> 美国卫生研究院文献>BMC Bioinformatics >Bias in random forest variable importance measures: Illustrations sources and a solution

【2h】

Bias in random forest variable importance measures: Illustrations sources and a solution

机译：森林随机变量重要性衡量中的偏见：插图来源和解决方案

代理获取

本网站仅为用户提供外文OA文献查询和代理获取服务，本网站没有原文。下单后我们将采用程序或人工为您竭诚获取高质量的原文，但由于OA文献来源多样且变更频繁，仍可能出现获取不到、文献不完整或与标题不符等情况，如果获取不到我们将提供退款服务。请知悉。

页面导航

摘要
著录项
相似文献
相关主题

摘要

BackgroundVariable importance measures for random forests have been receiving increased attention as a means of variable selection in many classification tasks in bioinformatics and related scientific fields, for instance to select a subset of genetic markers relevant for the prediction of a certain disease. We show that random forest variable importance measures are a sensible means for variable selection in many applications, but are not reliable in situations where potential predictor variables vary in their scale of measurement or their number of categories. This is particularly important in genomics and computational biology, where predictors often include variables of different types, for example when predictors include both sequence data and continuous variables such as folding energy, or when amino acid sequence data show different numbers of categories.

机译：背景技术在生物信息学和相关科学领域的许多分类任务中，作为随机选择变量的一种手段，可变森林的各种重要措施已受到越来越多的关注，例如，选择与某种疾病的预测相关的遗传标记的子集。我们表明，随机森林变量重要性度量是在许多应用中进行变量选择的明智方法，但在潜在预测变量的测量规模或类别数量变化的情况下并不可靠。这在基因组学和计算生物学中尤其重要，其中预测变量通常包括不同类型的变量，例如，当预测变量既包含序列数据又包含连续变量（例如折叠能量）时，或者当氨基酸序列数据显示不同类别的数量时。

著录项

期刊名称 BMC Bioinformatics
作者
Carolin Strobl; Anne-Laure Boulesteix; Achim Zeileis; Torsten Hothorn;
展开▼
作者单位

展开▼
年(卷),期 2007(8),-1
年度 2007
页码 25
总页数 21
原文格式 PDF
正文语种
中图分类应用微生物学;生化遗传学;生化药理学;
关键词

相似文献

外文文献
中文文献
专利

1. Bias in random forest variable importance measures: Illustrations, sources and a solution [J] . Carolin Strobl, Anne-Laure Boulesteix, Achim Zeileis, BMC Bioinformatics . 2007,第1期

机译：森林随机变量重要性衡量中的偏见：插图，来源和解决方案
2. Bias in random forest variable importance measures: Illustrations, sources and a solution [J] . Carolin Strobl, Anne-Laure Boulesteix, Achim Zeileis, BMC Bioinformatics . 2007,第1期

机译：森林随机变量重要性衡量中的偏见：插图，来源和解决方案
3. Bias in the intervention in prediction measure in random forests: illustrations and recommendations [J] . Nembrini Stefano Bioinformatics . 2019,第13期

机译：随机林中预测措施的干预中的偏见：插图和建议
4. Random Forests with Latent Variables to Foster Feature Selection in the Context of Highly Correlated Variables. Illustration with a Bioinformatics Application [C] . Christine Sinoquet, Kamel Mekhnacha International Symposium on Intelligent Data Analysis . 2018

机译：随机森林具有潜在变量，以促进高度相关变量的上下文中的特征选择。与生物信息学应用的插图
5. Using the C-index to measure prediction accuracy and variable importance of random forests, with application to tissue microarray data. [D] . Huang, Yunda. 2004

机译：使用C指数测量随机森林的预测准确性和可变重要性，并将其应用于组织微阵列数据。
6. Intervention in prediction measure: a new approach to assessing variable importance for random forests [O] . Irene Epifanio 2017

机译：干预预测措施：评估随机森林变量重要性的新方法
7. Bias in random forest variable importance measures: Illustrations, sources and a solution [O] . Hothorn Torsten, Zeileis Achim, Boulesteix Anne-Laure, 2007

机译：森林随机变量重要性衡量中的偏见：插图，来源和解决方案

Bias in random forest variable importance measures: Illustrations sources and a solution

摘要

著录项

相似文献

相关主题

期刊订阅