...
首页> 外文期刊>IEEE Transactions on Parallel and Distributed Systems >Almost certain fault diagnosis through algorithm-based fault tolerance
【24h】

Almost certain fault diagnosis through algorithm-based fault tolerance

机译:通过基于算法的容错能力几乎可以确定故障

获取原文
获取原文并翻译 | 示例
           

摘要

Algorithm-based fault tolerance has been proposed as a technique to detect incorrect computations in multiprocessor systems. In algorithm-based fault tolerance, processors produce data elements that are checked by concurrent error detection mechanisms. We investigate the efficacy of this approach for diagnosis of processor faults. Because checks are performed on data elements, the problem of location of data errors must first be solved. We propose a probabilistic model for the faults and errors in a multiprocessor system and use it to evaluate the probabilities of correct error location and fault diagnosis. We investigate the number of checks that are necessary to guarantee error location with high probability. We also give specific check assignments that accomplish this goal. We then consider the problem of fault diagnosis when the locations of erroneous data elements are known. Previous work on fault diagnosis required that the data sets produced by different processors be disjoint. We show, for the first time, that fault diagnosis is possible with high probability, even in systems where processors combine to produce individual data elements.
机译:已经提出了基于算法的容错能力,作为检测多处理器系统中错误计算的技术。在基于算法的容错能力中,处理器生成数据元素,这些数据元素由并发错误检测机制检查。我们调查这种方法对处理器故障诊断的功效。由于对数据元素执行检查,因此必须首先解决数据错误的位置问题。我们为多处理器系统中的故障和错误提出了一个概率模型,并使用它来评估正确的错误位置和故障诊断的概率。我们调查必要的检查次数,以确保高概率确定错误位置。我们还提供实现此目标的特定检查任务。然后,当已知错误数据元素的位置时,我们将考虑故障诊断问题。以前的故障诊断工作要求将不同处理器产生的数据集分离。我们首次证明,即使在处理器结合在一起产生单独的数据元素的系统中,故障诊断也是很有可能的。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号