If two translation systems differ differ in performanceon a test set, can we trust that this indicatesa difference in true system quality? To answer thisquestion, we describe bootstrap resampling methodsto compute statistical significance of test results,and validate them on the concrete example of theBLEU score. Even for small test sizes of only 300sentences, our methods may give us assurances thattest result differences are real.
展开▼