Skip to content
AI-grafen
FAI engineeringScientific method· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Negative results and publication bias

Be able to value negative results and understand why published figures are skewed.

Prerequisites

Intuition

A hundred groups test the same idea. Five get positive results (some of them by chance). Those five publish. The other 95 give up.

The literature now shows that the idea works. That is publication bias, and it makes published effect sizes systematically overestimated.

In ML it is reinforced by two more things:

  • Benchmark chasing: only results that beat SOTA are published, no matter how many variants were tried.
  • Reproductions are not rewarded: showing that something does not hold rarely gets published.

The consequence is that a published result should be regarded as a hypothesis until it has been independently replicated.

Formal

A useful negative result is different from «it did not work». It has to answer three questions:

  1. Was it implemented correctly? Sanity checks: do you reproduce the original's result on the original's task? If not, your negative result is about your implementation, not about the method.
  2. Did the experiment have enough power? With three seeds and σ = 1 % you cannot detect an effect of 0.5 %. «No effect» and «too weak an experiment» are different conclusions.
  3. What was the scope? The method did not work on your data, at your scale, with your hyperparameters. That is often an entirely legitimate and valuable result — but it has to be said.

A formulation that holds:

«We could not reproduce the reported gain from X in our domain (Swedish legal text, 40 k examples, 5 seeds). The difference from the baseline was +0.2 pp [−0.9; 1.3]. We did reproduce the original's result on their dataset (+3.1 pp), which suggests that the implementation is correct and that the effect is domain specific.»

That sentence is more useful than most positive results, because it saves time for everybody who was considering the same thing.

Internally it matters even more: a team that documents only what worked will try the same failed idea again in a year. The journal (the node experimentlogg) is where negative results pay off.

Mastery means

  • Values negative results correctly
  • Explains publication bias and its consequences
  • Reports their own negative results usefully

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences