What does your eval number actually support?

A retrieval benchmark table usually reports a mean and nothing else. From the mean and the query count alone you can recover the interval it supports — and, for a claimed improvement, whether any outcome of that comparison could have reached significance. Nothing is sent anywhere; the arithmetic runs in this page.

how many the figure was averaged over
names starting recall are treated as 0/1
between 0 and 1
the figure the lift is claimed over
0/1 = one relevant doc per query; graded = MRR, nDCG