Hello, thank you for the great paper.
I followed the training procedure as described in the README and used the trained checkpoint to run the evaluation by modifying the paths in source ./evaluation/eval_gen_composition.sh.
However, as shown in the attached image, there is a noticeable performance drop.

To reproduce the results reported in the paper more accurately, should I rerun the evaluation using eval_mode=final instead of fast from the beginning?
Thank you!
Hello, thank you for the great paper.
I followed the training procedure as described in the README and used the trained checkpoint to run the evaluation by modifying the paths in source ./evaluation/eval_gen_composition.sh.
However, as shown in the attached image, there is a noticeable performance drop.

To reproduce the results reported in the paper more accurately, should I rerun the evaluation using eval_mode=final instead of fast from the beginning?
Thank you!