Skip to content

Question about the training and evaluation protocol for Table 1 and Table 2 #5

Description

@yywd6

Hi authors,

Thank you for sharing this interesting work and the code.

I have a question about the reported Real3D-AD results in Table 1 and Table 2.

In the paper, the main text says that Table 1 follows a one-class training protocol, where each model is trained on annotated data from a single category and evaluated on the remaining categories with averaged performance.

Table 2 also reports cross-category P-AUROC results, where each row corresponds to training on one category and testing on the other categories. I recalculated the row-wise means in Table 2, and they match the reported Mean column. However, when averaging all cross-category P-AUROC values in Table 2, I obtain about 81.39, while the BTP point-level mean reported in Table 1 is 84.5.

I also checked that the Table 1 BTP point-level mean 84.5 is the average of the 12 per-category values in Table 1.

Could you please clarify how the Table 1 BTP results were obtained?

Specifically, I would like to know:

  • For each category result in Table 1, which source category was used for training?
  • Is Table 1 averaged over multiple one-class training settings, or does it use a selected/best source category for each test category?
  • Are Table 1 and Table 2 produced under the same checkpoints and evaluation settings?

This clarification would be very helpful for reproducing the results correctly.

Thank you very much.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions