Skip to content

Inquire about model compression issues #11

Description

@zhouyijava

Your article 'PB-LLM: PARTALLY BINARIZE LARGE LANGUAGE MODELS' has sparked my great interest in the methods of model compression. I used PB-LLM code on GitHub to quantify LLaMA, but the size of the model did not change. If it is convenient, I would like to ask why the quantified model size is the same as the original model?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions