Your article 'PB-LLM: PARTALLY BINARIZE LARGE LANGUAGE MODELS' has sparked my great interest in the methods of model compression. I used PB-LLM code on GitHub to quantify LLaMA, but the size of the model did not change. If it is convenient, I would like to ask why the quantified model size is the same as the original model?
Your article 'PB-LLM: PARTALLY BINARIZE LARGE LANGUAGE MODELS' has sparked my great interest in the methods of model compression. I used PB-LLM code on GitHub to quantify LLaMA, but the size of the model did not change. If it is convenient, I would like to ask why the quantified model size is the same as the original model?