Historical Python and MATLAB workflow for collecting Douban book information from doulists and tags.
Copyright (C) 2020 Jing Wang
Run the stages from the repository root:
- Review Notes (Douban note IDs used to discover doulists) and Tags_raw (raw tag input).
- Run
python main.pyto prepare the source lists and collect book information. This stage generates inputs such asDoulists_IDandTags_uniqueand writes the collected results. - Open MATLAB in the repository directory and run
mainto import, sort, and organize the collected results using main.m.
Published results: Douban-books-2020.
These historical scripts document the original workflow. Compatibility with current Douban pages has not been verified; the newer implementation is linked below.
- Python 3.7.
- MATLAB R2018a.
These versions document the original environment, rather than a current compatibility guarantee.
- Douban-books-2020: Historical 2020 results distributed as UTF-8 BOM CSV files.
- douban-books-ranking: Newer standalone Python implementation with an online ranking and published JSON data.
See LICENSE for the GNU General Public License v3.0.
Jing Wang
wangjing@xynu.edu.cn
yuzhounh@163.com
2020-7-5 18:25:16