I'm working in data mining and big data domain.
-
University of Information Technology - VNU HCM
- Ho Chi Minh city
-
19:01
(UTC +07:00)
Stars
The Python Code Tutorials
Apache Spark - A unified analytics engine for large-scale data processing
Source-LDA: Enhancing probabilistic topic models using prior knowledge sources (ICDE 2017)
Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.
Ansible playbook that installs a Hadoop cluster, with HBase, Hive, Presto for analytics, and Ganglia, Smokeping, Fluentd, Elasticsearch and Kibana for monitoring and centralized log indexing.