Deep learning system for automated polyp detection and tracking in colonoscopy images and videos, with a focus on minimizing missed detections over maximizing precision.
This project tackles polyp detection in colonoscopy images using YOLOv8m-seg with an instance segmentation approach. The core challenge: cross-dataset generalization — a model that performs well on its training distribution but fails on data from different hospitals or camera systems is not clinically useful.
Approach: A systematic engineering process — baseline training, error analysis, and evidence-based improvements — guided by published literature (YOLO-LAN, 2025).
| Metric | Optimal-Threshold Baseline (conf=0.30) | Final Model | Change |
|---|---|---|---|
| Kvasir-SEG Recall | 86.7% | 89.8% | +3.1% |
| CVC-ClinicDB Recall | 75.2% | 81.4% | +6.2% |
| Inference Speed | 45 FPS | 45 FPS | — |
Cross-dataset generalization: The final model achieves 81.4% recall on CVC-ClinicDB, a dataset never seen during training (+6.2% over baseline). "Baseline" here refers to the F1-optimal operating point (conf=0.30). At the true default threshold (conf=0.50), CVC-ClinicDB recall is 71.1% — see the Model Comparison table below for the full breakdown.
| Dataset | Images | Role | Source |
|---|---|---|---|
| Kvasir-SEG | 1,000 | Train / Val / Test | simula.no |
| CVC-ClinicDB | 612 | Cross-dataset evaluation only | Kaggle |
| Kvasir (normal) | 140 sampled | Negative samples for training | simula.no |
Data split (Kvasir-SEG): 70% train / 15% val / 15% test (stratified by polyp size)
Challenge: Kvasir-SEG and CVC-ClinicDB come from different hospitals, cameras, and patient populations. A model that only generalizes within one dataset is not suitable for real-world deployment.
- Architecture: YOLOv8m-seg (Ultralytics)
- Pretrained weights: COCO
- Input size: 640×640
- Task: Instance segmentation (bounding box + pixel-level mask per polyp)
- Optimizer: AdamW (lr=0.001)
- Epochs: 100 (early stopping, patience=15)
- Batch size: 8
- Mixed precision: FP16 (AMP)
- Augmentation: flipud=0.5, degrees=45, hsv_s=0.9, hsv_v=0.5 (stronger than default)
- Negative samples: 140 polyp-free frames added to training split (20% ratio)
- Operational threshold: 0.30 (selected via sweep in notebook 04)
- At the true default threshold (conf=0.50), CVC-ClinicDB recall was only 71.1%. Lowering to 0.30 recovered missed detections, raising recall to 75.2% — with no further gain below 0.30 (recall plateaus), and without meaningful precision loss on in-distribution data.
| Approach | Kvasir Recall | CVC Recall | Notes |
|---|---|---|---|
| Baseline (conf=0.50) | 0.867 | 0.711 | Standard YOLOv8m-seg (Default threshold) |
| Threshold Tuning (conf=0.30) | 0.867 | 0.752 | Same model, optimized threshold |
| Higher Resolution (imgsz=960) | 0.835 | 0.751 | Retrain at 960px — did not improve |
| Aug + Negatives (final) | 0.898 | 0.814 | Stronger augmentation + 140 negative frames |
Recall Comparison Across All Approaches:
Cross-Dataset Recall Improvement Journey:
Why Polyps Were Missed — Size Analysis:
This project followed a systematic diagnose → hypothesize → test → verify cycle:
Error analysis on CVC-ClinicDB false negatives revealed that missed polyps were 59% smaller (median area) than detected ones — a small-object detection problem, not a brightness or position problem.
Detected polyps — median area: 9.08% of image
Missed polyps — median area: 3.74% of image
Retrained at imgsz=960 with batch=4. Result: no improvement (miss rate increased from 11.3% to 14.1%). Conclusion: reducing batch size to fit the larger resolution hurt training stability more than the resolution gain helped.
Following YOLO-LAN (2025), which reported +15.2% mAP50:95 on the same dataset using stronger augmentation and negative samples:
- Added
flipud=0.5,degrees=45,hsv_s=0.9to augmentation - Added 140 polyp-free frames from Kvasir (normal categories) to training split
Result: CVC recall improved from 75.2% (optimal-threshold baseline) to 81.4% (+6.2 percentage points)
The final model supports real-time polyp tracking across video frames using ByteTrack, assigning consistent IDs to each polyp throughout the colonoscopy sequence.
Tracking statistics on 150-frame CVC-ClinicDB sequence:
- Detection rate: 84% of frames
- Inference speed: 45 FPS (RTX 4060)
Live Demo: Streamlit Cloud
- Upload colonoscopy image → instant polyp detection with segmentation mask
- Upload video clip → per-frame detection + ByteTrack polyp tracking
- Adjustable confidence threshold (default: 0.30)
- Model performance comparison table with visualizations
# Clone repository
git clone https://github.com/foroughm423/polyp-detection-safety-first.git
cd polyp-detection-safety-first
# Install dependencies
pip install -r requirements.txt
# Run app
streamlit run app/07_streamlit_app.pyNote: The trained model weights are hosted privately on Hugging Face Hub. The live demo above uses secure token authentication via Streamlit Secrets — no setup needed to use the demo. To run locally with your own weights, place a
best.ptfile atmodels/aug_neg/weights/best.pt(trained using the notebooks in this repo) and updateMODEL_PATHin07_streamlit_app.pyaccordingly.
polyp-detection-safety-first/
├── README.md
├── requirements.txt
├── .gitignore
│
├── app/
│ └── 07_streamlit_app.py # Streamlit web application
│
├── notebooks/
│ ├── 01_data_exploration.ipynb
│ ├── 02_data_preparation.ipynb
│ ├── 03_baseline_training.ipynb
│ ├── 04_safety_first_training.ipynb
│ ├── 04b_error_analysis.ipynb
│ ├── 04c_imgsz_retrain.ipynb
│ ├── 04d_augmentation_negatives.ipynb
│ ├── 05_final_evaluation.ipynb
│ └── 06_video_tracking.ipynb
│
├── configs/
│ ├── dataset.yaml
│ ├── dataset_aug_neg.yaml
│ └── dataset_cvc_eval.yaml
│
└── results/
├── figures/
│ ├── final_comparison.png
│ ├── recall_progression.png
│ ├── error_analysis_summary.png
│ ├── error_analysis_missed_samples.png
│ └── tracking_detections.png
└── metrics/
├── baseline_metrics.json
├── safety_first_metrics.json
├── aug_neg_metrics.json
└── final_summary.json
- Python 3.10+
- PyTorch 2.6 + CUDA 12.4
- YOLOv8m-seg (Ultralytics)
- ByteTrack (multi-object tracking)
- OpenCV (video processing)
- Streamlit (web deployment)
- Hugging Face Hub (secure model hosting)
- Test on additional colonoscopy datasets (SUN-SEG, ETIS-Larib)
- Experiment with YOLOv9 or YOLOv10 architectures
- Add Grad-CAM visualizations for model interpretability
- Fine-tune on combined Kvasir-SEG + CVC-ClinicDB for better generalization
- Deploy mobile-friendly version
- Kvasir-SEG Dataset — Jha et al., 2020
- CVC-ClinicDB Dataset — Bernal et al., 2015
- YOLOv8 — Ultralytics, 2023
- ByteTrack — Zhang et al., 2022
- YOLO-LAN — augmentation + negative samples methodology, 2025
- Live Demo: Streamlit Cloud
- Source Code: GitHub
- Training Notebooks (with outputs):
This project is licensed under the MIT License. See LICENSE for details.
Forough Ghayyem 📫 GitHub | LinkedIn | Kaggle
- Kvasir-SEG dataset: SimulaMet, Oslo
- CVC-ClinicDB dataset: Computer Vision Center, Barcelona
- Model hosting: Hugging Face Hub
⚠️ Medical Disclaimer: This model is for research and educational purposes only. It should not be used as a substitute for professional medical diagnosis. Always consult a qualified gastroenterologist for proper evaluation of colonoscopy findings.
---