ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.
image-captioning segmentation object-detection nodes video-understanding vlm custom-nodes img2text llm mllm llava llama-cpp comfyui moondream grounding-dino gguf joytag florence-2 sam2 qwen3-vl
-
Updated
Aug 8, 2026 - Python