Stars
cursor-byok is a local implementation of Cursor's backend. https://github.com/leookun/cursor-byok/releases
This repository contains a regularly updated paper list for LLMs-reasoning-in-latent-space.
[CVPR 2026 Highlight] Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
One Discrete Word for Visual Reasoning Overtakes Agentic and Latent Methods
The official implementation of "CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization"
[ICML2026 Spotlight] UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
Elevate your AI research writing, no more tedious polishing ✨
NanaDraw turns complex scientific ideas into clear, expressive visuals you can use right away. Powered by Nano Banana, it generates editable illustrations in formats like SVG, PPT, and XML.
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
【Zotero AI 管家】调用大模型,自动精读论文库里的论文,总结为Zotero笔记。支持主流大模型平台!您只需像往常一样把文献丢进 Zotero, 管家会自动帮您精读论文,将文章揉碎了总结为笔记,让您“十分钟完全了解”这篇论文!
Official Implementation of "Imagination Helps Visual Reasoning, But Not Yet in Latent Space"
[ACL'26 Oral] Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
Official codebase for the paper Latent Visual Reasoning
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
[CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"
[CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
强化学习中文教程(蘑菇书🍄),在线阅读地址:https://datawhalechina.github.io/easy-rl/
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)