Skip to content
View Oscar-dzy's full-sized avatar
  • Shanghai AI Lab
  • Shanghai

Block or report Oscar-dzy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

本人的科研经验

13,684 674 Updated Jun 6, 2026

cursor-byok is a local implementation of Cursor's backend. https://github.com/leookun/cursor-byok/releases

Go 2,250 349 Updated Aug 11, 2026

This repository contains a regularly updated paper list for LLMs-reasoning-in-latent-space.

369 9 Updated Jun 20, 2026

[CVPR 2026 Highlight] Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding

Python 92 5 Updated Apr 9, 2026
Python 9 Updated May 26, 2026

One Discrete Word for Visual Reasoning Overtakes Agentic and Latent Methods

Python 138 Updated Jun 9, 2026

The official implementation of "CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization"

Python 5 Updated May 12, 2026

[ICML2026 Spotlight] UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

Python 165 1 Updated Jul 13, 2026

Elevate your AI research writing, no more tedious polishing ✨

32,850 2,424 Updated May 18, 2026

NanaDraw turns complex scientific ideas into clear, expressive visuals you can use right away. Powered by Nano Banana, it generates editable illustrations in formats like SVG, PPT, and XML.

TypeScript 115 14 Updated Apr 29, 2026

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

Python 7,296 834 Updated Aug 4, 2026

anti-老登,反登味的飞书机器人。拒绝内耗,从我做起,让职场再无登味

Python 29 Updated Apr 14, 2026

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python 77,410 6,521 Updated Aug 11, 2026
Jupyter Notebook 7 1 Updated Jun 18, 2026

这倒是提醒我了

Python 406 57 Updated Jul 26, 2026
Python 49 2 Updated Jun 4, 2026

JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.

Python 2,188 160 Updated Aug 5, 2026

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,775 1,831 Updated Jan 30, 2026

【Zotero AI 管家】调用大模型,自动精读论文库里的论文,总结为Zotero笔记。支持主流大模型平台!您只需像往常一样把文献丢进 Zotero, 管家会自动帮您精读论文,将文章揉碎了总结为笔记,让您“十分钟完全了解”这篇论文!

TypeScript 1,599 83 Updated Aug 10, 2026

Official Implementation of "Imagination Helps Visual Reasoning, But Not Yet in Latent Space"

4 Updated Feb 27, 2026

A paper list of Awesome Latent Space.

956 41 Updated Jul 13, 2026

Glerium's Blog

Markdown 3 Updated Aug 11, 2026

[ACL'26 Oral] Interleaved Latent Visual Reasoning with Selective Perceptual Modeling

Python 66 4 Updated Aug 5, 2026

Official codebase for the paper Latent Visual Reasoning

Python 172 10 Updated Oct 22, 2025

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python 163,916 34,216 Updated Aug 11, 2026

[CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"

Python 216 10 Updated Mar 19, 2026

[CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Python 294 17 Updated Aug 2, 2025

强化学习中文教程(蘑菇书🍄),在线阅读地址:https://datawhalechina.github.io/easy-rl/

Jupyter Notebook 14,540 2,265 Updated Dec 30, 2025

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 74,012 9,057 Updated Aug 10, 2026

MUVR_Eval

Python 2 Updated May 16, 2025
Next