-
Shanghai Innovation Institute
Stars
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Context Compression in Long-Horizon LLM-based Agents
[COLM 2025] Official PyTorch implementation of "Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models"
[CVPR 2026] This repository is the official implementation of MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
A Reproduction of GDM's Nested Learning Paper
[ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
Official repo for "GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization"
[ICML 2026]A framework to compare low-bit integer and float-point formats
MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
[ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).
This is the official code repository for the paper: Towards General Continuous Memory for Vision-Language Models.
[ICLR 2026] LightMem: Lightweight and Efficient Memory-Augmented Generation
A high-throughput and memory-efficient inference and serving engine for LLMs
SPEC-RL: Accelerating On-Policy Reinforcement Learning via Speculative Rollouts
The world's first open-source multimodal creative assistant This is a substitute for Canva and Manus that prioritizes privacy and is usable locally.
Sing-box精装桶五合一协议VPS专用脚本:三大独家功能!自签/acme双证书切换、Argo固定临时双隧道(可共存)、Psiphon赛风VPN(30个国家)分流功能、本地IP订阅生成
siiRL: Shanghai Innovation Institute RL Framework for Advanced LLMs and Multi-Agent Systems
[NeurIPS 2025] Efficient Reasoning Vision Language Models
PyTorch Code for Energy-Based Transformers paper -- generalizable reasoning and scalable learning
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
[TMLR 2025] Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enablin…