Skip to content
View sotayang's full-sized avatar
❓
Seeking
❓
Seeking

Block or report sotayang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sotayang/README.md

Hi, I'm Zhenyu Yang πŸ‘‹

Zhenyu's GitHub activity statistics

Total Stars

  • πŸŽ“ Fourth-year Ph.D. student (2022–2027) at CASIA
  • πŸ”­ Building online video understanding systems and multimodal agents
  • 🀝 Always happy to discuss research and collaboration

Teaching video AI to watch, remember, reason β€” and respond on time.


More...

πŸ“‘ Current transmission

camera stream ──► temporal memory ──► multimodal reasoning ──► act at the right moment
                         β–²                                      β”‚
                         └──────────── learn from feedback β”€β”€β”€β”€β”€β”€β”˜
  • Streaming perception: understand an unbounded video as it arrives
  • Proactive interaction: decide whether and when a response is needed
  • Efficient memory: preserve useful history without replaying everything
  • Multimodal reasoning: connect visual events, language, and long-term context
πŸ“Š Open the lab dashboard
Zhenyu's contribution activity graph
PythonΒ  PyTorchΒ  C++Β  LinuxΒ  DockerΒ  GitΒ  GitHubΒ  LaTeXΒ  VS Code
🐍 Release the contribution snake
Contribution grid snake animation

Research should not only understand what happened β€” it should be ready for what happens next.

Pinned Loading

  1. Awesome-Streaming-Video-Understanding Awesome-Streaming-Video-Understanding Public

    πŸ”₯πŸ”₯πŸ”₯ [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding β€” Building Always-on, Real-time Video AI πŸ€–

    441 27

  2. LDRE LDRE Public

    [SIGIR'2024 Best Paper Honorable Mention] Official repository for "LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval"

    Python 105 8

  3. LiveStar LiveStar Public

    [NeurIPS'2025] Official repository for "LiveStar: Live Streaming Assistant for Real-World Online Video Understanding"

    Python 156 7

  4. SVBench SVBench Public

    [ICLR'2025 Spotlight] Official repository for "SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding"

    Python 121 2

  5. SEIZE SEIZE Public

    [ACM MM'2024] Official repository for "Semantic Editing Increment Benefits Zero-Shot Composed Image Retrieval"

    Python 43 2

  6. yuxumin/ViQ yuxumin/ViQ Public

    [ECCV2026] ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

    Python 83 5