Skip to content
View jlidw's full-sized avatar
👀
👀
  • The Hong Kong University of Science and Technology
  • Hong Kong SAR, China

Block or report jlidw

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Lightweight, open-source AI agent for your tools, chats, and workflows.

Python 46,142 8,157 Updated Jul 24, 2026
Python 418 44 Updated Jul 23, 2026

Synchronize Codex session provider metadata across rollout files and SQLite state.

C# 2,772 123 Updated Jul 23, 2026

Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.

Swift 25,035 2,064 Updated Jul 24, 2026

Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.5, Grok 4.3, Claude model thro…

Go 44,517 6,973 Updated Jul 24, 2026

Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA

TypeScript 123,970 18,567 Updated Jul 15, 2026

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…

Python 13,783 1,236 Updated Jul 22, 2026

LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost…

Python 58,492 50,245 Updated Jul 24, 2026

A live reading list for LLM data synthesis (Updated to July, 2025).

491 39 Updated Apr 9, 2026

Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models

Python 55 7 Updated Sep 19, 2025

The agent that grows with you

Python 219,584 41,682 Updated Jul 24, 2026

The best agent harness.

TypeScript 13,088 732 Updated Jul 23, 2026

A benchmark for LLMs on complicated tasks in the terminal

Python 2,481 560 Updated Jul 11, 2026

KIRA

Python 918 107 Updated May 29, 2026

PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

Python 1,296 147 Updated Jul 2, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 383,961 80,664 Updated Jul 24, 2026

SkillsBench evaluates how well skills work and how effective agents are at using them.

PDDL 1,571 344 Updated Jul 23, 2026

Framework for evaluating and improving agents

Python 3,444 1,372 Updated Jul 23, 2026

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

Python 733 65 Updated May 17, 2026

One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.

TypeScript 36,138 3,014 Updated Jul 23, 2026

Multi-Agent Harness for Production AI

Python 10,764 5,668 Updated May 29, 2026

MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.

Python 451 40 Updated Jun 12, 2026

Salesforce Enterprise Deep Research

Python 1,194 190 Updated Jun 2, 2026

MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.

Python 593 87 Updated Jun 23, 2026

A Survey of Reinforcement Learning for Large Reasoning Models

TeX 2,468 131 Updated Nov 9, 2025

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Python 494 68 Updated Oct 7, 2025

基于多智能体LLM的中文金融交易框架 - TradingAgents中文增强版

Python 30,585 6,442 Updated Jul 24, 2026

The evaluation benchmark on MCP servers

Python 251 16 Updated Sep 3, 2025
Python 2 Updated Nov 3, 2025

Tongyi Deep Research, the Leading Open-source Deep Research Agent

Python 19,711 1,507 Updated Feb 27, 2026
Next