Stars
Offical repository for NeurIPS 2025 paper "From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring".
An interface library for RL post training with environments.
CLI security scanner built for the agentic era. Detects CI/CD misconfigs, agent permission risks, MCP tool injection, hardcoded secrets, and DMCA-flagged AI dependencies.
“ToolHazard: Scalable Red Teaming of LLM Agents via Adversarial Tool-Interactive Environment and Task Synthesis”
AI Security Scanner - Test your AI systems for prompt injection and extraction vulnerabilities
A debugging framework for agentic AI systems: diagnose failures, attribute root causes, recover with evidence, and validate fixes through reruns.
SafeLine is a self-hosted WAF(Web Application Firewall) / reverse proxy to protect your web apps from attacks and exploits.
Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
Echo Agent 是一个可自托管、长期运行、持续学习的 AI Agent,面向个人与团队的私有自动化场景。它可以部署在自有服务器上,统一连接模型、工具、记忆、权限与消息入口。内置四层认知记忆、遗忘曲线与矛盾检测机制,能够在跨会话任务中持续沉淀上下文,并保持长期记忆的质量。针对命令执行、文件操作等高风险行为,它提供基于 LLM 的审批与解释机制,为关键操作建立可审计、可追溯的安全边界。原生…
The official implementation of the work "TopicAttack: An Indirect Prompt Injection Attack via Topic Transition"
POC管理平台已集成公开POC仓库内容,可自行添加POC到平台实现集中管理,快速搜索,验证
Code for paper "ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents"
FinanceHarness: Autonomous Financial Deep Research Framework
The Official Repository for Paper "HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?"
Installable AI coding workflow for risk-based routing, scoped sub-agents, and evidence-backed completion.
Silero VAD: pre-trained enterprise-grade Voice Activity Detector