thu-coai
Pinned Loading
Repositories
- EmbodiedAct Public
- Setox Public
- Syncred-Bench Public
SYNCRED-BENCH: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation
- lasa-multilingual-safety Public
【ACL 2026】LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
- AISafetyLab Public
AISafetyLab: A comprehensive framework covering safety attack, defense, evaluation and paper list.
- CROPI Public
[ACL'26] Official Repository for for paper "Data-Efficient RLVR via Off-Policy Influence Guidance"
- IF-RewardBench Public
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation (ACL 2026)
- Survive-at-All-Costs Public
Survive at All Costs: Exploring LLM's Risky Behaviors under Survival Pressure
- IF-CRITIC Public
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation (ACL 2026)
- MIR-SafetyBench Public
MIR-SafetyBench: Evaluating Multi-image Reasoning Safety of Multimodal Large Language Models
Top languages
Loading…
Most used topics
Loading…