CaML Benchmarks Collection CaML benchmark datasets: TAC, ANIMA, MORU, WAS + holdout/audit sets. Leaderboard: compassionbench.com • 10 items • Updated 6 days ago
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 23 days ago • 5
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning Paper • 2605.16301 • Published Jun 3 • 1
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 23 days ago • 5
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 23 days ago • 5