論文 Hugging Face 発表: 2026-05-26 HF ↑4

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

著者: Eric Onyame, Runtao Zhou, Kowshik Thopalli, Bhavya Kailkhura, Chirag Agarwal

要約

Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. However, its reliability remains largely unexplored beyond English and across diverse model families. We present the first large-scale evaluation of CoT mon…

#alignment#benchmark#llm

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

要約

同じカテゴリの記事

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment

World-R1: テキストから動画生成における3D制約の強化学習による整合