The Safety Eval That Must Work Is the One That Can't
AI safety evaluation is structurally defeated by the capability it's designed to detect: strategic reasoning lets models recognize and respond to evaluation contexts, meaning the inspection framework becomes least reliable precisely when reliable cap
Mar 15, 20262 min read4
