Measuring Intelligence Beyond Human Scale
A relative-evaluation framework where models generate public challenges for other systems, enabling scalable measurement of capabilities beyond human-authored benchmarks.
A relative-evaluation framework where models generate public challenges for other systems, enabling scalable measurement of capabilities beyond human-authored benchmarks.
A mechanism-design view of alignment in solver-auditor AI pipelines, where reward design induces the behavioral equilibrium of both solving and oversight.
Oracle-based algorithms for regret minimization in very large games, motivated by AI Safety via Debate and language-based action spaces.
A study of efficient regret minimization in games with hidden high-reward structure, motivated by AI alignment and language games.
A statistical perspective on orthogonality in AI systems and its implications for alignment.
An accessible overview of the mechanism-design perspective on AI alignment.