Anthropic and Redwood Research released the Conceptual Reasoning Index, aggregating benchmarks that score how models judge conceptual arguments, stay logically consistent, and handle decision-theoretic puzzles where empirical feedback is scarce. Their August 12 write-up argues those skills matter for AI governance work that cannot be hill-climbed with cheap unit tests. Through August 10, Claude Opus 5 led near 73.6, still below an estimated ceiling around 91. Live scores sit at conceptualreasoning.ai. That matters if labs start optimizing for philosophy-grade reasoning. The caveat is that the index is new and access to the main LMCA dataset still requires a request form.