A Communications Medicine study introduces CLAD, scoring fifteen large language models on 198 U.S. malpractice cases across 3,072 simulated consults. Newer systems look more legally defensible: GPT-5.2 averaged 0.71 versus 0.34 for GPT-4o. That gain tracks procedure volume and cost (Spearman rho 0.95): GPT-5.2 recommended about 9.3 procedures at roughly $1,118 Medicare cost per consult, versus 1.3 procedures and $221 for GPT-4o. Shorter prompts cut length far more than cost, and the pattern replicated on UK cases. Caveat: simulated liability scores are not clinical outcomes, so hospitals still need prospective trials before trusting the tradeoff.