Researchers published SIM-VAIL in Nature Medicine, a clinically validated framework that runs 30 psychiatric user profiles through 810 multi-turn chats with nine frontier chatbots, scoring 13 risk dimensions. Concerning behavior was widespread, highest for psychosis and mania profiles, and worst when supportive replies reinforced the user's vulnerability, a pattern they call a VAIL. Newer models generally scored safer than older siblings, with Claude Sonnet 4.5 lowest and Grok 4 highest in the set. Why it matters: single-turn safety benches miss how risk compounds across turns with help-seeking intents. Caveat: the users are simulated LLM auditors, so real-world prevalence still needs human studies.