Nature Medicine published lessons from deploying ChatEHR, a large language model system inside Stanford Medicine. Authors including Nigam Shah argue that benchmark-style scores are a poor monitor once clinicians drive open-ended chats, and that hospitals need new methods to watch live performance, including unsupported claims spotted during deployment. Stanford University owns ChatEHR; the authors are Stanford employees. The piece reframes evaluation as an operations problem for medical centers, not a leaderboard race. Caveat: the article is a deployment essay rather than a randomized outcome trial, so it guides monitoring practice more than it proves clinical benefit.