ByteDance Seed launched SeedRealtime, a native audio-video full-duplex model that fuses sound, vision, and text in one end-to-end stack for continuous real-time chat. Unlike cascaded ASR-plus-vision pipelines, it keeps perceiving while it replies, aiming to cut turn-taking glitches such as interrupting mid-sentence, lagging after a pause, or reacting to background chatter. Lab evaluations say rhythm problems fell by about half versus cascaded baselines. The model is live in Doubao video calls. Why it matters: China shipped multimodal full-duplex beyond speech-only demos. Caveat: claims rest on ByteDance tests, and noisy multi-party scenes remain hard.