Frontline Lab summary and source
The repost says OpenMOSS released MOSS-VL, an 11B open vision-language model for real-time video understanding that “perceives while speaking.”
This brief preserves the original source so the summary and editorial context can be checked independently.
Source attributionX · @_akhaliq
Open the original source