The 360 AI Research Institute has unveiled its MoSA (Motion-Grounded Segment Anything) model, which aims to enhance AI’s ability to perceive and understand objects in motion without the need for human-generated labels. Accepted at the ECCV 2026 conference, MoSA demonstrates the potential of using over 10,000 hours of unlabelled video to create more than 21 million pseudo-labels, marking a significant advancement in self-supervised learning techniques. This development aligns with the broader shift in AI research towards improving precision, controllability, and usability in visual technologies, moving beyond the traditional reliance on large datasets of hand-labeled images.
Meta: Meta develops large-scale AI models for vision and other domains as part of its technology research efforts. Its Segment Anything Model (SAM) demonstrated object outlining in images but relied on extensive hand-labeled datasets. The company’s prior work provides the baseline that newer approaches like MoSA seek to improve upon by eliminating manual labeling requirements.
MoSA: MoSA, short for Motion-Grounded Segment Anything, is a model designed to enable AI to recognize and segment objects in video by analyzing motion patterns. Accepted at ECCV 2026 under the ‘See Precisely’ theme, it tests self-supervised learning from unlabeled video footage to generate training signals. This contributes to the broader goal of creating more accurate and controllable AI perception systems for real-world applications.
ECCV 2026: ECCV 2026 is the European Conference on Computer Vision, a leading academic venue for research in computer vision and related AI fields. It has accepted the MoSA paper from 360 AI Research Institute as part of its 2026 program. The conference showcases advancements that push AI toward greater precision in visual understanding tasks.
360 AI Research Institute: 360 AI Research Institute conducts research on advancing AI systems for more precise perception, editing, and generation capabilities. Its H1 2026 papers focus on directions like ‘See Precisely’ to improve how AI understands and interacts with visual content. The institute’s work, including the MoSA paper, explores training approaches that reduce reliance on manual data labeling.
Precision Emphasis: AI development is shifting focus from simply increasing model size toward methods that enhance accuracy, controllability, and practical usability in visual tasks.
Research Dissemination: Acceptance at major computer vision conferences like ECCV highlights innovations in label-efficient training techniques for segmentation and object understanding.
Self-Supervised Learning: Motion cues in ordinary video footage can serve as a foundation for training AI perception models without requiring human-provided labels.
