← 论文 15

RADIO-ViPE:面向动态环境 open-vocabulary semantic SLAM 的在线紧耦合多模态融合

scored
↗ 原文 ↗ PDF · Hugging Face Daily
📋 摘要 ⭐ SLAM与视觉语言定位主题,与SE for AI、可信AI、公平性测试等方向几乎无交集。 RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments
中文
本文提出 RADIO-ViPE(Reduce All Domains Into One — Video Pose Engine),一种在线 semantic SLAM 系统,针对动态环境下的几何感知 open-vocabulary grounding 问题,将任意自然语言查询与局部化的 3D 区域和物体相关联。与依赖标定、有位姿 RGB-D 输入的现有方法不同,RADIO-ViPE 直接在原始单目 RGB 视频流上运行,无需相机内参、深度传感器或位姿初始化。方法上,系统将源自 agglomerative foundation models(如 RADIO)的多模态视觉-语言 embeddings 与几何场景信息进行紧耦合,并在初始化、优化以及 factor graph 连接环节中实现该耦合,以提升地图在多模态间的一致性。优化过程嵌入自适应 robust kernels,用以同时处理主动运动物体以及由 agent 移动造成的场景元素位移(例如 ego-centric 会话中被重排的家具)。实验表明,RADIO-ViPE 在动态 TUM-RGBD 基准上取得 state-of-the-art 结果,同时在依赖标定数据与静态场景假设的离线 open-vocabulary 方法面前保持具有竞争力的性能。作者认为该工作弥合了真实部署中的关键缺口,可为自主机器人及无约束 in-the-wild 视频流提供鲁棒的 open-vocabulary 语义 grounding。
English abstract
We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary grounding, associating arbitrary natural language queries with localized 3D regions and objects in dynamic environments. Unlike existing approaches that require calibrated, posed RGB-D input, RADIO-ViPE operates directly on raw monocular RGB video streams, requiring no prior camera intrinsics, depth sensors, or pose initialization. The system tightly couples multi-modal embeddings -- spanning vision and language -- derived from agglomerative foundation models (e.g., RADIO) with geometric scene information. This coupling takes place in initialization, optimization and factor graph connections to improve the consistency of the map from multiple modalities. The optimization is wrapped within adaptive robust kernels, designed to handle both actively moving objects and agent-displaced scene elements (e.g., furniture rearranged during ego-centric session). Experiments demonstrate that RADIO-ViPE achieves state-of-the-art results on the dynamic TUM-RGBD benchmark while maintaining competitive performance against offline open-vocabulary methods that rely on calibrated data and static scene assumptions. RADIO-ViPE bridges a critical gap in real-world deployment, enabling robust open-vocabulary semantic grounding for autonomous robotics and unconstrained in-the-wild video streams. Project page: https://be2rlab.github.io/radio_vipe
加载中…
点文件 → 加为 tab;按 Esc 关闭
Esc
输入名称、URL、路径或标签...
选择 Enter 打开 Enter 新标签