← 论文 30

探究image editing模型中的visual planning能力

scored
↗ 原文 ↗ PDF · Hugging Face Daily
📋 摘要 ⭐ 聚焦图像编辑模型的视觉规划与抽象谜题基准,与SE for AI、公平/形式化测试方向擦边相关。 Probing Visual Planning in Image Editing Models
中文
本文探讨image editing模型中的visual planning能力,针对当前主流verbal-centric方法难以处理复杂空间推理、以及fully visual方法因step-by-step planning-by-generation导致计算低效的问题,提出EAR(editing-as-reasoning)范式,将visual planning重构为单步图像变换任务。为剥离视觉识别因素、聚焦内在推理能力,作者构建了程序化生成的抽象puzzle数据集AMAZE,涵盖经典的Maze与Queen两类互补的visual planning问题,其抽象特性便于对autoregressive与diffusion-based模型在pixel-wise fidelity与逻辑正确性两方面进行自动评估。作者评测了主流闭源与开源editing模型,结果显示:zero-shot设置下各模型均表现不佳;在基础规模上fine-tune后,模型能显著泛化到更大的in-domain规模以及out-of-domain的尺度与几何形态。然而,即便在高端硬件上运行的最佳模型,其推理效率仍不及人类求解者的zero-shot水平,揭示了神经网络在visual reasoning方面与人类之间持续存在的差距。与已有以语言为中心或逐步生成式的visual planning研究不同,本文以单步编辑作为推理载体,并通过抽象puzzle设计将推理与识别解耦评估。
English abstract
Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherently visual problem is often tackled through a verbal-centric lens. While recent research demonstrates the promise of fully visual approaches, they suffer from significant computational inefficiency due to the step-by-step planning-by-generation paradigm. In this work, we present EAR, an editing-as-reasoning paradigm that reformulates visual planning as a single-step image transformation. To isolate intrinsic reasoning from visual recognition, we employ abstract puzzles as probing tasks and introduce AMAZE, a procedurally generated dataset that features the classical Maze and Queen problems, covering distinct, complementary forms of visual planning. The abstract nature of AMAZE also facilitates automatic evaluation of autoregressive and diffusion-based models in terms of both pixel-wise fidelity and logical validity. We assess leading proprietary and open-source editing models. The results show that they all struggle in the zero-shot setting, finetuning on basic scales enables remarkable generalization to larger in-domain scales and out-of-domain scales and geometries. However, our best model that runs on high-end hardware fails to match the zero-shot efficiency of human solvers, highlighting a persistent gap in neural visual reasoning.
加载中…
点文件 → 加为 tab;按 Esc 关闭
Esc
输入名称、URL、路径或标签...
选择 Enter 打开 Enter 新标签