📋 摘要
⭐ 时尚图像CNN分类与可解释性,与SE/AI测试、公平性测试等核心兴趣几乎无交集。
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing
中文
本文针对 fashion AI 系统在编码特定时装屋、编辑及历史时期审美逻辑时缺乏可解释性的问题,提出 FASH-iCNN——一个多模态 CNN probing 系统。作者基于 1991-2024 年间 15 个时装屋的 87,547 张 Vogue 走秀图像进行训练,使系统能够在给定服装照片时,识别其所属时装屋、所处年代及其反映的色彩传统。实验表明,仅基于服装的模型在 14 个时装屋上的 house identity top-1 准确率达到 78.2%,decade 识别 top-1 达 88.6%,跨 34 年的 specific year 识别 top-1 达 58.3%,平均误差仅 2.2 年。通过对视觉通道的 probing,作者揭示了显著的信号分离现象:移除颜色仅造成 house identity 准确率下降 10.6pp,而移除纹理则导致 37.6pp 的下降,表明 texture 与 luminance 是编辑性身份的主要载体。与既有工作不同,FASH-iCNN 将编辑文化视为信号本身而非背景噪声,使用户能够审视预测背后所编码的时装屋、编辑及历史时刻,从而实现 editorial fashion identity 的可检视性。
English abstract
Fashion AI systems routinely encode the aesthetic logic of specific houses, editors, and historical moments without disclosing it. We present FASH-iCNN, a multimodal system trained on 87,547 Vogue runway images across 15 fashion houses spanning 1991-2024 that makes this cultural logic inspectable. Given a photograph of a garment, the system recovers which house produced it, which era it belongs to, and which color tradition it reflects. A clothing-only model identifies the fashion house at 78.2% top-1 across 14 houses, the decade at 88.6% top-1, and the specific year at 58.3% top-1 across 34 years with a mean error of just 2.2 years. Probing which visual channels carry this signal reveals a sharp dissociation: removing color costs only 10.6pp of house identity accuracy, while removing texture costs 37.6pp, establishing texture and luminance as the primary carriers of editorial identity. FASH-iCNN treats editorial culture as the signal rather than background noise, identifying which houses, eras, and color traditions shaped each output so that users can see not just what the system predicts but which houses, editors, and historical moments are encoded in that prediction.
加载中…
点文件 → 加为 tab;按 Esc 关闭