PerceptionDLM
arxiv.org原文 ↗
PerceptionDLM 用 multimodal diffusion language model 做并行区域感知,目标是在多个图像区域上同时生成描述。作者引入 efficient prompting 与 structured attention masking,并构建 ParaDLC-Bench,把 DLC-Bench 扩展到每图多个 region masks,以同时评估 caption 质量和推理效率。它的关键贡献是让 region captioning 不再按区域顺序串行处理,而是利用 DLM 的并行解码特性。
–浏览
评论 · Comments