每日 Harness 开源 · Source
返回本期 · Back to 2026-06-23

论文 · Papers2026-06-23 · Tuesday, June 23, 2026

PerceptionDLM

arxiv.org原文 ↗

PerceptionDLM
PerceptionDLM 用 multimodal diffusion language model 做并行区域感知,目标是在多个图像区域上同时生成描述。作者引入 efficient prompting 与 structured attention masking,并构建 ParaDLC-Bench,把 DLC-Bench 扩展到每图多个 region masks,以同时评估 caption 质量和推理效率。它的关键贡献是让 region captioning 不再按区域顺序串行处理,而是利用 DLM 的并行解码特性。
浏览

评论 · Comments