Abstract
Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we present ChartGaze, a new eye-tracking dataset that captures human gaze patterns during chart reasoning tasks. Through a systematic comparison of human and model attention, we find that LVLMs often diverge from human gaze, leading to reduced interpretability and accuracy. To address this, we propose a gaze-guided attention refinement that aligns image-text attention with human fixations. Our approach improves both answer accuracy and attention alignment, yielding gains of up to 2.56 percentage points across multiple models. These results demonstrate the promise of incorporating human gaze to enhance both the reasoning quality and interpretability of chart-focused LVLMs.
Abstract (translated)
图表是用于传达和表示信息的重要视觉媒介。虽然大型视觉-语言模型(LVLM)在处理图表问答(CQA)任务上已经取得了一些进展,但当这些模型关注图表中的无关区域时,该任务仍然具有挑战性。在这项工作中,我们提出了ChartGaze,这是一个新的眼动追踪数据集,用于捕捉人们在执行图表推理任务时的眼球运动模式。通过系统地比较人类和模型的注意力焦点,我们发现LVLM通常与人类的目光偏离,导致可解释性和准确性降低。为了应对这一问题,我们提出了一种基于目光引导的注意力优化方法,该方法能够将图像-文本注意力与人类的注视点对齐。我们的方法在提高答案准确率的同时也改善了注意机制的一致性,在多种模型上获得了高达2.56个百分点的性能提升。这些结果展示了将人类的目光纳入图表中心的LVLM以增强其推理质量和可解释性的潜力。
URL
https://arxiv.org/abs/2509.13282