Paper Reading AI Learner

Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models

2026-04-03 16:28:03
Yunfei Bai, Amit Dhanda, Shekhar Jain

Abstract

The recent advancements in Vision Language Models (VLMs) have demonstrated progress toward true intelligence requiring robust reasoning capabilities. Beyond pattern recognition, linguistic reasoning must integrate with visual comprehension, particularly for Chart Question Answering (CQA) tasks involving complex data visualizations. Current VLMs face significant limitations in CQA, including imprecise numerical extraction, difficulty interpreting implicit visual relationships, and inadequate attention mechanisms for capturing spatial relationships in charts. In this work, we address these challenges by presenting Chart-RL, a novel reinforcement learning framework that enhances VLMs chart understanding through feedback-driven policy optimization of visual perception and logical inference. Our key innovation includes a comprehensive framework integrating Reinforcement Learning (RL) from Policy Optimization techniques along with adaptive reward functions, that demonstrates superior performance compared to baseline foundation models and competitive results against larger state-of-the-art architectures. We also integrated Parameter-Efficient Fine-Tuning through Low-Rank Adaptation (LoRA) in the RL framework that only requires single GPU configurations while preserving performance integrity. We conducted extensive benchmarking across open-source, proprietary, and state-of-the-art closed-source models utilizing the ChartQAPro dataset. The RL fine-tuned Qwen3-VL-4B-Instruct model achieved an answer accuracy of 0.634, surpassing the 0.580 accuracy of the Qwen3-VL-8B-Instruct foundation model despite utilizing half the parameter count, while simultaneously reducing inference latency from 31 seconds to 9 seconds.

Abstract (translated)

视觉语言模型(VLMs)的最新进展已展现出向真正智能迈进的态势,而这种智能需要强大的推理能力。超越模式识别,语言推理必须与视觉理解相结合,尤其在涉及复杂数据可视化的图表问答(CQA)任务中。当前VLMs在CQA方面面临显著局限,包括数值提取不精确、难以解读隐含视觉关系,以及对图表中空间关系的注意力机制不足。本研究通过提出Chart-RL——一种新颖的强化学习框架,来应对这些挑战。该框架通过视觉感知与逻辑推理的反馈驱动策略优化,增强了VLMs的图表理解能力。我们的核心创新在于整合了基于策略优化的强化学习(RL)技术与自适应奖励函数的综合框架,相较于基线基础模型展现出更优性能,并与大型先进架构相比也取得了具有竞争力的结果。我们还通过低秩适配(LoRA)在RL框架中集成了参数高效微调,仅需单GPU配置即可保持性能完整性。我们利用ChartQAPro数据集,对开源、专有及最先进的闭源模型进行了广泛基准测试。经RL微调的Qwen3-VL-4B-Instruct模型达到了0.634的答案准确率,超越了参数量减半的Qwen3-VL-8B-Instruct基础模型(0.580准确率),同时将推理延迟从31秒降至9秒。

URL

https://arxiv.org/abs/2604.03157

PDF

https://arxiv.org/pdf/2604.03157.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot