Paper Reading AI Learner

VisDocSketcher: Towards Scalable Visual Documentation with Agentic Systems

2025-09-15 14:02:29
Lu\'is F. Gomes, Xin Zhou, David Lo, Rui Abreu

Abstract

Visual documentation is an effective tool for reducing the cognitive barrier developers face when understanding unfamiliar code, enabling more intuitive comprehension. Compared to textual documentation, it provides a higher-level understanding of the system structure and data flow. Developers usually prefer visual representations over lengthy textual descriptions for large software systems. Visual documentation is both difficult to produce and challenging to evaluate. Manually creating it is time-consuming, and currently, no existing approach can automatically generate high-level visual documentation directly from code. Its evaluation is often subjective, making it difficult to standardize and automate. To address these challenges, this paper presents the first exploration of using agentic LLM systems to automatically generate visual documentation. We introduce VisDocSketcher, the first agent-based approach that combines static analysis with LLM agents to identify key elements in the code and produce corresponding visual representations. We propose a novel evaluation framework, AutoSketchEval, for assessing the quality of generated visual documentation using code-level metrics. The experimental results show that our approach can valid visual documentation for 74.4% of the samples. It shows an improvement of 26.7-39.8% over a simple template-based baseline. Our evaluation framework can reliably distinguish high-quality (code-aligned) visual documentation from low-quality (non-aligned) ones, achieving an AUC exceeding 0.87. Our work lays the foundation for future research on automated visual documentation by introducing practical tools that not only generate valid visual representations but also reliably assess their quality.

Abstract (translated)

视觉文档是降低开发人员理解陌生代码时的认知障碍的有效工具,使得理解更加直观。与文本文档相比,它提供了对系统结构和数据流的高层次理解。对于大型软件系统,开发者通常更喜欢使用图表而非冗长的文字描述。然而,生成视觉文档既困难又具挑战性,手动创建耗时且目前没有方法能够直接从代码自动生成高质量的视觉文档。此外,评估其质量往往主观性强,难以标准化和自动化。为应对这些挑战,本文首次探索了利用代理LLM系统(即具有代理能力的语言模型)自动生成视觉文档的方法。 我们介绍了VisDocSketcher——首个结合静态分析与LLM代理的技术,用于识别代码中的关键元素并生成相应的可视化表示的代理基方法。同时,我们提出了一种新颖的评估框架AutoSketchEval,该框架使用基于代码级别的指标来评定生成的视觉文档质量。实验结果显示,我们的方法能够为74.4%的样本创建有效的视觉文档,并且相较于简单的模板基础线相比,性能提升了26.7-39.8%。 此外,我们的评估框架可以可靠地区分高质量(与代码对齐)和低质量(非对齐)的视觉文档,准确率AUC超过0.87。本研究为未来自动化生成高质量视觉文档的研究奠定了基础,不仅提供了实用的工具来创建有效的可视化表示,还能够可靠地评定其质量和有效性。

URL

https://arxiv.org/abs/2509.11942

PDF

https://arxiv.org/pdf/2509.11942.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot