Abstract
Large Language Models (LLMs) have revolutionized natural language processing with their remarkable capabilities in text generation and reasoning. However, these models face critical challenges when deployed in real-world applications, including hallucination generation, outdated knowledge, and limited domain expertise. Retrieval And Structuring (RAS) Augmented Generation addresses these limitations by integrating dynamic information retrieval with structured knowledge representations. This survey (1) examines retrieval mechanisms including sparse, dense, and hybrid approaches for accessing external knowledge; (2) explore text structuring techniques such as taxonomy construction, hierarchical classification, and information extraction that transform unstructured text into organized representations; and (3) investigate how these structured representations integrate with LLMs through prompt-based methods, reasoning frameworks, and knowledge embedding techniques. It also identifies technical challenges in retrieval efficiency, structure quality, and knowledge integration, while highlighting research opportunities in multimodal retrieval, cross-lingual structures, and interactive systems. This comprehensive overview provides researchers and practitioners with insights into RAS methods, applications, and future directions.
Abstract (translated)
大型语言模型(LLMs)通过其在文本生成和推理方面的卓越能力彻底革新了自然语言处理。然而,这些模型在实际应用中面临一些关键挑战,包括幻觉产生、知识过时以及领域专业知识的局限性。检索与结构化(RAS)增强型生成旨在通过整合动态信息检索与结构化知识表示来解决这些问题。这份调查报告: 1. 检视了用于访问外部知识的检索机制,包括稀疏检索、密集检索和混合方法; 2. 探讨了将非结构化文本转换为组织良好的表示的方法,如分类法构建、层次分类和信息提取等文本结构技术;以及 3. 研究这些结构化的表示如何通过基于提示的方法、推理框架和技术知识嵌入与LLMs集成。 该报告还指出了检索效率、结构质量及知识整合方面存在的技术挑战,并强调了多模态检索、跨语言结构和交互式系统等研究机会。这份全面的概览为研究人员和从业者提供了关于RAS方法、应用以及未来发展方向的见解。
URL
https://arxiv.org/abs/2509.10697