Abstract
Realistic and controllable garment visualization is critical for fashion e-commerce, where users expect personalized previews under diverse poses and lighting conditions. Existing methods often rely on predefined poses, limiting semantic flexibility and illumination adaptability. To address this, we introduce FashionPose, the first unified text-to-pose-to-relighting generation framework. Given a natural language description, our method first predicts a 2D human pose, then employs a diffusion model to generate high-fidelity person images, and finally applies a lightweight relighting module, all guided by the same textual input. By replacing explicit pose annotations with text-driven conditioning, FashionPose enables accurate pose alignment, faithful garment rendering, and flexible lighting control. Experiments demonstrate fine-grained pose synthesis and efficient, consistent relighting, providing a practical solution for personalized virtual fashion display.
Abstract (translated)
逼真的服装可视化对于时装电子商务至关重要,用户期望在不同的姿势和光照条件下获得个性化的预览效果。现有的方法通常依赖于预定义的姿势,这限制了语义灵活性和照明适应性。为了解决这些问题,我们引入了FashionPose,这是首个统一的文本到姿态再到重新布光生成框架。给定自然语言描述后,我们的方法首先预测2D人体姿态,然后使用扩散模型生成高保真的人体图像,并最终应用一个轻量级的重新布光模块,所有这些步骤都由相同的文本输入指导。通过用基于文本驱动的条件取代显式的姿势标注,FashionPose能够实现准确的姿态对齐、忠实的服装渲染和灵活的光照控制。实验结果表明,Fine-grained姿态合成及高效的、一致性的重新布光可以为个性化的虚拟时装展示提供实用解决方案。
URL
https://arxiv.org/abs/2507.13311