Paper Reading AI Learner

FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization

2025-07-17 17:30:29
Chuancheng Shi, Yixiang Chen, Burong Lei, Jichao Chen

Abstract

Realistic and controllable garment visualization is critical for fashion e-commerce, where users expect personalized previews under diverse poses and lighting conditions. Existing methods often rely on predefined poses, limiting semantic flexibility and illumination adaptability. To address this, we introduce FashionPose, the first unified text-to-pose-to-relighting generation framework. Given a natural language description, our method first predicts a 2D human pose, then employs a diffusion model to generate high-fidelity person images, and finally applies a lightweight relighting module, all guided by the same textual input. By replacing explicit pose annotations with text-driven conditioning, FashionPose enables accurate pose alignment, faithful garment rendering, and flexible lighting control. Experiments demonstrate fine-grained pose synthesis and efficient, consistent relighting, providing a practical solution for personalized virtual fashion display.

Abstract (translated)

逼真的服装可视化对于时装电子商务至关重要,用户期望在不同的姿势和光照条件下获得个性化的预览效果。现有的方法通常依赖于预定义的姿势,这限制了语义灵活性和照明适应性。为了解决这些问题,我们引入了FashionPose,这是首个统一的文本到姿态再到重新布光生成框架。给定自然语言描述后,我们的方法首先预测2D人体姿态,然后使用扩散模型生成高保真的人体图像,并最终应用一个轻量级的重新布光模块,所有这些步骤都由相同的文本输入指导。通过用基于文本驱动的条件取代显式的姿势标注,FashionPose能够实现准确的姿态对齐、忠实的服装渲染和灵活的光照控制。实验结果表明,Fine-grained姿态合成及高效的、一致性的重新布光可以为个性化的虚拟时装展示提供实用解决方案。

URL

https://arxiv.org/abs/2507.13311

PDF

https://arxiv.org/pdf/2507.13311.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot