Pith. sign in

REVIEW 2 major objections 1 minor 81 references

AvatarMix composes outfits by directly mixing head and body from two Gaussian avatars using mesh retargeting and refinement modules.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-28 10:20 UTC pith:ISBZ5OFR

load-bearing objection AvatarMix gives a direct composition route for 3D Gaussian outfit transfer using mesh retargeting plus two diffusion fixes, but the SOTA claim sits on unshown experiments and an optional correction step. the 2 major comments →

arxiv 2606.03506 v1 pith:ISBZ5OFR submitted 2026-06-02 cs.CV cs.GR

AvatarMix: Identity-Preserving Cross-Avatar Composition for Outfit Personalization

classification cs.CV cs.GR
keywords 3D avataroutfit transferGaussian avataridentity preservationmesh retargetingdiffusion modelavatar compositionpersonalization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that composing the head from one high-fidelity 3D Gaussian avatar and the body from another can personalize outfits while keeping both the outfit details and the user's body identity intact. This matters because previous methods either degrade quality when lifting 2D changes to 3D or create intersections when modeling layers separately. The approach uses targeted fixes at the seam and optionally across the body to make the join look natural after reshaping. If successful, it offers a simpler path to realistic 3D avatar customization without the typical trade-offs in fidelity.

Core claim

AvatarMix introduces a compositional paradigm that directly composes the head and body from two high-fidelity Gaussian avatars. This bypasses quality degradation and intersection artifacts by avoiding 2D-to-3D lifting and layered modeling. A two-tier refinement strategy with SeamFix for hair and neck joins and optional FullbodyFix for garment appearance is applied to 3D-consistent renders. Mesh-based retargeting adapts the clothed body to the user's physique, enabling robust handling of diverse body shapes and achieving state-of-the-art results in outfit fidelity and identity preservation.

What carries the argument

The two-tier refinement strategy of SeamFix, a localized diffusion module for artifact-free joins at hair and neck, and FullbodyFix for restoring garment appearance after retargeting, operating on renders from mesh-retargeted Gaussian avatars.

Load-bearing premise

The two input avatars are already high-fidelity Gaussian representations and the mesh retargeting can be applied to clothed bodies without introducing unfixable appearance degradation.

What would settle it

Visual inspection or metric scores on test cases showing visible seams at the neck or loss of outfit details after composition and refinement would indicate the method does not achieve seamless and faithful results.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Outfit personalization avoids intersection artifacts by not using separate clothing layers.
  • Body reshaping preserves appearance through adaptation of robust mesh retargeting to clothed Gaussians.
  • Refinements on 3D-consistent renders limit multi-view artifacts compared to 2D methods.
  • Direct composition maintains outfit quality without the degradation common in lifting approaches.
  • The method enables handling of diverse body shapes while keeping identity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Similar composition techniques could be tested for transferring other features like hairstyles or accessories across avatars.
  • Applying the method to avatars generated from single images rather than high-fidelity scans could test its robustness in less controlled settings.
  • Integration with animation pipelines might allow dynamic outfit changes during motion without re-rendering issues.
  • Quantitative comparisons on standard benchmarks for 3D avatar editing would help validate the claimed improvements over baselines.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript presents AvatarMix, a compositional method for 3D outfit personalization that directly composes the head and body from two high-fidelity 3D Gaussian avatars. It addresses challenges in outfit transfer by using a mesh-based Gaussian representation to enable body reshaping via mesh retargeting, while introducing SeamFix for seamless joins and an optional FullbodyFix for restoring appearance after retargeting. The paper claims this approach achieves state-of-the-art performance in outfit fidelity and identity preservation, offering a new paradigm that avoids intersection artifacts and quality degradation.

Significance. If the results hold, this work offers a new compositional paradigm for 3D avatar editing that maintains outfit quality and 3D consistency better than lifting-based or layered approaches. The mesh-based retargeting on Gaussians and refinement on already-consistent renders are potentially impactful strengths for virtual try-on applications.

major comments (2)
  1. [Abstract] Abstract: The SOTA claim in identity preservation and outfit fidelity rests on mesh retargeting of clothed Gaussian bodies. The text states that retargeting can degrade the clothed body, with FullbodyFix described as optional to restore garment appearance. This makes preservation conditional rather than guaranteed; without quantitative evidence (e.g., ablation on FullbodyFix necessity, multi-view error metrics, or cases where the fix is not applied) the central claim is undermined. The stress-test concern about unfixable distortions applies directly.
  2. [Method] Method (FullbodyFix and retargeting description): The optional status of FullbodyFix is load-bearing for the identity-preservation guarantee. If retargeting on diverse clothed bodies produces distortions (stretching, folds, texture shifts) that the diffusion fix cannot reliably correct across views, the claim fails. The manuscript should either integrate the fix or provide tests showing robustness without it.
minor comments (1)
  1. [Abstract] Abstract: The assertion of 'extensive experiments' demonstrating SOTA lacks any quantitative details, baselines, or error analysis. A one-sentence summary of key metrics would improve the abstract.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the detailed and constructive feedback. The comments correctly identify that the optional status of FullbodyFix requires stronger justification to support the identity-preservation and SOTA claims. We address each point below and will revise the manuscript accordingly.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The SOTA claim in identity preservation and outfit fidelity rests on mesh retargeting of clothed Gaussian bodies. The text states that retargeting can degrade the clothed body, with FullbodyFix described as optional to restore garment appearance. This makes preservation conditional rather than guaranteed; without quantitative evidence (e.g., ablation on FullbodyFix necessity, multi-view error metrics, or cases where the fix is not applied) the central claim is undermined. The stress-test concern about unfixable distortions applies directly.

    Authors: We acknowledge that describing FullbodyFix as optional without accompanying quantitative ablations leaves the central claim vulnerable, as the referee notes. The manuscript text does state that retargeting can degrade appearance and positions the fix as optional for cases where degradation occurs. To resolve this, we will add an ablation study reporting identity and outfit fidelity metrics (including multi-view consistency) with and without FullbodyFix across diverse body shapes and garments. We will also revise the abstract to clarify that the reported SOTA results use the full pipeline including refinement where needed. revision: yes

  2. Referee: [Method] Method (FullbodyFix and retargeting description): The optional status of FullbodyFix is load-bearing for the identity-preservation guarantee. If retargeting on diverse clothed bodies produces distortions (stretching, folds, texture shifts) that the diffusion fix cannot reliably correct across views, the claim fails. The manuscript should either integrate the fix or provide tests showing robustness without it.

    Authors: The referee is correct that the optional status is load-bearing and that robustness without the fix must be demonstrated if it remains optional. Our current experiments indicate that mesh retargeting preserves appearance in many cases, but we agree additional evidence is required. We will revise the method section to present FullbodyFix as an integrated, recommended component of the pipeline rather than purely optional, and include new experiments quantifying performance with and without it, including stress tests on diverse clothed bodies and multi-view error metrics. revision: yes

Circularity Check

0 steps flagged

No derivation chain or equations; method is engineering composition

full rationale

The paper presents AvatarMix as a compositional pipeline that directly combines head and body from two pre-existing high-fidelity Gaussian avatars, then applies existing mesh retargeting plus optional diffusion-based fixes (SeamFix, FullbodyFix). No equations, first-principles derivations, fitted parameters, or predictions appear in the provided text. The central claims rest on the engineering choice of mesh-based Gaussians to enable retargeting, not on any self-referential reduction or self-citation chain that would make the result equivalent to its inputs by construction. This is a standard applied CV method paper whose validity is evaluated by external experiments rather than internal definitional closure.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Review performed on abstract only; no explicit free parameters, axioms, or invented entities are stated. Standard assumptions of Gaussian splatting and diffusion models are implicit but not detailed.

pith-pipeline@v0.9.1-grok · 5784 in / 1046 out tokens · 27353 ms · 2026-06-28T10:20:16.056612+00:00 · methodology

0 comments
read the original abstract

Existing 3D avatar outfit transfer methods face distinct challenges: approaches that lift 2D edits to 3D often suffer from outfit or identity quality degradation, while those that separately model body and clothing layers are prone to intersection artifacts. We introduce AvatarMix, a compositional paradigm that bypasses these issues by directly composing the head and body from two high-fidelity Gaussian avatars. While this paradigm inherently preserves outfit quality and avoids intersections, it introduces challenges in creating a seamless join and maintaining appearance fidelity after body reshaping. To this end, we propose a two-tier refinement strategy: SeamFix, a localized diffusion module that refines hair and neck to ensure an artifact-free join, and an optional full-body refinement, FullbodyFix, that restores garment appearance when retargeting degrades the clothed body. Both operate on renders from an already 3D-consistent Gaussian avatar, which limits multi-view artifacts compared to 2D-to-3D lifting. To preserve the user's body identity, our mesh-based Gaussian representation enables the adaptation of a robust mesh retargeting technique, precisely reshaping the clothed body to the user's physique and robustly handling diverse body shapes. Extensive experiments demonstrate that our method achieves state-of-the-art results in outfit fidelity and identity preservation, providing a new perspective for realistic 3D outfit personalization. Project page: https://larsph.github.io/avatarmix/

Figures

Figures reproduced from arXiv: 2606.03506 by Yoshihiro Kanamori, Yuki Endo, Zhaorong Wang.

Figure 1
Figure 1. Figure 1: AvatarMix performs free-viewpoint outfit personalization by composing a user’s identity cues (head–neck, body shape and scale, [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of AvatarMix. Given multi-view images of a User and a Model, we first employ Mesh-Based Avatar Reconstruction (Sec. 3.1) with semantic segmentation. We then perform Cross-Avatar Geometric Composition (Sec. 3.2) by aligning the user’s head and neck to the Model’s pose (Head & Neck Alignment) and reshaping the Model’s clothed body via our GSReshape module (Body Alignment & Reshaping) so that the bod… view at source ↗
Figure 3
Figure 3. Figure 3: Training strategy for SeamFix and FullbodyFix. Top: starting from two avatars A and B, we deliberately introduce seg￾mentation errors by using the original 4D-Dress with SAM voting and perform a first head-swap A→B, which produces Gaussian ar￾tifacts at the head, neck, and hands. After re-segmentation and a second head-swap B→A, we obtain double-swapped avatars that are pixel-aligned with the ground-truth … view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison with TIP-Editor and VTON360. For each user–model pair, we show the input user and model images (front/back under two lighting conditions), followed by three-view outputs of TIP-Editor, VTON360, and AvatarMix. Zoomed insets highlight faces and garment regions, and red dashed boxes mark typical failure cases of existing methods, including view inconsistency, unnatural garment wrinkles,… view at source ↗
Figure 5
Figure 5. Figure 5: Ablation of diffusion refinement and GSReshape. Left: rows show compositions without SeamFix, with SeamFix, without FullbodyFix, and with FullbodyFix for three user–model pairs, illustrating how SeamFix cleans head–neck seams and Full￾bodyFix restores garment appearance while preserving face and outfit details. Right: for two subjects, we compare the reference model body, composition without GSReshape, and… view at source ↗
Figure 6
Figure 6. Figure 6: GSReshape pipeline overview. From left to right: starting from the model’s low-resolution clothed mesh (top left) and SMPL-X mesh (bottom left), we project the SMPL mesh to skeleton, inflating the SMPL mesh while jointly optimizing the clothed mesh, following the retargeting method of Huang et al. [29]. After retargeting, we compute vertex offsets between input and retargeted clothed mesh, and transfer the… view at source ↗
Figure 7
Figure 7. Figure 7: Hand-aware skin tightness examples. First row: high fit weight produces Gaussian artifacts (left) versus our hand shape preserving method (right). Second row: low fit weight creates glove-like hands (left) versus our approach (right). Our semantic weighting strategy achieves better balance between visual fidelity and robustness. used in GSReshape. During the subsequent mesh retarget￾ing, let Vg = {vi} and … view at source ↗
Figure 8
Figure 8. Figure 8: Additional ablation on GSReshape. We visualize the effect of our body reshaping module by comparing the model avatars without GSReshape versus with GSReshape. As shown in the with GSReshape results, the garment adapts smoothly to the user’s body shape while preserving details after the body reshap￾ing. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Additional comparisons with THUman2.0. We compare AvatarMix with baselines on more user-model pairs, demonstrating superior preservation of identity and outfit across diverse views. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

81 extracted references · 7 canonical work pages

  1. [1]

    Robust skin weights transfer via weight inpainting

    Rinat Abdrashitov, Kim Raichstat, Jared Monsen, and David Hill. Robust skin weights transfer via weight inpainting. InSIGGRAPH Asia 2023 Technical Communications, pages 25:1–25:4. ACM, 2023. 5

  2. [2]

    MET3R: measuring multi-view consistency in generated images

    Mohammad Asim, Christopher Wewer, Thomas Wimmer, Bernt Schiele, and Jan Eric Lenssen. MET3R: measuring multi-view consistency in generated images. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025, pages 6034–6044. Computer Vision Foundation / IEEE, 2025. 8, 14

  3. [3]

    Driving-signal aware full-body avatars.ACM TOG, 40(4):1–17, 2021

    Timur Bagautdinov, Chenglei Wu, Tomas Simon, Fabian Prada, Takaaki Shiratori, Shih-En Wei, Weipeng Xu, Yaser Sheikh, and Jason Saragih. Driving-signal aware full-body avatars.ACM TOG, 40(4):1–17, 2021. 3

  4. [4]

    Multi-garment net: Learning to dress 3D people from images

    Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, and Gerard Pons-Moll. Multi-garment net: Learning to dress 3D people from images. InCVPR, pages 5420–5430, 2019. 3

  5. [5]

    Robust treatment of collisions, contact and friction for cloth anima- tion

    Robert Bridson, Ronald Fedkiw, and John Anderson. Robust treatment of collisions, contact and friction for cloth anima- tion. InProceedings of the 29th annual conference on Com- puter graphics and interactive techniques, pages 594–603,

  6. [6]

    Design preserving garment transfer.ACM TOG, 31(4):36:1–36:11, 2012

    R ´emi Brouet, Alla Sheffer, Laurence Boissieux, and Marie- Paule Cani. Design preserving garment transfer.ACM TOG, 31(4):36:1–36:11, 2012. 3

  7. [7]

    GS- VTON: Controllable 3D virtual try-on with Gaussian splat- ting, 2024

    Yukang Cao, Masoud Hadi, Liang Pan, and Ziwei Liu. GS- VTON: Controllable 3D virtual try-on with Gaussian splat- ting, 2024. 2, 3, 13

  8. [8]

    Segment anything in 3D with NeRFs

    Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang, Wei Shen, Lingxi Xie, Xiaopeng Zhang, and Qi Tian. Segment anything in 3D with NeRFs. InNeurIPS, 2023. 3

  9. [9]

    GaussianVTON: 3D human virtual try-on via multi-stage Gaussian splatting editing with image prompting, 2024

    Haodong Chen, Yongle Huang, Haojian Huang, Xiangsheng Ge, and Dian Shao. GaussianVTON: 3D human virtual try-on via multi-stage Gaussian splatting editing with image prompting, 2024. 2, 3, 13

  10. [10]

    GGAvatar: Reconstructing garment- separated 3D Gaussian splatting avatars from monocular video

    Jingxuan Chen. GGAvatar: Reconstructing garment- separated 3D Gaussian splatting avatars from monocular video. InACM MMAsia, pages 80:1–80:7, 2024. 2, 3

  11. [11]

    Taoa- vatar: Real-time lifelike full-body talking avatars for aug- mented reality via 3D Gaussian splatting

    Jianchuan Chen, Jingchuan Hu, Gaige Wang, Zhonghua Jiang, Tiansong Zhou, Zhiwen Chen, and Chengfei Lv. Taoa- vatar: Real-time lifelike full-body talking avatars for aug- mented reality via 3D Gaussian splatting. InCVPR, pages 10723–10734, 2025. 3

  12. [12]

    Gaussianeditor: Swift and controllable 3D editing with Gaussian splatting

    Yiwen Chen, Zilong Chen, Chi Zhang, Feng Wang, Xiaofeng Yang, Yikai Wang, Zhongang Cai, Lei Yang, Huaping Liu, and Guosheng Lin. Gaussianeditor: Swift and controllable 3D editing with Gaussian splatting. InCVPR, pages 21476– 21485. IEEE, 2024. 3

  13. [13]

    Splatformer: Point trans- former for robust 3D Gaussian splatting

    Yutong Chen, Marko Mihajlovic, Xiyi Chen, Yiming Wang, Sergey Prokudin, and Siyu Tang. Splatformer: Point trans- former for robust 3D Gaussian splatting. InICLR. OpenRe- view.net, 2025. 3

  14. [14]

    Cosseggaussians: Compact and swift scene segmenting 3D gaussians with dual feature fusion, 2024

    Bin Dou, Tianyu Zhang, Zhaohui Wang, Yongjia Ma, and Zejian Yuan. Cosseggaussians: Compact and swift scene segmenting 3D gaussians with dual feature fusion, 2024. 3

  15. [15]

    Black, and Timo Bolkart

    Yao Feng, Jinlong Yang, Marc Pollefeys, Michael J. Black, and Timo Bolkart. Capturing and animation of body and clothing from monocular video. InSIGGRAPH Asia Confer- ence Proceedings, 2022. 3

  16. [16]

    Learning disentangled avatars with hybrid 3D representations.arXiv preprint arXiv:2309.06441, 2023

    Yao Feng, Weiyang Liu, Timo Bolkart, Jinlong Yang, Marc Pollefeys, and Michael J Black. Learning disentangled avatars with hybrid 3D representations.arXiv preprint arXiv:2309.06441, 2023. 3

  17. [17]

    Parser-free virtual try-on via distilling appearance flows

    Yuying Ge, Yibing Song, Ruimao Zhang, Chongjian Ge, Wei Liu, and Ping Luo. Parser-free virtual try-on via distilling appearance flows. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 8485–8493, 2021. 3

  18. [18]

    Taming the power of diffusion models for high-quality virtual try-on with appearance flow

    Junhong Gou, Siyu Sun, Jianfu Zhang, Jianlou Si, Chen Qian, and Liqing Zhang. Taming the power of diffusion models for high-quality virtual try-on with appearance flow. arXiv preprint arXiv:2308.06101, 2023. 3

  19. [19]

    3D human avatar reconstruction with neural fields: A recent survey.Image and Vision Computing, 154: 105341, 2025

    Meiying Gu, Jiahe Li, Yuchen Wu, Haonan Luo, Jin Zheng, and Xiao Bai. 3D human avatar reconstruction with neural fields: A recent survey.Image and Vision Computing, 154: 105341, 2025. 2, 3

  20. [20]

    Hirshberg, Alexander Weiss, and Michael J

    Peng Guan, Loretta Reiss, David A. Hirshberg, Alexander Weiss, and Michael J. Black. DRAPE: dressing any person. ACM TOG, 31(4):35:1–35:10, 2012. 3

  21. [21]

    Vid2avatar-pro: Authentic avatar from videos in the wild via universal prior

    Chen Guo, Junxuan Li, Yash Kant, Yaser Sheikh, Shunsuke Saito, and Chen Cao. Vid2avatar-pro: Authentic avatar from videos in the wild via universal prior. InCVPR, pages 5559–

  22. [22]

    Computer Vision Foundation / IEEE, 2025. 3, 15

  23. [23]

    Subspace clothing simulation using adaptive bases.ACM TOG, 33(4):1–9, 2014

    Fabian Hahn, Bernhard Thomaszewski, Stelian Coros, Robert W Sumner, Forrester Cole, Mark Meyer, Tony DeRose, and Markus Gross. Subspace clothing simulation using adaptive bases.ACM TOG, 33(4):1–9, 2014. 3

  24. [24]

    Viton: An image-based virtual try-on network

    Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S Davis. Viton: An image-based virtual try-on network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7543–7552, 2018. 3

  25. [25]

    Efros, Aleksander Holynski, and Angjoo Kanazawa

    Ayaan Haque, Matthew Tancik, Alexei A. Efros, Aleksander Holynski, and Angjoo Kanazawa. Instruct-nerf2nerf: Edit- ing 3d scenes with instructions. InICCV, pages 19683– 19693. IEEE, 2023. 14

  26. [26]

    VTON 360: High-fidelity virtual try-on from any viewing direction

    Zijian He, Yuwei Ning, Yipeng Qin, Guangrun Wang, Sibei Yang, Liang Lin, and Guanbin Li. VTON 360: High-fidelity virtual try-on from any viewing direction. InCVPR, pages 26388–26398, 2025. 2, 3, 6, 13

  27. [27]

    Gauhuman: Articu- lated Gaussian splatting from monocular human videos

    Shoukang Hu, Tao Hu, and Ziwei Liu. Gauhuman: Articu- lated Gaussian splatting from monocular human videos. In CVPR, pages 20418–20431. IEEE, 2024. 2

  28. [28]

    Dreamwaltz: Make a scene with complex 3D animatable avatars.NeurIPS, 36, 2024

    Yukun Huang, Jianan Wang, Ailing Zeng, He Cao, Xianbiao Qi, Yukai Shi, Zheng-Jun Zha, and Lei Zhang. Dreamwaltz: Make a scene with complex 3D animatable avatars.NeurIPS, 36, 2024. 3 9

  29. [29]

    Tech: Text-guided reconstruction of lifelike clothed humans

    Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Ji- axiang Tang, Deng Cai, and Justus Thies. Tech: Text-guided reconstruction of lifelike clothed humans. In2024 Interna- tional Conference on 3D Vision (3DV), pages 1531–1542. IEEE, 2024. 3

  30. [30]

    Intersection- free garment retargeting

    Zizhou Huang, Chrystiano Ara ´ujo, Andrew Kunz, Denis Zorin, Daniele Panozzo, and Victor Zordan. Intersection- free garment retargeting. InACM SIGGRAPH Conference Papers, New York, NY , USA, 2025. 2, 3, 4, 5, 8, 12, 13

  31. [31]

    GaussianBlock: Building part-aware compositional and editable 3D scene by primitives and gaus- sians

    Shuyi Jiang, Qihao Zhao, Hossein Rahmani, De Wen Soh, Jun Liu, and Na Zhao. GaussianBlock: Building part-aware compositional and editable 3D scene by primitives and gaus- sians. InICLR, 2025. 2, 3

  32. [32]

    Virtual try-on by replacing the person in image.Journal of Computer-Aided Design & Computer Graphics, 27(9):1694–1700, 2015

    Li Jun, Zhang Mingmin, and Pan Zhigeng. Virtual try-on by replacing the person in image.Journal of Computer-Aided Design & Computer Graphics, 27(9):1694–1700, 2015. 3

  33. [33]

    3D Gaussian splatting for real-time radiance field rendering.ACM TOG, 42(4):139:1–139:14,

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian splatting for real-time radiance field rendering.ACM TOG, 42(4):139:1–139:14,

  34. [34]

    Gala: Generating animatable layered assets from a sin- gle scan

    Taeksoo Kim, Byungjun Kim, Shunsuke Saito, and Hanbyul Joo. Gala: Generating animatable layered assets from a sin- gle scan. InCVPR, 2024. 3

  35. [35]

    Deepwrin- kles: Accurate and realistic clothing modeling

    Zorah Lahner, Daniel Cremers, and Tony Tung. Deepwrin- kles: Accurate and realistic clothing modeling. InECCV, pages 667–684, 2018. 3

  36. [36]

    DiffAvatar: Simulation-Ready Garment Optimization with Differen- tiable Simulation

    Yifei Li, Hsiao yu Chen, Egor Larionov, Nikolaos Sarafi- anos, Wojciech Matusik, and Tuur Stuyck. DiffAvatar: Simulation-Ready Garment Optimization with Differen- tiable Simulation. InCVPR, 2024. 3

  37. [37]

    Animatable Gaussians: Learning pose-dependent Gaussian maps for high-fidelity human avatar modeling

    Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Animatable Gaussians: Learning pose-dependent Gaussian maps for high-fidelity human avatar modeling. InCVPR, pages 19711–19722. IEEE, 2024. 2, 3, 15

  38. [38]

    Graphonomy: Universal image parsing via graph rea- soning and transfer.IEEE Trans

    Liang Lin, Yiming Gao, Ke Gong, Meng Wang, and Xiaodan Liang. Graphonomy: Universal image parsing via graph rea- soning and transfer.IEEE Trans. Pattern Anal. Mach. Intell., 44(5):2504–2518, 2022. 6

  39. [39]

    LayGA: Layered Gaussian avatars for animatable clothing transfer

    Siyou Lin, Zhe Li, Zhaoqi Su, Zerong Zheng, Hongwen Zhang, and Yebin Liu. LayGA: Layered Gaussian avatars for animatable clothing transfer. InACM SIGGRAPH Con- ference Papers, 2024. 2, 3, 13, 15

  40. [40]

    Creating your ed- itable 3D photorealistic avatar with tetrahedron-constrained Gaussian splatting

    Hanxi Liu, Yifang Men, and Zhouhui Lian. Creating your ed- itable 3D photorealistic avatar with tetrahedron-constrained Gaussian splatting. InCVPR, pages 15976–15986, 2025. 2

  41. [41]

    Deceptive-nerf/3dgs: Diffusion- generated pseudo-observations for high-quality sparse-view reconstruction

    Xinhang Liu, Jiaben Chen, Shiu-hong Kao, Yu-Wing Tai, and Chi-Keung Tang. Deceptive-nerf/3dgs: Diffusion- generated pseudo-observations for high-quality sparse-view reconstruction. InECCV, 2024. 3

  42. [42]

    3dgs- enhancer: Enhancing unbounded 3D Gaussian splatting with view-consistent 2d diffusion priors.arXiv preprint arXiv:2410.16266, 2024

    Xi Liu, Chaoyi Zhou, and Siyu Huang. 3dgs- enhancer: Enhancing unbounded 3D Gaussian splatting with view-consistent 2d diffusion priors.arXiv preprint arXiv:2410.16266, 2024. 3

  43. [43]

    Qianli Ma, Jinlong Yang, Anurag Ranjan, Sergi Pujades, Gerard Pons-Moll, Siyu Tang, and Michael J. Black. Learn- ing to dress 3D people in generative clothing. InCVPR, pages 6468–6477. Computer Vision Foundation / IEEE,

  44. [44]

    Learn- ing to transfer texture from clothing images to 3D humans

    Aymen Mir, Thiemo Alldieck, and Gerard Pons-Moll. Learn- ing to transfer texture from clothing images to 3D humans. InCVPR, pages 7023–7034, 2020. 3

  45. [45]

    Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rab- bat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herv ´e J´egou, Julien Mairal, P...

  46. [46]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. InCVPR, pages 10975– 10985. Computer Vision Foundation / IEEE, 2019. 5

  47. [47]

    Clothcap: Seamless 4d clothing capture and retarget- ing.ACM TOG, 36(4):1–15, 2017

    Gerard Pons-Moll, Sergi Pujades, Sonny Hu, and Michael J Black. Clothcap: Seamless 4d clothing capture and retarget- ing.ACM TOG, 36(4):1–15, 2017. 3

  48. [48]

    Splattingavatar: Realistic real-time human avatars with mesh-embedded Gaussian splatting

    Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang. Splattingavatar: Realistic real-time human avatars with mesh-embedded Gaussian splatting. InCVPR, pages 1606–

  49. [49]

    Garment personalization via identity transfer.IEEE Computer Graphics and Applications, 33(4):62–72, 2013

    Roy Shilkrot, Daniel Cohen-Or, Ariel Shamir, and Ligang Liu. Garment personalization via identity transfer.IEEE Computer Graphics and Applications, 33(4):62–72, 2013. 2, 3

  50. [50]

    Caphy: Cap- turing physical properties for animatable human avatars

    Zhaoqi Su, Liangxiao Hu, Siyou Lin, Hongwen Zhang, Shengping Zhang, Justus Thies, and Yebin Liu. Caphy: Cap- turing physical properties for animatable human avatars. In ICCV, 2023. 3

  51. [51]

    Fenerf: Face editing in neural radiance fields

    Jingxiang Sun, Xuan Wang, Yong Zhang, Xiaoyu Li, Qi Zhang, Yebin Liu, and Jue Wang. Fenerf: Face editing in neural radiance fields. InCVPR, pages 7662–7672. IEEE,

  52. [52]

    Learning semantic- aware disentangled representation for flexible 3d human body editing

    Xiaokun Sun, Qiao Feng, Xiongzheng Li, Jinsong Zhang, Yu-Kun Lai, Jingyu Yang, and Kun Li. Learning semantic- aware disentangled representation for flexible 3d human body editing. InCVPR, pages 16985–16994. IEEE, 2023. 3

  53. [53]

    Mega: Hybrid mesh-Gaussian head avatar for high-fidelity rendering and head editing

    Cong Wang, Di Kang, Heyi Sun, Shen-Han Qian, Zixuan Wang, Linchao Bao, and Song-Hai Zhang. Mega: Hybrid mesh-Gaussian head avatar for high-fidelity rendering and head editing. InCVPR, pages 26274–26284. Computer Vi- sion Foundation / IEEE, 2025. 3

  54. [54]

    Gaussianeditor: Editing 3D gaussians delicately with text instructions

    Junjie Wang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie, and Qi Tian. Gaussianeditor: Editing 3D gaussians delicately with text instructions. InCVPR, pages 20902–20911. IEEE,

  55. [55]

    Ruihe Wang, Yukang Cao, Kai Han, and Kwan-Yee K. Wong. A survey on 3D human avatar modeling – from re- construction to generation, 2024. 2, 3

  56. [56]

    4d-dress: A 4d dataset of real-world human clothing with 10 semantic annotations

    Wenbo Wang, Hsuan-I Ho, Chen Guo, Boxiang Rong, Artur Grigorev, Jie Song, Juan Jose Zarate, and Otmar Hilliges. 4d-dress: A 4d dataset of real-world human clothing with 10 semantic annotations. InCVPR, pages 550–560. IEEE, 2024. 4

  57. [57]

    Neus2: Fast learning of neural implicit surfaces for multi-view recon- struction

    Yiming Wang, Qin Han, Marc Habermann, Kostas Dani- ilidis, Christian Theobalt, and Lingjie Liu. Neus2: Fast learning of neural implicit surfaces for multi-view recon- struction. InICCV, pages 3272–3283. IEEE, 2023. 4

  58. [58]

    Gs- fix3d: Diffusion-guided repair of novel views in Gaussian splatting, 2025

    Jiaxin Wei, Stefan Leutenegger, and Simon Schaefer. Gs- fix3d: Diffusion-guided repair of novel views in Gaussian splatting, 2025. 3

  59. [59]

    Reid, Philip Torr, and Victor Adrian Prisacariu

    Jing Wu, Jia-Wang Bian, Xinghui Li, Guangrun Wang, Ian D. Reid, Philip Torr, and Victor Adrian Prisacariu. Gaussctrl: Multi-view consistent text-driven 3D Gaussian splatting editing. InECCV, pages 55–71. Springer, 2024. 3

  60. [60]

    DIFIX3D+: improving 3D reconstructions with single-step diffusion models

    Jay Zhangjie Wu, Yuxuan Zhang, Haithem Turki, Xuanchi Ren, Jun Gao, Mike Zheng Shou, Sanja Fidler, Zan Gojcic, and Huan Ling. DIFIX3D+: improving 3D reconstructions with single-step diffusion models. InCVPR, pages 26024– 26035. Computer Vision Foundation / IEEE, 2025. 3, 8, 12

  61. [61]

    Modeling clothing as a separate layer for an animatable hu- man avatar.ACM TOG, 40(6):1–15, 2021

    Donglai Xiang, Fabian Prada, Timur Bagautdinov, Weipeng Xu, Yuan Dong, He Wen, Jessica Hodgins, and Chenglei Wu. Modeling clothing as a separate layer for an animatable hu- man avatar.ACM TOG, 40(6):1–15, 2021. 3

  62. [62]

    Dressing avatars: Deep photorealistic appearance for physically simu- lated clothing.ACM TOG, 41(6):1–15, 2022

    Donglai Xiang, Timur Bagautdinov, Tuur Stuyck, Fabian Prada, Javier Romero, Weipeng Xu, Shunsuke Saito, Jing- fan Guo, Breannan Smith, Takaaki Shiratori, et al. Dressing avatars: Deep photorealistic appearance for physically simu- lated clothing.ACM TOG, 41(6):1–15, 2022. 3

  63. [63]

    Dreamvton: Customizing 3D virtual try- on with personalized diffusion models.arXiv preprint arXiv:2407.16511, 2024

    Zhenyu Xie, Haoye Dong, Yufei Gao, Zehua Ma, and Xi- aodan Liang. Dreamvton: Customizing 3D virtual try- on with personalized diffusion models.arXiv preprint arXiv:2407.16511, 2024. 3

  64. [64]

    Oot- diffusion: Outfitting fusion based latent diffusion for control- lable virtual try-on.arXiv preprint arXiv:2403.01779, 2024

    Yuhao Xu, Tao Gu, Weifeng Chen, and Chengcai Chen. Oot- diffusion: Outfitting fusion based latent diffusion for control- lable virtual try-on.arXiv preprint arXiv:2403.01779, 2024. 3

  65. [65]

    Gaussianob- ject: High-quality 3D object reconstruction from four views with Gaussian splatting.ACM TOG, 43(6):199:1–199:13,

    Chen Yang, Sikuang Li, Jiemin Fang, Ruofan Liang, Lingxi Xie, Xiaopeng Zhang, Wei Shen, and Qi Tian. Gaussianob- ject: High-quality 3D object reconstruction from four views with Gaussian splatting.ACM TOG, 43(6):199:1–199:13,

  66. [66]

    Makeup prior models for 3D facial makeup estimation and applications

    Xingchao Yang, Takafumi Taketomi, Yuki Endo, and Yoshi- hiro Kanamori. Makeup prior models for 3D facial makeup estimation and applications. InCVPR, pages 2165–2175. IEEE, 2024. 3

  67. [67]

    Gaussian grouping: Segment and edit anything in 3D scenes

    Mingqiao Ye, Martin Danelljan, Fisher Yu, and Lei Ke. Gaussian grouping: Segment and edit anything in 3D scenes. InECCV, pages 162–179. Springer, 2024. 3

  68. [68]

    Omniseg3d: Omniversal 3D segmentation via hierarchical contrastive learning

    Haiyang Ying, Yixuan Yin, Jinzhi Zhang, Fan Wang, Tao Yu, Ruqi Huang, and Lu Fang. Omniseg3d: Omniversal 3D segmentation via hierarchical contrastive learning. InCVPR, pages 20612–20622. IEEE, 2024. 3

  69. [69]

    Function4d: Real-time human vol- umetric capture from very sparse consumer RGBD sensors

    Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qiong- hai Dai, and Yebin Liu. Function4d: Real-time human vol- umetric capture from very sparse consumer RGBD sensors. InCVPR, pages 5746–5756, 2021. 6

  70. [70]

    3D hu- man body reshaping with anthropometric modeling

    Yanhong Zeng, Jianlong Fu, and Hongyang Chao. 3D hu- man body reshaping with anthropometric modeling. InIn- ternational Conference on Internet Multimedia Computing and Service, pages 96–107. Springer, 2017. 3

  71. [71]

    FATE: full-head Gaussian avatar with textural editing from monoc- ular video

    Jiawei Zhang, Zijian Wu, Zhiyang Liang, Yicheng Gong, Dongfang Hu, Yao Yao, Xun Cao, and Hao Zhu. FATE: full-head Gaussian avatar with textural editing from monoc- ular video. InCVPR, pages 5535–5545. Computer Vision Foundation / IEEE, 2025. 3

  72. [72]

    M3d-vton: A monocular-to-3D virtual try-on network

    Fuwei Zhao, Zhenyu Xie, Michael Kampffmeyer, Haoye Dong, Songfang Han, Tianxiang Zheng, Tao Zhang, and Xi- aodan Liang. M3d-vton: A monocular-to-3D virtual try-on network. InCVPR, pages 13239–13249, 2021. 3

  73. [73]

    Image-based clothes changing system.Com- putational Visual Media, 3(4):337–347, 2017

    Zhao-Heng Zheng, Hao-Tian Zhang, Fang-Lue Zhang, and Tai-Jiang Mu. Image-based clothes changing system.Com- putational Visual Media, 3(4):337–347, 2017. 3

  74. [74]

    Avatarmakeup: Realistic makeup transfer for 3D animatable head avatars.CoRR, abs/2507.02419, 2025

    Yiming Zhong, Xiaolin Zhang, Ligang Liu, Yao Zhao, and Yunchao Wei. Avatarmakeup: Realistic makeup transfer for 3D animatable head avatars.CoRR, abs/2507.02419, 2025. 3

  75. [75]

    Pt-vton: an image-based virtual try-on network with progressive pose attention transfer, 2021

    Hanhan Zhou, Tian Lan, and Guru Venkataramani. Pt-vton: an image-based virtual try-on network with progressive pose attention transfer, 2021. 3

  76. [76]

    Parametric reshaping of human bodies in images.ACM TOG, 29(4):126:1–126:10, 2010

    Shizhe Zhou, Hongbo Fu, Ligang Liu, Daniel Cohen-Or, and Xiaoguang Han. Parametric reshaping of human bodies in images.ACM TOG, 29(4):126:1–126:10, 2010. 3

  77. [77]

    Tryondiffusion: A tale of two unets

    Luyang Zhu, Dawei Yang, Tyler Zhu, Fitsum Reda, William Chan, Chitwan Saharia, Mohammad Norouzi, and Ira Kemelmacher-Shlizerman. Tryondiffusion: A tale of two unets. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 4606–4615,

  78. [78]

    Dagsm: Disentangled avatar generation with gs-enhanced mesh.arXiv preprint arXiv:2411.15205, 2024

    Jingyu Zhuang, Di Kang, Linchao Bao, Liang Lin, and Guanbin Li. Dagsm: Disentangled avatar generation with gs-enhanced mesh.arXiv preprint arXiv:2411.15205, 2024. 3

  79. [79]

    Tip-editor: An accurate 3D editor fol- lowing both text-prompts and image-prompts.ACM TOG, 43(4):121:1–121:12, 2024

    Jingyu Zhuang, Di Kang, Yan-Pei Cao, Guanbin Li, Liang Lin, and Ying Shan. Tip-editor: An accurate 3D editor fol- lowing both text-prompts and image-prompts.ACM TOG, 43(4):121:1–121:12, 2024. 2, 3, 6

  80. [80]

    Driv- able 3D Gaussian avatars, 2025

    Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito, Michael Zollh ¨ofer, Justus Thies, and Javier Romero. Driv- able 3D Gaussian avatars, 2025. 3 11 AvatarMix: Identity-Preserving Cross-Avatar Composition for Outfit Personalization Supplementary Material Zhaorong Wang, Yoshihiro Kanamori, Yuki Endo University of Tsukuba zhaorong.wang1997@gmail.com,{kana...

Showing first 80 references.