| dc.contributor.author | Chekhmestruk, R. | en |
| dc.contributor.author | Voitsekhovska, О. | en |
| dc.contributor.author | Чехместрук, Р. Ю. | en |
| dc.contributor.author | Войцеховська, О. В. | en |
| dc.date.accessioned | 2026-09-14T12:13:10Z | |
| dc.date.available | 2026-09-14T12:13:10Z | |
| dc.date.issued | 2026 | en |
| dc.identifier.citation | Chekhmestruk R., Voitsekhovska O. Self-Supervised Multimodal 3-D Garment Reconstruction a Single Consumer Image for Energy-Efficient Virtual Try-On Systems // Electronic Letters on Computer Vision and Image Analysis. 2026. Vol. 25, іss. 1. P. 60-82. URI: https://elcvia.cvc.uab.cat/article/view/2276. | en |
| dc.identifier.issn | 5108-5097 | en |
| dc.identifier.uri | https://ir.lib.vntu.edu.ua/handle/123456789/54391 | |
| dc.description.abstract | Accurate 3-D reconstruction of garments a single consumer-grade image remains a critical barrier to truly immersive and resource-aware virtual try-on systems. We introduce a self-supervised, multimodal pipeline that fuses visual tokens extracted by a Vision Transformer with textual garment descriptors to synthesise high-fidelity cloth geometry and texture while operating within the stringent power envelope of mobile neural-processing units (NPUs). A hybrid latent-diffusion module generates pseudo-meshes that supervise a lightweight INT8-quantised Mesh-Autoencoder, thereby eliminating the dependence on large annotated 3-D-scan corpora. To compensate for limited real data we construct SyntheCloth-300K, a dataset blending CLO-3D captures with PhysX-driven synthetic variations, and use it for joint visual–textual training. On the DeepFashion3D benchmark our method reduces Chamfer-Distance by 18% and improves SSIM by 0.03 over DressCode-NeRF, while sustaining 21 FPS at 0.32mJvertex−1 on a Snapdragon 8 Gen 3 — tripling the energy efficiency of prior art. Qualitative results reveal robust reconstruction of fine pleats and fabric drape, even under severe self-occlusion. The proposed framework thus bridges computer vision, physically based graphics, and embedded optimisation, laying the groundwork for next-generation, on-device virtual fitting applications. | uk_UA |
| dc.language.iso | en_US | en_US |
| dc.publisher | Universitat Autonoma de Barcelona | en |
| dc.relation.ispartof | Electronic Letters on Computer Vision and Image Analysis. Vol. 25, іss. 1 : 60-82. | en |
| dc.subject | Words: multimodal learning; self-supervised diffusion; garment reconstruction; energy-efficient inference; virtual try-on | en |
| dc.title | Self-Supervised Multimodal 3-D Garment Reconstruction from a Single Consumer Image for Energy-Efficient Virtual Try-On Systems | en |
| dc.type | Article, Scopus-WoS | |
| dc.relation.references | https://elcvia.cvc.uab.cat/article/view/2276 | en |
| dc.identifier.doi | https://doi.org/10.5565/rev/elcvia.2276 | en |
| dc.identifier.orcid | https://orcid.org/0000-0002-5362-8796 | en |