<link rel="stylesheet" href="styles.f3b1fba60ec7970c.css">

Self-Supervised Multimodal 3-D Garment Reconstruction from a Single Consumer Image for Energy-Efficient Virtual Try-On Systems

dc.contributor.authorChekhmestruk, R.en
dc.contributor.authorVoitsekhovska, O.en
dc.date.accessioned2026-09-14T12:13:10Z
dc.date.available2026-09-14T12:13:10Z
dc.date.issued2026
dc.description.abstractAccurate 3-D reconstruction of garments a single consumer-grade image remains a critical barrier to truly immersive and resource-aware virtual try-on systems. We introduce a self-supervised, multimodal pipeline that fuses visual tokens extracted by a Vision Transformer with textual garment descriptors to synthesise high-fidelity cloth geometry and texture while operating within the stringent power envelope of mobile neural-processing units (NPUs). A hybrid latent-diffusion module generates pseudo-meshes that supervise a lightweight INT8-quantised Mesh-Autoencoder, thereby eliminating the dependence on large annotated 3-D-scan corpora. To compensate for limited real data we construct SyntheCloth-300K, a dataset blending CLO-3D captures with PhysX-driven synthetic variations, and use it for joint visual–textual training. On the DeepFashion3D benchmark our method reduces Chamfer-Distance by 18% and improves SSIM by 0.03 over DressCode-NeRF, while sustaining 21 FPS at 0.32mJvertex−1 on a Snapdragon 8 Gen 3 — tripling the energy efficiency of prior art. Qualitative results reveal robust reconstruction of fine pleats and fabric drape, even under severe self-occlusion. The proposed framework thus bridges computer vision, physically based graphics, and embedded optimisation, laying the groundwork for next-generation, on-device virtual fitting applications.en
dc.identifier.citationChekhmestruk R., Voitsekhovska O. Self-Supervised Multimodal 3-D Garment Reconstruction a Single Consumer Image for Energy-Efficient Virtual Try-On Systems // Electronic Letters on Computer Vision and Image Analysis. 2026. Vol. 25, іss. 1. P. 60-82. URI: https://elcvia.cvc.uab.cat/article/view/2276.en
dc.identifier.doihttps://doi.org/10.5565/rev/elcvia.2276
dc.identifier.issn5108-5097
dc.identifier.orcidhttps://orcid.org/0000-0002-5362-8796
dc.identifier.urihttps://ir.lib.vntu.edu.ua/handle/123456789/54391
dc.language.isoen_USen_US
dc.publisherUniversitat Autonoma de Barcelonaen
dc.relation.ispartofElectronic Letters on Computer Vision and Image Analysis. Vol. 25, іss. 1 : 60-82.uk
dc.relation.urihttps://elcvia.cvc.uab.cat/article/view/2276
dc.subjectWords: multimodal learningen
dc.subjectself-supervised diffusionen
dc.subjectgarment reconstructionen
dc.subjectenergy-efficient inferenceen
dc.subjectvirtual try-onen
dc.titleSelf-Supervised Multimodal 3-D Garment Reconstruction from a Single Consumer Image for Energy-Efficient Virtual Try-On Systemsen
dc.typeArticle, Scopus-WoS
dc.typeArticle

Файли

Контейнер файлів

Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
208050.pdf
Розмір:
16,59 MB
Формат:
Adobe Portable Document Format

Ліцензійна угода

Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
license.txt
Розмір:
129 B
Формат:
Plain Text
Опис: