dorsal/arxiv
View SchemaDisentangled Concept Representation for Text-to-image Person Re-identification
| Authors | Giyeol Kim, Chanho Eom |
|---|---|
| Categories | |
| ArXiv ID | 2601.10053vv1 |
| URL | https://arxiv.org/abs/2601.10053 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Text-to-image person re-identification (TIReID) aims to retrieve person images from a large gallery given free-form textual descriptions. TIReID is challenging due to the substantial modality gap between visual appearances and textual expressions, as well as the need to model fine-grained correspondences that distinguish individuals with similar attributes such as clothing color, texture, or outfit style. To address these issues, we propose DiCo (Disentangled Concept Representation), a novel framework that achieves hierarchical and disentangled cross-modal alignment. DiCo introduces a shared slot-based representation, where each slot acts as a part-level anchor across modalities and is further decomposed into multiple concept blocks. This design enables the disentanglement of complementary attributes (\textit{e.g.}, color, texture, shape) while maintaining consistent part-level correspondence between image and text. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid demonstrate that our framework achieves competitive performance with state-of-the-art methods, while also enhancing interpretability through explicit slot- and block-level representations for more fine-grained retrieval results.
{
"annotation_id": "a1594dfb-ec10-49ff-9162-69f09058ea6d",
"date_created": "2026-02-17T05:53:23.797000Z",
"date_modified": "2026-02-17T05:53:23.797000Z",
"file_hash": "3965ccf8ab623abbe1548e113015e4eb38ef44458e994a235baf977cf08d4a59",
"private": false,
"record": {
"abstract": "Text-to-image person re-identification (TIReID) aims to retrieve person images from a large gallery given free-form textual descriptions. TIReID is challenging due to the substantial modality gap between visual appearances and textual expressions, as well as the need to model fine-grained correspondences that distinguish individuals with similar attributes such as clothing color, texture, or outfit style. To address these issues, we propose DiCo (Disentangled Concept Representation), a novel framework that achieves hierarchical and disentangled cross-modal alignment. DiCo introduces a shared slot-based representation, where each slot acts as a part-level anchor across modalities and is further decomposed into multiple concept blocks. This design enables the disentanglement of complementary attributes (\\textit{e.g.}, color, texture, shape) while maintaining consistent part-level correspondence between image and text. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid demonstrate that our framework achieves competitive performance with state-of-the-art methods, while also enhancing interpretability through explicit slot- and block-level representations for more fine-grained retrieval results.",
"arxiv_id": "2601.10053",
"authors": [
"Giyeol Kim",
"Chanho Eom"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Disentangled Concept Representation for Text-to-image Person Re-identification",
"url": "https://arxiv.org/abs/2601.10053",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "2bb8cf2c-9755-48f7-8374-a5ac100fe144",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}