REVIEW 2 major objections 44 references
Depth-only person re-identification with temporal transformers and Hungarian matching achieves competitive performance while hiding identifiable features.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-26 08:52 UTC pith:XUHMDAPY
load-bearing objection The paper applies known components to depth Re-ID but supplies no numbers and leaves the privacy claim untested. the 2 major comments →
Privacy-Preserving Person Re-Identification from Temporal Sequences with Transformer and Hungarian Optimization
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that depth images, which obscure facial and identifiable features, can support effective person re-identification when processed as temporal sequences through a Transformer encoder and matched using the Hungarian algorithm for global cost minimization, achieving competitive CMC and mAP scores compared to state-of-the-art methods on top-view datasets.
What carries the argument
A Transformer encoder that processes temporal sequences of depth frames, paired with the Hungarian algorithm that minimizes the global cost in the distance matrix for associating multiple views of individuals.
Load-bearing premise
Depth images inherently obscure facial and other identifiable features sufficiently to constitute a privacy-preserving solution while still enabling effective feature extraction for re-identification across views.
What would settle it
Running the depth-only model on the TVPR2, GODPR, and BIWI RGBD-ID datasets and checking if its CMC and mAP scores are within a small margin of the RGB-based state-of-the-art methods; a large gap would disprove competitive performance.
If this is right
- Depth-only Re-ID provides a privacy-preserving alternative that performs competitively on CMC and mAP metrics.
- Incorporating temporal information via Transformer improves capture of dynamic movement patterns.
- Batch hard triplet loss enhances discriminative features by focusing on hard samples.
- The method works on both depth-only and RGB-D inputs across multiple datasets.
Where Pith is reading between the lines
- This technique could be applied to other privacy-sensitive tracking scenarios such as in retail or healthcare monitoring.
- Combining it with edge devices might enable on-site processing to further reduce data exposure risks.
- Future work might test robustness to different lighting or occlusion levels beyond the evaluated datasets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a privacy-preserving person re-identification method that processes temporal sequences of depth (and optionally RGB) images with a Transformer encoder, uses batch-hard triplet loss for discriminative features, and applies the Hungarian algorithm to solve multi-view association via global cost minimization on a distance matrix. It evaluates depth-only and RGB-D variants on the top-view datasets TVPR2, GODPR, and BIWI RGBD-ID, asserting that depth-only re-identification achieves competitive CMC and mAP scores relative to state-of-the-art methods while inherently preserving privacy by obscuring facial and other identifiable features.
Significance. If the empirical claims hold with verifiable numbers and ablations, the work would provide a concrete demonstration that depth sequences can support competitive Re-ID performance, offering a practical route to privacy-aware surveillance systems. The combination of Transformer temporal modeling with Hungarian matching is a reasonable technical choice for the association problem, and the explicit focus on depth-only evaluation is a strength.
major comments (2)
- [Abstract] Abstract: the central claim that 'depth-only re-identification can achieve competitive performance compared to state-of-the-art methods' is asserted without any numerical CMC, mAP, baseline comparisons, error bars, or ablation results. This absence prevents verification of the primary empirical contribution.
- [Abstract] Abstract (privacy claim): the statement that depth images 'inherently obscures facial and other identifiable features' is presented as sufficient for privacy preservation, yet the manuscript supplies no analysis of whether body shape, height, or gait dynamics extractable from the temporal depth sequences fed to the Transformer remain identifying. All cited datasets are top-view, where even RGB already limits facial visibility, so the incremental privacy benefit is not demonstrated.
Simulated Author's Rebuttal
We thank the referee for the detailed feedback on our manuscript. We address each major comment below and will incorporate revisions to strengthen the abstract and related sections.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that 'depth-only re-identification can achieve competitive performance compared to state-of-the-art methods' is asserted without any numerical CMC, mAP, baseline comparisons, error bars, or ablation results. This absence prevents verification of the primary empirical contribution.
Authors: We agree that the abstract would be strengthened by including quantitative results. In the revised manuscript, we will update the abstract to report key CMC and mAP scores for the depth-only model on TVPR2, GODPR, and BIWI RGBD-ID, along with comparisons to relevant baselines from the literature. This will directly support the claim of competitive performance. revision: yes
-
Referee: [Abstract] Abstract (privacy claim): the statement that depth images 'inherently obscures facial and other identifiable features' is presented as sufficient for privacy preservation, yet the manuscript supplies no analysis of whether body shape, height, or gait dynamics extractable from the temporal depth sequences fed to the Transformer remain identifying. All cited datasets are top-view, where even RGB already limits facial visibility, so the incremental privacy benefit is not demonstrated.
Authors: We acknowledge that the current abstract does not provide a detailed privacy analysis. While depth inherently excludes color and texture cues, we recognize that shape and gait information could remain. In revision, we will expand the abstract and add a short discussion paragraph clarifying the privacy advantages in top-view settings and noting potential residual identifiers, to better demonstrate the incremental benefit. revision: yes
Circularity Check
No circularity; empirical evaluation on public datasets with standard components
full rationale
The paper presents an empirical method using Transformer on temporal depth sequences, Hungarian matching, and batch-hard triplet loss, evaluated via CMC and mAP on TVPR2, GODPR, and BIWI RGBD-ID. No derivation chain, fitted parameters renamed as predictions, or self-citation load-bearing steps appear in the provided text. Claims rest on external benchmark results rather than self-referential definitions or ansatzes imported from prior author work. The privacy assertion is an unverified modeling assumption, not a circular reduction in the derivation.
Axiom & Free-Parameter Ledger
read the original abstract
Person re-identification (Re-ID) is a crucial task in surveillance and human behavior analysis, often used in public spaces such as transport hubs. Traditional RGB-based Re-ID methods raise privacy concerns and are highly sensitive to lighting variations and occlusion. In this paper, we propose a novel Re-ID approach that leverages depth images, which inherently obscures facial and other identifiable features, making it a privacy-preserving solution. Our method addresses the association problem between multiple views of individuals by applying the Hungarian algorithm, optimizing the matching process through minimization of the global cost across the distance matrix. We further enhance the approach by introducing temporal sequences of frames as input to a Transformer encoder architecture, which exploits both RGB and depth modalities. This architecture captures dynamic movement patterns, improving feature extraction and re-identification accuracy. Additionally, we employ batch hard triplet loss to enhance discriminative feature learning by focusing on the hardest samples. We evaluate both depth-only and RGB-D models on several top-view datasets, including TVPR2, GODPR, and BIWI RGBD-ID. Our results demonstrate that depth-only re-identification can achieve competitive performance compared to state-of-the-art methods, as measured by standard metrics such as Cumulative Matching Characteristics (CMC) and Mean Average Precision (mAP), while prioritizing privacy preservation.
Figures
Reference graph
Works this paper leans on
-
[1]
Ahmed, M
E. Ahmed, M. Jones, and T. K. Marks. An improved deep learning architecture for person re-identification. In2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3908–3916, June 2015. ISSN: 1063-6919
2015
-
[2]
P. P. Busto and J. Gall. Open Set Domain Adaptation. In2017 IEEE International Conference on Computer Vision (ICCV), pages 754–763, Venice, Oct. 2017. IEEE
2017
-
[3]
G. Chen, C. Lin, L. Ren, J. Lu, and J. Zhou. Self-Critical Attention Learning for Person Re-Identification. In2019 IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 9636–9645, Seoul, Korea (South), Oct. 2019. IEEE
2019
-
[4]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, June
-
[5]
Fuentes-Jimenez, C
D. Fuentes-Jimenez, C. L. Gutierrez, J. M. Guarasa, C. Luna, and D. Pizarro. Depth Person detection database (GFPD), 2020
2020
-
[6]
Gong and T
S. Gong and T. Xiang. Person Re-identification. In S. Gong and T. Xiang, editors,Visual Analysis of Behaviour: From Pixels to Semantics, pages 301–313. Springer, London, 2011
2011
-
[7]
F. M. Hafner, A. Bhuyian, J. F. P. Kooij, and E. Granger. Cross-modal distillation for RGB-depth person re-identification.Computer Vision and Image Understanding, 216:103352, Feb. 2022
2022
-
[8]
K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, Las Vegas, NV , USA, June 2016. IEEE
2016
-
[9]
In Defense of the Triplet Loss for Person Re-Identification
A. Hermans, L. Beyer, and B. Leibe. In Defense of the Triplet Loss for Person Re-Identification, Nov. 2017. arXiv:1703.07737 [cs]
work page internal anchor Pith review Pith/arXiv arXiv 2017
- [10]
-
[11]
H. W. Kuhn. The Hungarian method for the assignment problem. Naval Research Logistics Quarterly, 2(1-2):83–97, 1955. eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/nav.3800020109
-
[12]
A. R. Lejbolle, K. Nasrollahi, B. Krogh, and T. B. Moeslund. Multimodal Neural Network for Overhead Person Re-Identification. In2017 International Conference of the Biometrics Special Interest Group (BIOSIG), pages 1–5, Darmstadt, Germany, Sept. 2017. IEEE
2017
-
[13]
A. R. Lejbølle, B. Krogh, K. Nasrollahi, and T. B. Moeslund. Attention in Multimodal Neural Networks for Person Re-identification. In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 292–2928, June 2018. ISSN: 2160-7516
2018
-
[14]
A. R. Lejbølle, K. Nasrollahi, B. Krogh, and T. B. Moeslund. Person Re-Identification Using Spatial and Layer-Wise Attention.IEEE Transactions on Information Forensics and Security, 15:1216–1231,
-
[15]
Conference Name: IEEE Transactions on Information Forensics and Security
-
[16]
Q. Leng, M. Ye, and Q. Tian. A Survey of Open-World Person Re- Identification.IEEE Transactions on Circuits and Systems for Video Technology, 30(4):1092–1108, Apr. 2020. Conference Name: IEEE Transactions on Circuits and Systems for Video Technology
2020
-
[17]
W. Li, R. Zhao, T. Xiao, and X. Wang. DeepReID: Deep Filter Pairing Neural Network for Person Re-identification. In2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 152– 159, Columbus, OH, USA, June 2014. IEEE
2014
-
[18]
Liciotti, M
D. Liciotti, M. Paolanti, E. Frontoni, A. Mancini, and P. Zingaretti. Person Re-identification Dataset with RGB-D Camera in a Top-View Configuration. In K. Nasrollahi, C. Distante, G. Hua, A. Cavallaro, T. B. Moeslund, S. Battiato, and Q. Ji, editors,Video Analytics. Face and Facial Expression Recognition and Audience Measurement, pages 1–11, Cham, 2017. ...
2017
-
[19]
C. A. Luna, C. Losada-Guti ´errez, D. Fuentes-Jimenez, and M. Mazo. People re-identification using depth and intensity information from an overhead camera.Expert Systems with Applications, 182:115287, Nov. 2021
2021
-
[20]
Martini, M
M. Martini, M. Paolanti, and E. Frontoni. Open-World Person Re- Identification With RGBD Camera in Top-View Configuration for Retail Applications.IEEE Access, 8:67756–67765, 2020. Conference Name: IEEE Access
2020
-
[21]
Mukhtar and M
H. Mukhtar and M. U. G. Khan. CMOT: A cross-modality transformer for RGB-D fusion in person re-identification with online learning capabilities.Knowledge-Based Systems, 283:111155, Jan. 2024
2024
-
[22]
Munaro, A
M. Munaro, A. Fossati, A. Basso, E. Menegatti, and L. Van Gool. One-Shot Person Re-identification with a Consumer Depth Camera. In S. Gong, M. Cristani, S. Yan, and C. C. Loy, editors,Person Re- Identification, pages 161–181. Springer, London, 2014
2014
-
[23]
F. Pala, R. Satta, G. Fumera, and F. Roli. Multimodal Person Reidentification Using RGB-D Cameras.IEEE Transactions on Circuits and Systems for Video Technology, 26(4):788–799, Apr. 2016. Conference Name: IEEE Transactions on Circuits and Systems for Video Technology
2016
-
[24]
Paolanti, R
M. Paolanti, R. Pierdicca, R. Pietrini, M. Martini, and E. Frontoni. SeSAME: Re-identification-based ambient intelligence system for museum environment.Pattern Recognition Letters, 161:17–23, Sept. 2022
2022
-
[25]
Paolanti, R
M. Paolanti, R. Pietrini, A. Mancini, E. Frontoni, and P. Zingaretti. Deep understanding of shopper behaviours and interactions using RGB-D vision.Machine Vision and Applications, 31(7):66, Sept. 2020
2020
-
[26]
Paolanti, L
M. Paolanti, L. Romeo, D. Liciotti, R. Pietrini, A. Cenci, E. Frontoni, and P. Zingaretti. Person Re-Identification with RGB-D Camera in Top-View Configuration through Multiple Nearest Neighbor Clas- sifiers and Neighborhood Component Features Selection.Sensors, 18(10):3471, Oct. 2018. Number: 10 Publisher: Multidisciplinary Digital Publishing Institute
2018
-
[27]
H. Rao, C. Leung, and C. Miao. Hierarchical Skeleton Meta-Prototype Contrastive Learning with Hard Skeleton Mining for Unsupervised Person Re-identification.International Journal of Computer Vision, 132(1):238–260, Jan. 2024
2024
- [28]
-
[29]
Rao and C
H. Rao and C. Miao. TranSG: Transformer-Based Skeleton Graph Prototype Contrastive Learning with Structure-Trajectory Prompted Reconstruction for Person Re-Identification. In2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22118–22128, Vancouver, BC, Canada, June 2023. IEEE
2023
-
[30]
L. Ren, J. Lu, J. Feng, and J. Zhou. Multi-modal uniform deep learning for RGB-D person re-identification.Pattern Recognition, 72:446–457, Dec. 2017
2017
-
[31]
Ristani, F
E. Ristani, F. Solera, R. Zou, R. Cucchiara, and C. Tomasi. Per- formance Measures and a Data Set for Multi-target, Multi-camera Tracking. In G. Hua and H. J ´egou, editors,Computer Vision – ECCV 2016 Workshops, pages 17–35, Cham, 2016. Springer International Publishing
2016
-
[32]
Schroff, D
F. Schroff, D. Kalenichenko, and J. Philbin. FaceNet: A unified em- bedding for face recognition and clustering. In2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 815–823, Boston, MA, USA, June 2015. IEEE
2015
-
[33]
C. Si, Y . Jing, W. Wang, L. Wang, and T. Tan. Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning. In V . Ferrari, M. Hebert, C. Sminchisescu, and Y . Weiss, editors, Computer Vision – ECCV 2018, volume 11205, pages 106–121. Springer International Publishing, Cham, 2018. Series Title: Lecture Notes in Computer Science
2018
-
[34]
Y . Sun, L. Zheng, Y . Yang, Q. Tian, and S. Wang. Beyond Part Models: Person Retrieval with Refined Part Pooling (and A Strong Convolutional Baseline). In V . Ferrari, M. Hebert, C. Sminchisescu, and Y . Weiss, editors,Computer Vision – ECCV 2018, volume 11208, pages 501–518. Springer International Publishing, Cham, 2018. Series Title: Lecture Notes in C...
2018
-
[35]
Szegedy, S
C. Szegedy, S. Ioffe, V . Vanhoucke, and A. Alemi. Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learn- ing.Proceedings of the AAAI Conference on Artificial Intelligence, 31(1), Feb. 2017. Number: 1
2017
-
[36]
M. K. Uddin, A. Lam, H. Fukuda, Y . Kobayashi, and Y . Kuno. Depth Guided Attention for Person Re-identification. In D.-S. Huang and P. Premaratne, editors,Intelligent Computing Methodologies, pages 110–120, Cham, 2020. Springer International Publishing
2020
-
[37]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is All you Need. In I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vish- wanathan, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017
2017
-
[38]
Wu, W.-S
A. Wu, W.-S. Zheng, and J.-H. Lai. Robust Depth-Based Person Re- Identification.IEEE Transactions on Image Processing, 26(6):2588– 2603, June 2017. Conference Name: IEEE Transactions on Image Processing
2017
-
[39]
Wu, W.-S
A. Wu, W.-S. Zheng, H.-X. Yu, S. Gong, and J. Lai. RGB-Infrared Cross-Modality Person Re-identification. In2017 IEEE International Conference on Computer Vision (ICCV), pages 5390–5399, Venice, Oct. 2017. IEEE
2017
-
[40]
J. Wu, J. Jiang, M. Qi, C. Chen, and J. Zhang. An End-to-end Heterogeneous Restraint Network for RGB-D Cross-modal Person Re-identification.ACM Trans. Multimedia Comput. Commun. Appl., 18(4):109:1–109:22, Mar. 2022
2022
-
[41]
M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. H. Hoi. Deep Learning for Person Re-Identification: A Survey and Outlook. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6):2872–2893, June 2022. Conference Name: IEEE Transactions on Pattern Analysis and Machine Intelligence
2022
-
[42]
C. Zhao, X. Lv, Z. Zhang, W. Zuo, J. Wu, and D. Miao. Deep Fusion Feature Representation Learning With Hard Mining Center-Triplet Loss for Person Re-Identification.IEEE Transactions on Multimedia, 22(12):3180–3195, Dec. 2020. Conference Name: IEEE Transactions on Multimedia
2020
-
[43]
Zheng, L
L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian. Scalable Person Re-identification: A Benchmark. In2015 IEEE International Conference on Computer Vision (ICCV), pages 1116–1124, Santiago, Chile, Dec. 2015. IEEE
2015
-
[44]
Zheng, L
Z. Zheng, L. Zheng, and Y . Yang. Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in Vitro. In2017 IEEE International Conference on Computer Vision (ICCV), pages 3774–3782, Venice, Oct. 2017. IEEE
2017
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.