Pith. sign in

REVIEW 2 major objections 2 minor 58 references

HyperVision provides the first pre-trained backbone for ground-based hyperspectral images that unifies varying sensor channels and learns representations without labels.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 19:37 UTC pith:M6JRDK2Z

load-bearing objection HyperVision is the first claimed pre-trained backbone for ground-based hyperspectral imaging via channel-adaptive embedding and SAM2/HyperFree pseudo-labeling plus RGB distillation, but the reported gains rest on unverified experimental details. the 2 major comments →

arxiv 2605.17286 v2 pith:M6JRDK2Z submitted 2026-05-17 cs.CV

HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

classification cs.CV
keywords hyperspectral imagingpre-trained backbonechannel-adaptive embeddingunsupervised representation learningsemantic segmentationobject trackingsalient object detectioncross-modal distillation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper claims that hyperspectral imaging can move from task-specific models to a single reusable backbone by solving three barriers at once: mismatched spectral bands across sensors, scarce and inconsistent labels, and small dataset scale. It does so through a channel-adaptive embedding that turns any input into a shared token space plus an unsupervised pre-training stage that fuses spatial cues from one model, spectral material cues from another, and semantic knowledge distilled from an RGB model. If this holds, a backbone trained once on 15k images can be dropped into new sensors and new tasks by training only a small head, delivering measurable gains on segmentation, tracking, and detection.

Core claim

HyperVision is the first ground-based hyperspectral pre-trained backbone. It maps heterogeneous spectral inputs into a unified space with a channel-adaptive dynamic embedding mechanism. It then learns via an unsupervised framework that creates pseudo-labels by combining spatial structure from SAM2 with fine-grained spectral material information from HyperFree and transfers additional semantics through cross-modal distillation from a pre-trained RGB model. After training on 15k images from 26 datasets, only the task head needs updating; the backbone stays frozen yet outperforms prior task-specific methods on semantic segmentation, object tracking, and salient object detection under different

What carries the argument

Channel-adaptive dynamic embedding mechanism that projects arbitrary numbers of spectral bands into a common token space, paired with the multi-source pseudo-labeling and cross-modal distillation pipeline that supplies training signals.

Load-bearing premise

The fused pseudo-labels from spatial and spectral sources plus the RGB distillation produce training signals consistent and rich enough to learn representations that generalize across sensors and tasks.

What would settle it

Train the same architecture on a new collection of hyperspectral images that use spectral bands absent from the original 26 datasets; if head-only adaptation no longer exceeds task-specific baselines on segmentation accuracy, the claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Head-only adaptation suffices for state-of-the-art results on hyperspectral semantic segmentation, object tracking, and salient object detection.
  • The backbone generalizes across sensor configurations without any parameter updates to its core layers.
  • Pre-training on combined data from 26 datasets compensates for the small scale and limited diversity of individual hyperspectral collections.
  • Unsupervised signals can replace manual labels for learning generalizable spatial-spectral features.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same channel-adaptive idea could be tested on satellite or airborne hyperspectral data where band sets also differ between instruments.
  • If the pseudo-label quality holds, the method offers a route to pre-train backbones for other label-scarce modalities such as thermal or multispectral video.
  • Real-world deployment becomes simpler when a single backbone can be reused across hardware generations without full retraining.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript introduces HyperVision as the first ground-based hyperspectral pre-trained backbone. It addresses varying sensor configurations via a channel-adaptive dynamic embedding, and tackles label scarcity through an unsupervised framework that fuses spatial structure from SAM2 with spectral information from HyperFree, augmented by cross-modal distillation from a pre-trained RGB model. Pre-trained on 15k images from 26 datasets, the model claims state-of-the-art results on semantic segmentation, object tracking, and salient object detection using only head-only adaptation, with reported relative gains of 16.3% in Acc_M, 2.1% in AUC, and 35.5% MAE reduction.

Significance. If the reported gains prove robust under standard experimental controls, the work would provide the first general-purpose pre-trained backbone for ground-based hyperspectral imaging, directly addressing sensor heterogeneity and data scarcity that have limited progress in the field. The explicit commitment to public release of code and pre-trained weights on GitHub is a clear strength that supports reproducibility and follow-on research.

major comments (2)
  1. [unsupervised representation learning framework description] The central empirical claims depend on the quality and consistency of the multi-source pseudo-labels; the manuscript provides no quantitative validation (e.g., agreement with held-out ground truth or inter-source consistency metrics) for the fusion of SAM2 and HyperFree outputs across the 26 datasets used in pre-training.
  2. [experimental results section] No information is supplied on experimental protocols, baseline implementations, statistical testing, error bars, or dataset splits for the three downstream tasks; without these, it is not possible to assess whether the reported relative improvements are statistically reliable or comparable to prior work.
minor comments (2)
  1. [abstract and results] Notation for Acc_M, AUC, and MAE should be defined at first use with explicit reference to the evaluation metrics employed in each task.
  2. [channel-adaptive dynamic embedding mechanism] The channel-adaptive embedding weights are listed among free parameters; clarify whether these are learned during pre-training or fixed after initialization.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback on our manuscript. We address each major comment below and commit to revisions that enhance the clarity and rigor of the work without altering its core contributions.

read point-by-point responses
  1. Referee: [unsupervised representation learning framework description] The central empirical claims depend on the quality and consistency of the multi-source pseudo-labels; the manuscript provides no quantitative validation (e.g., agreement with held-out ground truth or inter-source consistency metrics) for the fusion of SAM2 and HyperFree outputs across the 26 datasets used in pre-training.

    Authors: We agree that explicit quantitative validation of the pseudo-label fusion would improve transparency. While downstream task results provide indirect evidence of label quality, we will add in the revision: (i) inter-source consistency metrics (e.g., IoU between SAM2 and HyperFree masks on overlapping regions) computed across the 26 datasets, and (ii) agreement scores against any available held-out ground truth subsets. These additions will be placed in a new subsection under the pre-training framework description. revision: yes

  2. Referee: [experimental results section] No information is supplied on experimental protocols, baseline implementations, statistical testing, error bars, or dataset splits for the three downstream tasks; without these, it is not possible to assess whether the reported relative improvements are statistically reliable or comparable to prior work.

    Authors: We acknowledge the omission of these details in the submitted manuscript. In the revised version we will expand the experimental section to include: full dataset split specifications, exact baseline re-implementation details (including hyper-parameters), statistical significance tests (e.g., paired t-tests), and error bars (standard deviation over multiple runs). This will allow direct assessment of reliability and comparability. revision: yes

Circularity Check

0 steps flagged

No significant circularity

full rationale

The paper presents an empirical architecture (channel-adaptive embedding) and unsupervised pretraining pipeline (multi-source pseudo-labeling from SAM2 + HyperFree plus RGB distillation) whose outputs are measured by downstream task performance after head-only fine-tuning. These results are obtained from training on 15k images and evaluating on separate task-specific benchmarks; they do not reduce by any equation or definition in the paper to quantities that are tautologically equivalent to the inputs or to self-citations. No self-definitional, fitted-input, or uniqueness-theorem circularity is present.

Axiom & Free-Parameter Ledger

2 free parameters · 2 axioms · 0 invented entities

The central claim depends on the quality of the generated pseudo-labels and the effectiveness of the distillation transfer; these rest on domain assumptions rather than external benchmarks or formal derivations. Model parameters are fitted during pre-training on the collected 15k images, but no additional ad-hoc constants are described in the abstract.

free parameters (2)
  • channel-adaptive embedding weights
    Learned parameters that map variable-length spectral inputs into a fixed token space during pre-training.
  • pseudo-label fusion parameters
    Any thresholds, weights, or selection rules used to combine SAM2 and HyperFree outputs.
axioms (2)
  • domain assumption Pseudo-labels produced by fusing SAM2 spatial cues and HyperFree spectral cues are sufficiently accurate and unbiased for representation learning.
    Invoked to address label scarcity and inconsistency in the unsupervised framework.
  • domain assumption Semantic representations transferred from a pre-trained RGB vision model remain useful when applied to hyperspectral inputs.
    Invoked to compensate for limited scene diversity.

pith-pipeline@v0.9.1-grok · 5865 in / 1574 out tokens · 33783 ms · 2026-06-30T19:37:05.027383+00:00 · methodology

0 comments
read the original abstract

While hyperspectral imaging provides rich spatial-spectral information across hundreds of narrow wavelength bands for precise material identification, ground-based hyperspectral pre-trained backbones remain absent, constrained by varying spectral configurations across sensors, the scarcity and inconsistency of labels, and the limited scale and scene diversity of existing datasets. To address these challenges and enable universal perception, we propose HyperVision, the first ground-based hyperspectral pre-trained backbone. First, to handle varying spectral configurations, HyperVision adopts a channel-adaptive dynamic embedding mechanism to map heterogeneous inputs into a unified token space. Second, we develop an unsupervised representation learning framework. Specifically, to address label scarcity and inconsistency, a multi-source pseudo-labeling method is introduced to fuse spatial structures from SAM2 and fine-grained spectral material information from HyperFree. Furthermore, to enrich scene diversity and compensate for limited dataset scale, a cross-modal knowledge distillation mechanism is utilized to transfer rich semantic representations from a pre-trained RGB vision model to our backbone. Pre-trained on a collection of 15k images from 26 diverse ground-based datasets, HyperVision demonstrates exceptional generalization. Requiring only efficient head-only adaptation without adjusting backbone parameters, it achieves state-of-the-art performance compared to task-specific methods across three downstream tasks under varying sensor configurations, yielding up to a 16.3% relative improvement in hyperspectral semantic segmentation $\mathrm{Acc}_{\mathrm{M}}$, a 2.1% relative gain in object tracking AUC, and a 35.5% reduction in salient object detection MAE. The source code and pre-trained model will be publicly available on https://github.com/lronkitty/HyperVision .

Figures

Figures reproduced from arXiv: 2605.17286 by Chengrong Chen, Diqi Chen, Fengchao Xiong, Guanyiman Fu, Jianfeng Lu, Jingtao Li, Jun Zhou, Xiangyu Liu, Yan Xu, Zhuanfeng Li, Zihang Cheng.

Figure 1
Figure 1. Figure 1: Comparison of airborne and ground-based HSI modeling. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of existing HSI modeling using pre-trained models for downstream [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The architecture of HyperVision. hyperspectral tasks [29, 48], we extend the embedding stage with a two-branch design that processes dynamically assembled layers in parallel. Specifically, X is patchified in step p and split into a key-channel component Xk and an intermediate-cube component Xc. Denoting the dictionary-based weight construction as g(b,β), we use two separate dictionaries βk and βc to proces… view at source ↗
Figure 4
Figure 4. Figure 4: Unsupervised representation learning with HyperVision. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Visualization of pseudo-masks generated by SAM2 and HyperFree. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Illustration of the prompt-driven segmentation pipeline. [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Visual comparisons of hyperspectral semantic segmentation results. [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 7
Figure 7. Figure 7: Visual comparisons of hyperspectral semantic segmentation results. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Visual comparisons of hyperspectral object tracking results. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Visual comparisons of salient object detection results. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 9
Figure 9. Figure 9: Visual comparisons of salient object detection results. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: t-SNE visualization comparing the feature representations of HyperFree and the [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 2 canonical work pages

  1. [1]

    Sparse Recovery of Hyperspectral Signal from Nat- ural RGB Images

    Boaz Arad and Ohad Ben-Shahar. Sparse Recovery of Hyperspectral Signal from Nat- ural RGB Images. InProc. Eur . Conf. Comput. Vis. (ECCV), pages 19–34, 2016. ISBN 978-3-319-46478-7

  2. [2]

    Mohamed Mansoor Roomi

    Boaz Arad, Radu Timofte, Rony Yahel, Nimrod Morag, Amir Bernat, Yuanhao Cai, Jing Lin, Zudi Lin, Haoqian Wang, Yulun Zhang, Hanspeter Pfister, Luc Van Gool, Shuai Liu, Yongqiang Li, Chaoyu Feng, Lei Lei, Jiaojiao Li, Songcheng Du, Chaox- iong Wu, Yihong Leng, Rui Song, Mingwei Zhang, Chongxing Song, Shuyi Zhao, Zhiqiang Lang, Wei Wei, Lei Zhang, Renwei Di...

  3. [3]

    NTIRE 2022 Spectral Demosaicing Challenge and Data Set

    Boaz Arad, Radu Timofte, Rony Yahel, Nimrod Morag, Amir Bernat, Yaqi Wu, Xun Wu, Zhihao Fan, Chenjie Xia, Feng Zhang, Shuai Liu, Yongqiang Li, Chaoyu Feng, Lei Lei, Mingwei Zhang, Kai Feng, Xun Zhang, Jiaxin Yao, Yongqiang Zhao, Suina Ma, Fan He, Yangyang Dong, Shufang Yu, Difa Qiu, Jinhui Liu, Mengzhao Bi, Beibei Song, WenFang Sun, Jiesi Zheng, Bowen Zha...

  4. [4]

    Labeled Hyperspectral and RGB Images of Several Tree Species

    Ryan Brown and Josh Moser. Labeled Hyperspectral and RGB Images of Several Tree Species. 2021

  5. [5]

    SAM 3: Segment anything with concepts

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu, Tsung-Han Wu, Yu Zhou, Liliane ...

  6. [6]

    Statistics of real-world hyperspectral images

    Ayan Chakrabarti and Todd Zickler. Statistics of real-world hyperspectral images. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 193–200, 2011

  7. [7]

    Encoder-Decoder with Atrous Separable Convolution for Semantic Image Seg- mentation

    Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Seg- mentation. InProc. Eur . Conf. Comput. Vis. (ECCV), 2018

  8. [8]

    SENSE: Hyperspectral video object tracker via fusing material and motion cues.Inf

    Yuzeng Chen, Qiangqiang Yuan, Yuqi Tang, Yi Xiao, Jiang He, and Zhenqi Liu. SENSE: Hyperspectral video object tracker via fusing material and motion cues.Inf. Fusion, 109:102395, 2024. ISSN 1566-2535. 16FU ET AL.: HYPERVISION

  9. [9]

    Foster and Adam Reeves

    David H. Foster and Adam Reeves. Colour constancy failures expected in colourful environments. InProc. R. Soc. B Biol. Sci., volume 289, page 20212483, 2022

  10. [10]

    Visible – Near infrared hyperspectral dataset of healthy and in- fected apple tree leaves images for the monitoring of apple fire blight.Data in Brief, 50:109532, 2023

    Belal Gaci, Florent Abdelghafour, Maxime Ryckewaert, Silvia Mas-Garcia, Marine Louargant, Florence Verpont, Yohana Laloum, Aude Moronvalle, Ryad Bendoula, and Jean-Michel Roger. Visible – Near infrared hyperspectral dataset of healthy and in- fected apple tree leaves images for the monitoring of apple fire blight.Data in Brief, 50:109532, 2023. ISSN 2352-3409

  11. [11]

    Gaidel, V .V

    A.V . Gaidel, V .V . Podlipnov, Ivliev Nikolay, R.A. Paringer, P.A. Ishkin, S.V . Mashkov, and R.V . Skidanov. Agricultural plant hyperspectral imaging dataset.Comput. Opt., 47, 2023

  12. [12]

    CBFF- Net: A New Framework for Efficient and Accurate Hyperspectral Object Tracking

    Long Gao, Pan Liu, Yan Jiang, Weiying Xie, Jie Lei, Yunsong Li, and Qian Du. CBFF- Net: A New Framework for Efficient and Accurate Hyperspectral Object Tracking. IEEE Trans. Geosci. Remote Sens., 61:1–14, 2023

  13. [13]

    Victoria Martínez, and Unai Martinez-Corral

    Jon Gutiérrez-Zaballa, Koldo Basterretxea, Javier Echanobe, M. Victoria Martínez, and Unai Martinez-Corral. HSI-Drive v2.0: More Data for New Challenges in Scene Un- derstanding for Autonomous Driving. InProc. IEEE Symp. Ser . Comput. Intell. (SSCI), pages 207–214, 2023

  14. [14]

    A Hyperspectral and RGB Dataset for Build- ing Façade Segmentation

    Nariman Habili, Ernest Kwan, Weihao Li, Christfried Webers, Jeremy Oorloff, Mo- hammad Ali Armin, and Lars Petersson. A Hyperspectral and RGB Dataset for Build- ing Façade Segmentation. InProc. Eur . Conf. Comput. Vis. (ECCV) Workshops, pages 258–267, 2023. ISBN 978-3-031-25082-8

  15. [15]

    Hyper-Drive: Visible-Short Wave Infrared Hyperspectral Imaging Datasets for Robots in Unstructured Environments

    Nathaniel Hanson, Benjamin Pyatski, Samuel Hibbard, Charles DiMarzio, and Ta¸ skın Padır. Hyper-Drive: Visible-Short Wave Infrared Hyperspectral Imaging Datasets for Robots in Unstructured Environments. InProc. Workshop Hyperspectral Imaging Sig- nal Process.: Evolution Remote Sens. (WHISPERS), pages 1–5, 2023

  16. [16]

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. InProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016

  17. [17]

    Masked Autoencoders Are Scalable Vision Learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked Autoencoders Are Scalable Vision Learners. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), pages 16000–16009, 2022

  18. [18]

    SpectralGPT: Spectral Remote Sensing Foun- dation Model.IEEE Trans

    Danfeng Hong, Bing Zhang, Xuyang Li, Yuxuan Li, Chenyu Li, Jing Yao, Naoto Yokoya, Hao Li, Pedram Ghamisi, Xiuping Jia, Antonio Plaza, Paolo Gamba, Jon Atli Benediktsson, and Jocelyn Chanussot. SpectralGPT: Spectral Remote Sensing Foun- dation Model.IEEE Trans. Pattern Anal. Mach. Intell., 46(8):5227–5244, 2024

  19. [19]

    Spectral simulation and method design of camouflage textiles for concealment of hyperspectral imaging in UV-VIS-IR against multidimensional combat background.J

    Anowar Hossain. Spectral simulation and method design of camouflage textiles for concealment of hyperspectral imaging in UV-VIS-IR against multidimensional combat background.J. Text. Inst., 114(2):331–342, 2023

  20. [20]

    Spatial–Spectral Weighted and Regu- larized Tensor Sparse Correlation Filter for Object Tracking in Hyperspectral Videos

    Zengfu Hou, Wei Li, Jun Zhou, and Ran Tao. Spatial–Spectral Weighted and Regu- larized Tensor Sparse Correlation Filter for Object Tracking in Hyperspectral Videos. IEEE Trans. Geosci. Remote Sens., 60:1–12, 2022. FU ET AL.: HYPERVISION17

  21. [21]

    HSICityV2: Urban Scene Understanding via Hyperspectral Images, 2021

    Yuxing Huang, Tianqi Ren, Qiu Shen, Ying Fu, and Shaodi You. HSICityV2: Urban Scene Understanding via Hyperspectral Images, 2021

  22. [22]

    Hyperspectral adapter for semantic segmentation with vision foundation models.IEEE Robotics and Automation Letters, 11(3):3606–3613, 2026

    Juana Valeria Hurtado, Rohit Mohan, and Abhinav Valada. Hyperspectral adapter for semantic segmentation with vision foundation models.IEEE Robotics and Automation Letters, 11(3):3606–3613, 2026

  23. [23]

    Hyperspectral Image Dataset for Benchmarking on Salient Object Detection

    Nevrez Imamoglu, Yu Oishi, Xiaoqiang Zhang, Guanqun Ding, Yuming Fang, Toru Kouyama, and Ryosuke Nakamura. Hyperspectral Image Dataset for Benchmarking on Salient Object Detection. InProc. Int. Conf. Quality Multimedia Experience (QoMEX), pages 1–3, 2018

  24. [24]

    Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment anything. InProceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV), pages 4015–4026, October 2023

  25. [25]

    Hy- perFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery

    Jingtao Li, Yingyi Liu, Xinyu Wang, Yunning Peng, Chen Sun, Shaoyu Wang, Zhen- dong Sun, Tian Ke, Xiao Jiang, Tangwei Lu, Anran Zhao, and Yanfei Zhong. Hy- perFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 23048–23058, 2025

  26. [26]

    RGB-induced feature modulation network for hyperspectral image super-resolution.IEEE Transactions on Geoscience and Remote Sensing, 61:1–11, 2023

    Qiang Li, Maoguo Gong, Yuan Yuan, and Qi Wang. RGB-induced feature modulation network for hyperspectral image super-resolution.IEEE Transactions on Geoscience and Remote Sensing, 61:1–11, 2023. doi: 10.1109/TGRS.2023.3277486

  27. [27]

    SiamBAG: Band Attention Grouping- Based Siamese Object Tracking Network for Hyperspectral Videos.IEEE Trans

    Wei Li, Zengfu Hou, Jun Zhou, and Ran Tao. SiamBAG: Band Attention Grouping- Based Siamese Object Tracking Network for Hyperspectral Videos.IEEE Trans. Geosci. Remote Sens., 61:1–12, 2023

  28. [28]

    BAE-Net: A Band Attention Aware Ensemble Network for Hyperspectral Object Tracking

    Zhuanfeng Li, Fengchao Xiong, Jun Zhou, Jing Wang, Jianfeng Lu, and Yuntao Qian. BAE-Net: A Band Attention Aware Ensemble Network for Hyperspectral Object Tracking. InProc. IEEE Int. Conf. Image Process. (ICIP), pages 2106–2110, 2020

  29. [29]

    Material- Guided Siamese Fusion Network for Hyperspectral Object Tracking

    Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu, Jun Zhou, and Yuntao Qian. Material- Guided Siamese Fusion Network for Hyperspectral Object Tracking. InProc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), pages 2809–2813, 2022

  30. [30]

    Learning a Deep Ensemble Network With Band Importance for Hyperspectral Object Tracking

    Zhuanfeng Li, Fengchao Xiong, Jun Zhou, Jianfeng Lu, and Yuntao Qian. Learning a Deep Ensemble Network With Band Importance for Hyperspectral Object Tracking. IEEE Trans. Image Process., 32:2901–2914, 2023

  31. [31]

    Spectrum- Driven Mixed-Frequency Network for Hyperspectral Salient Object Detection.IEEE Trans

    Peifu Liu, Tingfa Xu, Huan Chen, Shiyun Zhou, Haolin Qin, and Jianan Li. Spectrum- Driven Mixed-Frequency Network for Hyperspectral Salient Object Detection.IEEE Trans. Multimedia, 26:5296–5310, 2024

  32. [32]

    Swin Transformer: Hierarchical Vision Transformer Using Shifted Win- dows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer Using Shifted Win- dows. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pages 10012–10022, 2021. 18FU ET AL.: HYPERVISION

  33. [33]

    SiamHYPER: Learning a hyperspectral object tracker from an rgb-based tracker.IEEE Trans

    Zhenqi Liu, Xinyu Wang, Yanfei Zhong, Meng Shu, and Chen Sun. SiamHYPER: Learning a hyperspectral object tracker from an rgb-based tracker.IEEE Trans. Image Process., 31:7116–7129, 2022

  34. [34]

    HSI Road: A Hyper Spectral Image Dataset For Road Segmentation

    Jiarou Lu, Huafeng Liu, Yazhou Yao, Shuyin Tao, Zhenming Tang, and Jianfeng Lu. HSI Road: A Hyper Spectral Image Dataset For Road Segmentation . InProc. IEEE Int. Conf. Multimedia Expo (ICME), pages 1–6, 2020

  35. [35]

    Nascimento, Kinjiro Amano, and David H

    Sérgio M.C. Nascimento, Kinjiro Amano, and David H. Foster. Spatial distributions of local illumination color in natural scenes.Vision Res., 120:39–44, 2016. ISSN 0042-6989

  36. [36]

    Context- Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation

    Zhenliang Ni, Xinghao Chen, Yingjie Zhai, Yehui Tang, and Yunhe Wang. Context- Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation. InProc. Eur . Conf. Comput. Vis. (ECCV), pages 239–255, 2025. ISBN 978-3-031-72943-0

  37. [37]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El- Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Jegou, Julien Mairal, Patrick ...

  38. [38]

    Zaiane, and Martin Jagersand

    Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood Dehghan, Osmar R. Zaiane, and Martin Jagersand. U2-Net: Going deeper with nested U-structure for salient object detection.Pattern Recognit., 106:107404, 2020. ISSN 0031-3203

  39. [39]

    HSOD- BIT-V2: A Challenging Benchmark for Hyperspectral Salient Object Detection.Proc

    Yuhao Qiu, Shuyan Bai, Tingfa Xu, Peifu Liu, Haolin Qin, and Jianan Li. HSOD- BIT-V2: A Challenging Benchmark for Hyperspectral Salient Object Detection.Proc. AAAI Conf. Artif. Intell., 39(6):6630–6638, 2025

  40. [40]

    Learning Transferable Visual Models From Natural Language Supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models From Natural Language Supervision. InProc. Int. Conf. Mach. Learn. (ICML), pages 8748–8763, 2021

  41. [41]

    SAM 2: Segment Anything in Images and Videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feichtenhofer. SAM 2: Segment Anything in Images and Videos. InProc. Int. Conf. Learn. Re...

  42. [42]

    A dataset for evaluating blood detection in hyperspectral images.F orensic Sci

    Michał Romaszewski, Przemysław Głomb, Arkadiusz Sochan, and Michał Cholewa. A dataset for evaluating blood detection in hyperspectral images.F orensic Sci. Int., 320: 110701, 2021. ISSN 0379-0738

  43. [43]

    U-Net: Convolutional Networks for Biomedical Image Segmentation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation. InProc. Int. Conf. Med. Image Comput. Comput.- Assist. Intervention (MICCAI), pages 234–241, 2015. ISBN 978-3-319-24574-4. FU ET AL.: HYPERVISION19

  44. [44]

    MobileNetV2: Inverted Residuals and Linear Bottlenecks

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. MobileNetV2: Inverted Residuals and Linear Bottlenecks. InProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018

  45. [45]

    Oriane Siméoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julie...

  46. [46]

    BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model

    Yiran Song, Qianyu Zhou, Xiangtai Li, Deng-Ping Fan, Xuequan Lu, and Lizhuang Ma. BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 3162–3173, 2024

  47. [47]

    HS3-Bench: A Benchmark and Strong Baseline for Hyperspectral Semantic Segmentation in Driving Scenarios

    Nick Theisen, Robin Bartsch, Dietrich Paulus, and Peer Neubert. HS3-Bench: A Benchmark and Strong Baseline for Hyperspectral Semantic Segmentation in Driving Scenarios. InProc. IEEE/RSJ Int. Conf. Intell. Robots Syst. (IROS), pages 5895–5901, 2024

  48. [48]

    Measuring the Ripeness of Fruit with Hyperspectral Imaging and Deep Learning

    Leon Amadeus Varga, Jan Makowski, and Andreas Zell. Measuring the Ripeness of Fruit with Hyperspectral Imaging and Deep Learning. InProc. Int. Joint Conf. Neural Netw. (IJCNN), pages 1–8, 2021

  49. [49]

    A Fast Neighborhood Grouping Method for Hyperspectral Band Selection.IEEE Trans

    Qi Wang, Qiang Li, and Xuelong Li. A Fast Neighborhood Grouping Method for Hyperspectral Band Selection.IEEE Trans. Geosci. Remote Sens., 59(6):5028–5039, 2021

  50. [50]

    PVT v2: Improved baselines with pyramid vision trans- former.Comput

    Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. PVT v2: Improved baselines with pyramid vision trans- former.Comput. Vis. Media, 8(3):415–424, 2022

  51. [51]

    100 radical innovation breakthroughs for the future, 2019

    Philine Warnke, Kerstin Cuhls, Ulrich Schmoch, Lea Daniel, Liviu Andreescu, Bianca Dragomir, Radu Gheorghiu, Catalina Baboschi, Adrian Curaj, Marjukka Parkkinen, and Osmo Kuusi. 100 radical innovation breakthroughs for the future, 2019

  52. [52]

    HyKo: A Spectral Dataset for Scene Understanding

    Christian Winkens, Florian Sattler, Veronika Adams, and Dietrich Paulus. HyKo: A Spectral Dataset for Scene Understanding. InProc. IEEE Int. Conf. Comput. Vis. (ICCV) Workshops, 2017

  53. [53]

    Material Based Object Tracking in Hyperspectral Videos.IEEE Trans

    Fengchao Xiong, Jun Zhou, and Yuntao Qian. Material Based Object Tracking in Hyperspectral Videos.IEEE Trans. Image Process., 29:3719–3733, 2020

  54. [54]

    Hyperspectral Object Tracking Challenge, 2025

    Fengchao Xiong, Jun Zhou, Hiep Quang Luong, Mina Zahiri, Rafal Muszynski, Wouter Charle, Yanfei Zhong, Pedram Ghamisi, and Jocelyn Chanussot. Hyperspectral Object Tracking Challenge, 2025

  55. [55]

    Fumihito Yasuma, Tomoo Mitsunaga, Daisuke Iso, and Shree K. Nayar. Generalized Assorted Pixel Camera: Postcapture Control of Resolution, Dynamic Range, and Spec- trum.IEEE Trans. Image Process., 19(9):2241–2253, 2010. 20FU ET AL.: HYPERVISION

  56. [56]

    Hyperspectral city v1

    Shaodi You, Erqi Huang, Shuaizhe Liang, Yongrong Zheng, Yunxiang Li, Fan Wang, Sen Lin, Qiu Shen, Xun Cao, Diming Zhang, et al. Hyperspectral city v1. 0 dataset and benchmark.arXiv preprint arXiv:1907.10270, 2019

  57. [57]

    Yanfei Zhong, Xin Hu, Chang Luo, Xinyu Wang, Ji Zhao, and Liangpei Zhang. Whu- hi: Uav-borne hyperspectral with high spatial resolution (h2) benchmark datasets and classifier for precise crop identification based on deep convolutional neural network with crf.Remote Sensing of Environment, 250:112012, 2020. ISSN 0034-4257

  58. [58]

    Visual Prompt Multi-Modal Tracking

    Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, and Huchuan Lu. Visual Prompt Multi-Modal Tracking. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 9516–9526, 2023