Pith. sign in

REVIEW 2 minor 52 references

GaussianMap replaces fixed BEV grids with learned Gaussian primitives that adaptively represent sparse map elements for online HD map construction.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-01 06:21 UTC pith:EBGOUG4B

load-bearing objection Gaussian primitives replace fixed BEV grids for map construction and the paper claims SOTA on nuScenes and Argoverse 2, but the abstract supplies no numbers or ablations to check the gains.

arxiv 2606.31177 v1 pith:EBGOUG4B submitted 2026-06-30 cs.CV

GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction

classification cs.CV
keywords HD map constructionGaussian representationBEV featuresonline mappingmulti-sensor fusionvectorized map predictionautonomous driving
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Existing online HD map methods encode scenes with uniform high-resolution BEV grids, which waste capacity on empty space while struggling with the fine localization needed for vectorized road elements. GaussianMap instead maintains a set of adaptive Gaussian primitives on the BEV plane, each carrying its own geometric parameters and feature vector, so representational effort concentrates only where map elements exist. A feed-forward encoder refines these primitives by modeling their interactions and fusing camera and LiDAR observations, after which the primitives are splatted into a BEV feature map for final vectorized decoding. Experiments on nuScenes and Argoverse 2 show this approach reaches state-of-the-art accuracy in both camera-only and camera-LiDAR settings.

Core claim

The paper introduces a Gaussian representation consisting of a modest number of learnable primitives on the BEV plane; each primitive encodes a flexible local region through explicit geometric properties together with an attached feature vector. A Gaussian encoder progressively refines the set via interaction modeling and multi-sensor aggregation, after which the primitives are splatted into a dense BEV feature map that is decoded into vectorized map elements.

What carries the argument

A collection of learnable Gaussian primitives on the BEV plane, each defined by geometric parameters and a feature vector, that are refined by a feed-forward encoder and then splatted into a BEV feature map.

Load-bearing premise

A modest number of learned Gaussian primitives can capture the fine-grained geometry and localization of map elements at least as accurately as fixed high-resolution grids.

What would settle it

An experiment in which GaussianMap produces higher localization error on map element vertices than a comparable fixed-grid baseline on the same nuScenes or Argoverse 2 validation splits.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Representational capacity is allocated only to map-relevant regions instead of uniform grids.
  • The same Gaussian encoder supports both camera-only and camera-LiDAR fusion inputs.
  • The splatted BEV feature map can be decoded into vectorized predictions without additional post-processing stages.
  • Performance gains appear on standard autonomous-driving map benchmarks without requiring extra sensor modalities.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same adaptive primitive mechanism could be applied to other sparse scene-understanding tasks such as lane detection or traffic-sign localization.
  • Because the primitives carry explicit geometric parameters, they may enable direct metric evaluation of map accuracy without rasterization.
  • Reducing the number of primitives further could trade accuracy for lower memory and faster inference on embedded hardware.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The manuscript introduces GaussianMap, an online HD map construction method that learns an adaptive set of Gaussian primitives on the BEV plane to represent the scene, each encoding local geometry and features. A feed-forward Gaussian encoder refines the primitives via interaction modeling and multi-sensor (camera/LiDAR) feature aggregation; the primitives are then splatted to a BEV feature map and decoded into vectorized map elements. The central claim is that this representation is more efficient than fixed-resolution BEV grids for sparse map elements and yields state-of-the-art results on nuScenes and Argoverse 2 in both camera-only and camera-LiDAR fusion settings.

Significance. If the reported performance gains hold under rigorous evaluation, the shift from dense fixed grids to adaptive Gaussian primitives could improve representational efficiency for HD mapping tasks. The explicit commitment to public code release supports reproducibility, which is a strength for this line of work.

minor comments (2)
  1. The abstract states SOTA results but does not reference specific quantitative tables or ablation studies; adding a sentence pointing to the main result table would improve readability for readers scanning the abstract.
  2. Notation for the Gaussian primitives (position, covariance, feature vector) should be introduced with explicit symbols in the method section to avoid ambiguity when describing the splatting operation.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive review and recommendation of minor revision. The recognition of the efficiency advantages of adaptive Gaussian primitives over fixed BEV grids for sparse map elements is appreciated.

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper proposes a feed-forward Gaussian encoder that produces adaptive primitives on the BEV plane, followed by splatting and vectorized decoding. All performance claims rest on empirical results from external benchmarks (nuScenes and Argoverse 2) rather than any internal derivation that reduces to fitted inputs or self-citations. No equations, uniqueness theorems, or ansatzes are shown that loop back to the method's own definitions or prior author work; the architecture is presented as a standard learned model whose validity is tested outside the training distribution.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 1 invented entities

Based solely on the abstract, the central claim rests on the domain assumption that map elements are spatially sparse and that adaptive Gaussians can allocate capacity more effectively than uniform grids; no free parameters or invented entities beyond the Gaussian primitives themselves are described.

axioms (1)
  • domain assumption Map elements are spatially sparse yet require fine-grained geometric localization, rendering uniform BEV grids redundant.
    Stated in the second sentence of the abstract as the motivation for moving away from fixed-resolution grids.
invented entities (1)
  • Gaussian primitives no independent evidence
    purpose: Adaptive local region encoding on the BEV plane with geometric properties and feature vectors
    Introduced as the core scene representation; no independent evidence outside the paper is provided in the abstract.

pith-pipeline@v0.9.1-grok · 5775 in / 1337 out tokens · 22892 ms · 2026-07-01T06:21:58.316383+00:00 · methodology

0 comments
read the original abstract

Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local vectorized maps from onboard sensor observations. Existing methods commonly adopt bird's-eye-view (BEV) features as the intermediate scene representation, encoding the surrounding space with fixed-resolution dense grids. However, map elements are spatially sparse yet require fine-grained geometric localization, making uniformly allocated BEV representations redundant and less effective for vectorized map prediction. In this work, we propose GaussianMap, an online HD map construction framework that learns an adaptive Gaussian representation of the surrounding scene. This representation consists of a set of Gaussian primitives on the BEV plane, each encoding a flexible local region with geometric properties and a feature vector, allowing the model to allocate representational capacity to map-relevant regions. To generate such a representation from sensor observations, we introduce a feed-forward Gaussian encoder that progressively refines these primitives through Gaussian interaction modeling and multi-sensor feature aggregation. The refined Gaussian representation is then splatted into a BEV feature map and decoded into vectorized map predictions. Extensive experiments on nuScenes and Argoverse 2 datasets demonstrate that GaussianMap achieves state-of-the-art performance in both camera-only and camera-LiDAR fusion settings. Our code will be made publicly available.

Figures

Figures reproduced from arXiv: 2606.31177 by Hongyu Lyu, Julie Stephany Berrio Perez, Mao Shan, Stewart Worrall.

Figure 1
Figure 1. Figure 1: Motivation of GaussianMap. Compared with grid-based BEV features that uniformly allocate representational capacity, GaussianMap represents the surrounding scene with adaptive Gaus￾sian primitives to focus on map-relevant regions and perseve fine￾grained geometry. Gaussian colors are used only for visualization. features as the intermediate scene representation. Although BEV features provide a structured to… view at source ↗
Figure 2
Figure 2. Figure 2: Overall framework of GaussianMap. The framework takes multi-camera images and an optional LiDAR point cloud as inputs. The Gaussian encoder generates an adaptive Gaussian representation on the BEV plane from the extracted sensor features. This representation is converted into a BEV feature map via Gaussian-to-BEV splatting, from which the map decoder produces the vectorized map prediction. For clarity, we … view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative results on the nuScenes validation set. (a) Camera-only setting. (b) Camera-LiDAR fusion setting. We compare the vectorized map predictions by GaussianMap with those from MapQR and MapQR+DAMap, using the ground truth as the reference. For GaussianMap, we additionally visualize the Gaussians alongside the predictions, retaining those with opacity higher than 0.4 for clarity. interaction modeling… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 2 canonical work pages

  1. [1]

    Online high-definition map construction for autonomous vehicles: A comprehensive survey,

    H. Lyu, J. S. Berrio Perez, Y . Huang, K. Li, M. Shan, and S. Worrall, “Online high-definition map construction for autonomous vehicles: A comprehensive survey,”Journal of Sensor and Actuator Networks, vol. 14, no. 1, p. 15, 2025

  2. [2]

    High-definition maps: Comprehensive survey, challenges, and future perspectives,

    G. Elghazaly, R. Frank, S. Harvey, and S. Safko, “High-definition maps: Comprehensive survey, challenges, and future perspectives,” IEEE Open Journal of Intelligent Transportation Systems, vol. 4, pp. 527–550, 2023

  3. [3]

    Loam: Lidar odometry and mapping in real-time

    J. Zhang, S. Singhet al., “Loam: Lidar odometry and mapping in real-time.” inRobotics: Science and systems, vol. 2, no. 9. Berkeley, CA, 2014, pp. 1–9

  4. [4]

    Lego-loam: Lightweight and ground- optimized lidar odometry and mapping on variable terrain,

    T. Shan and B. Englot, “Lego-loam: Lightweight and ground- optimized lidar odometry and mapping on variable terrain,” in2018 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, 2018, pp. 4758–4765

  5. [5]

    Long-term map maintenance pipeline for autonomous vehicles,

    J. S. Berrio, S. Worrall, M. Shan, and E. Nebot, “Long-term map maintenance pipeline for autonomous vehicles,”IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 8, pp. 10 427–10 440, 2021

  6. [6]

    Hdmapnet: An online hd map construction and evaluation framework,

    Q. Li, Y . Wang, Y . Wang, and H. Zhao, “Hdmapnet: An online hd map construction and evaluation framework,” inInternational Conference on Robotics and Automation (ICRA). IEEE, 2022, pp. 4628–4634

  7. [7]

    Vectormapnet: End-to-end vectorized hd map learning,

    Y . Liu, T. Yuan, Y . Wang, Y . Wang, and H. Zhao, “Vectormapnet: End-to-end vectorized hd map learning,” inInternational Conference on Machine Learning. PMLR, 2023, pp. 22 352–22 369

  8. [8]

    Maptr: Structured modeling and learning for online vectorized hd map construction,

    B. Liao, S. Chen, X. Wang, T. Cheng, Q. Zhang, W. Liu, and C. Huang, “Maptr: Structured modeling and learning for online vectorized hd map construction,” inThe Eleventh International Conference on Learning Representations, 2023

  9. [9]

    Maprf: Weakly supervised online hd map construction via nerf-guided self-training,

    H. Lyu, T. Monninger, J. S. B. Perez, M. Shan, Z. Ming, and S. Worrall, “Maprf: Weakly supervised online hd map construction via nerf-guided self-training,” inIEEE International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2026

  10. [10]

    Maptrv2: An end-to-end framework for online vectorized hd map construction,

    B. Liao, S. Chen, Y . Zhang, B. Jiang, Q. Zhang, W. Liu, C. Huang, and X. Wang, “Maptrv2: An end-to-end framework for online vectorized hd map construction,”International Journal of Computer Vision, vol. 133, no. 3, pp. 1352–1374, 2025

  11. [11]

    Leveraging enhanced queries of point sets for vectorized map construction,

    Z. Liu, X. Zhang, G. Liu, J. Zhao, and N. Xu, “Leveraging enhanced queries of point sets for vectorized map construction,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 461–477

  12. [12]

    Damap: Distance-aware mapnet for high quality hd map construction,

    J. Dong, C. Li, Y . Lin, J. Fu, S. Zhou, and N. Zheng, “Damap: Distance-aware mapnet for high quality hd map construction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 5285–5294

  13. [13]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, G. Drettakiset al., “3d gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023

  14. [14]

    Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,

    J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” inEuropean conference on computer vision. Springer, 2020, pp. 194–210

  15. [15]

    Bev- former: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai, “Bev- former: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 3, pp. 2020–2036, 2024

  16. [16]

    Bevfusion: Multi-task multi-sensor fusion with unified bird’s- eye view representation,

    Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, “Bevfusion: Multi-task multi-sensor fusion with unified bird’s- eye view representation,” in2023 IEEE international conference on robotics and automation (ICRA). IEEE, 2023, pp. 2774–2781

  17. [17]

    End-to-end vectorized hd- map construction with piecewise bezier curve,

    L. Qiao, W. Ding, X. Qiu, and C. Zhang, “End-to-end vectorized hd- map construction with piecewise bezier curve,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 13 218–13 228

  18. [18]

    Pivotnet: Vectorized pivot learning for end-to-end hd map construction,

    W. Ding, L. Qiao, X. Qiu, and C. Zhang, “Pivotnet: Vectorized pivot learning for end-to-end hd map construction,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3672–3682

  19. [19]

    Streammapnet: Streaming mapping network for vectorized online hd map construc- tion,

    T. Yuan, Y . Liu, Y . Wang, Y . Wang, and H. Zhao, “Streammapnet: Streaming mapping network for vectorized online hd map construc- tion,” inProceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, 2024, pp. 7356–7365

  20. [20]

    Mgmap: Mask-guided learning for online vectorized hd map construction,

    X. Liu, S. Wang, W. Li, R. Yang, J. Chen, and J. Zhu, “Mgmap: Mask-guided learning for online vectorized hd map construction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 812–14 821

  21. [21]

    Refdiffmap: Diffusion-guided progressive refinement for vectorized hd map con- struction,

    W. Gao, E. Chang, J. Fu, Z. Zhu, S. Chen, and N. Zheng, “Refdiffmap: Diffusion-guided progressive refinement for vectorized hd map con- struction,”IEEE Robotics and Automation Letters, vol. 11, no. 3, pp. 2554–2561, 2026

  22. [22]

    Admap: Anti-disturbance framework for vectorized hd map construction,

    H. Hu, F. Wang, Y . Wang, L. Hu, J. Xu, and Z. Zhang, “Admap: Anti-disturbance framework for vectorized hd map construction,” in European Conference on Computer Vision. Springer, 2024, pp. 311– 326

  23. [23]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021

  24. [24]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 20 310–20 320

  25. [25]

    Dreamgaussian: Generative gaussian splatting for efficient 3d content creation,

    J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng, “Dreamgaussian: Generative gaussian splatting for efficient 3d content creation,” in International Conference on Learning Representations, vol. 2024, 2024, pp. 33 879–33 896

  26. [26]

    Langsplat: 3d language gaussian splatting,

    M. Qin, W. Li, J. Zhou, H. Wang, and H. Pfister, “Langsplat: 3d language gaussian splatting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 051–20 060

  27. [27]

    Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy predic- tion,

    Y . Huang, W. Zheng, Y . Zhang, J. Zhou, and J. Lu, “Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy predic- tion,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 376–393

  28. [28]

    Gaussianformer-2: Probabilistic gaussian superposition for efficient 3d occupancy prediction,

    Y . Huang, A. Thammatadatrakoon, W. Zheng, Y . Zhang, D. Du, and J. Lu, “Gaussianformer-2: Probabilistic gaussian superposition for efficient 3d occupancy prediction,” inProceedings of the computer vision and pattern recognition conference, 2025, pp. 27 477–27 486

  29. [29]

    Gaussianbev: 3d gaussian representation meets perception models for bev segmentation,

    F. Chabot, N. Granger, and G. Lapouge, “Gaussianbev: 3d gaussian representation meets perception models for bev segmentation,” in2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 2250–2259

  30. [30]

    Toward real-world bev per- ception: Depth uncertainty estimation via gaussian splatting,

    S.-W. Lu, Y .-H. Tsai, and Y .-T. Chen, “Toward real-world bev per- ception: Depth uncertainty estimation via gaussian splatting,” inPro- ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 17 124–17 133

  31. [31]

    Gaussianad: Gaussian-centric end- to-end autonomous driving,

    W. Zheng, J. Wu, Y . Zheng, S. Zuo, Z. Xie, L. Yang, Y . Pan, Z. Hao, P. Jia, X. Langet al., “Gaussianad: Gaussian-centric end- to-end autonomous driving,”arXiv preprint arXiv:2412.10371, 2024

  32. [32]

    Pointpainting: Sequential fusion for 3d object detection,

    S. V ora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4604–4612

  33. [33]

    Pointaugmenting: Cross- modal augmentation for 3d object detection,

    C. Wang, C. Ma, M. Zhu, and X. Yang, “Pointaugmenting: Cross- modal augmentation for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 11 794–11 803

  34. [34]

    Multimodal virtual point 3d detection,

    T. Yin, X. Zhou, and P. Kr ¨ahenb¨uhl, “Multimodal virtual point 3d detection,”Advances in Neural Information Processing Systems, vol. 34, pp. 16 494–16 507, 2021

  35. [35]

    Multi-view 3d object detection network for autonomous driving,

    X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object detection network for autonomous driving,” inProceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2017, pp. 1907–1915

  36. [36]

    Centerfusion: Center-based radar and camera fusion for 3d object detection,

    R. Nabati and H. Qi, “Centerfusion: Center-based radar and camera fusion for 3d object detection,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 1527–1536

  37. [37]

    Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,

    X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1090–1099

  38. [38]

    Bevfusion: A simple and robust lidar-camera fusion framework,

    T. Liang, H. Xie, K. Yu, Z. Xia, Z. Lin, Y . Wang, T. Tang, B. Wang, and Z. Tang, “Bevfusion: A simple and robust lidar-camera fusion framework,”Advances in neural information processing systems, vol. 35, pp. 10 421–10 434, 2022

  39. [39]

    Mapfusion: A novel bev feature fusion network for multi-modal map construction,

    X. Hao, Y . Diao, M. Wei, Y . Yang, P. Hao, R. Yin, H. Zhang, W. Li, S. Zhao, and Y . Liu, “Mapfusion: A novel bev feature fusion network for multi-modal map construction,”Information Fusion, vol. 119, p. 103018, 2025

  40. [40]

    What really matters for robust multi-sensor hd map construction?

    X. Hao, Y . Zhao, Y . Ji, L. Dai, P. Hao, D. Li, S. Cheng, and R. Yin, “What really matters for robust multi-sensor hd map construction?” in2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2025, pp. 1298–1304

  41. [41]

    Sef-map: Subspace-decomposed expert fusion for robust multimodal hd map prediction,

    H. Fu, L. Zhang, H. Li, R. Hu, Z. Li, G. Liu, Z. Tan, L. Chen, H. Ye, and X. Hao, “Sef-map: Subspace-decomposed expert fusion for robust multimodal hd map prediction,” inInternational Conference on Robotics and Automation (ICRA). IEEE, 2026

  42. [42]

    Deformable detr: Deformable transformers for end-to-end object detection,

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations, 2021

  43. [43]

    Efficient and robust 2d-to-bev representation learning via geometry- guided kernel transformer,

    S. Chen, T. Cheng, X. Wang, W. Meng, Q. Zhang, and W. Liu, “Efficient and robust 2d-to-bev representation learning via geometry- guided kernel transformer,”arXiv preprint arXiv:2206.04584, 2022

  44. [44]

    nuscenes: A multimodal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer vision and pattern recognition, 2020, pp. 11 621–11 631

  45. [45]

    Argoverse 2: Next generation datasets for self-driving perception and forecasting,

    B. Wilson, W. Qi, T. Agarwal, J. Lambert, J. Singh, S. Khandelwal, B. Pan, R. Kumar, A. Hartnett, J. K. Pontes, D. Ramanan, P. Carr, and J. Hays, “Argoverse 2: Next generation datasets for self-driving perception and forecasting,” inProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks), 2021

  46. [46]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE Conference on Computer vision and pattern recognition, 2016, pp. 770–778

  47. [47]

    Second: Sparsely embedded convolutional detection,

    Y . Yan, Y . Mao, and B. Li, “Second: Sparsely embedded convolutional detection,”Sensors, vol. 18, no. 10, p. 3337, 2018

  48. [48]

    Online map vectorization for autonomous driving: A rasterization perspective,

    G. Zhang, J. Lin, S. Wu, Y . Song, Z. Luo, Y . Xue, S. Lu, and Z. Wang, “Online map vectorization for autonomous driving: A rasterization perspective,”Advances in Neural Information Processing Systems, 2023

  49. [49]

    Heightmapnet: Explicit height modeling for end-to-end hd map learning,

    W. Qiu, S. Pang, H. Zhang, J. Fang, and J. Xue, “Heightmapnet: Explicit height modeling for end-to-end hd map learning,” in2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 6022–6031

  50. [50]

    Himap: Hybrid representation learning for end-to-end vectorized hd map construction,

    Y . Zhou, H. Zhang, J. Yu, Y . Yang, S. Jung, S.-I. Park, and B. Yoo, “Himap: Hybrid representation learning for end-to-end vectorized hd map construction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 15 396–15 406

  51. [51]

    Relmap: Enhancing online map construction with class-aware spatial relation and semantic priors,

    T. Cai, Y . Zhang, Z. Zhou, Z. Huang, and J. Ma, “Relmap: Enhancing online map construction with class-aware spatial relation and semantic priors,” inInternational Conference on Robotics and Automation (ICRA). IEEE, 2026

  52. [52]

    Online vectorized hd map construction using geometry,

    Z. Zhang, Y . Zhang, X. Ding, F. Jin, and X. Yue, “Online vectorized hd map construction using geometry,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 73–90