Imitating Human Driving for Online High-Definition Map Construction
[Paper]
Abstract:
High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images provide crucial information. However, due to the modality and perspective differences among these data sources, existing methods often struggle to effectively align and fuse them, making online HD map construction still challenging. To address these issues, we propose Driver2Map, an online HD map construction model inspired by human drivers. Unlike existing HD map construction models that utilize only two modalities, our Driver2Map can simultaneously exploit three modalities. Specifically, we propose a ''two-stage alignment'' strategy to reduce spatial misalignment across different modalities. Additionally, we introduce ''Pose-Guided BEV Fusion'', a BEV (bird's-eye-view) generation module that leverages camera pose information to adaptively weight multi-view features, thereby effectively suppressing cross-view feature overlap during BEV generation. Also, we design a ''Pretrained Prior for Map Refinement'' module to refine the initial prediction by learning map structure priors, thus improving the HD map prediction under dynamic occlusions. Extensive experiments demonstrate that Driver2Map outperforms existing methods on both IoU and AP metrics.
- Create a new conda environment, and activate:
conda create -n hdmap python=3.9
conda activate hdmap
- Install all dependencies by running:
pip install -r requirement.txt
Note: CUDA 11.3 needed.
- Clone the Driver2Map project to your local machine and navigate to the project root directory.
git clone https://github.com/UserBits/Driver2Map.git
cd Driver2Map
-
Download nuScenes dataset and put it to
./dataset/folder -
Download the complementary satellite map tiles for nuscenes and put it to
./satmap/folder. -
Download the complementary SD map tiles for nuscenes and put it to
./sdmap/folder.
If you want to generate the Sat & SD map tiles by yourself, follow step 5&6.
-
Follow Generating map inputs to generate the complementary satellite/prior map patches for nuScenes. Place the generated files (for example, under
./satmap/prior_map_trainval/) in the project directory. -
In the same Generating map inputs section, run
data/make_data.pyto generate the complementary SD map patches. Place the generated files (for example, under./sdmap/) in the project directory.
The final folder structure
Driver2Map
|-- data/
|-- evaluation/
|-- icon/
|-- model/
|-- postprocess/
|-- preprocess/
|-- dataset/
│ ├── nuScenes_trainval/
│ │ ├── maps/
│ │ ├── samples/
│ │ ├── sweeps/
| | ├── v1.0-trainval/
|-- satmap
│ ├── map/
│ ├── prior_map_trainval/
│ │ ├── map_prior.json
│ ├── prior_map_test/
|-- sdmap
│ ├── boston-seaport.png
│ ├── singapore-onenorth.png
│ ├── singapore-hollandvillage.png
│ ├── singapore-queenstown.png
│ ├── v1.0-mini
│ ├── v1.0-trainval
│ ├── v1.0-test
This is a guide on how to train the backbone model, excluding the ''Pretrained Prior for Map Refinement'' module.
We recommend not modifying certain training configurations in train.py.
- To start training from epoch 0, run:
NCCL_TIMEOUT=360 CUDA_VISIBLE_DEVICES=1,2,3,4 python -m torch.distributed.launch --nproc_per_node=4 train.py --instance_seg --direction_pred --fusion_mode seg-masked-atten --align_fusion --dataroot ./dataset/nuScenes_trainval --prior_map_root ./satmap/prior_map_trainval --version v1.0-trainval --logdir [output dir] --nepochs=30
Note: replace [output dir] with your preferred directory for storing parameter checkpoints and training logs.
- To start training from one certain epoch, run:
NCCL_TIMEOUT=360 CUDA_VISIBLE_DEVICES=1,2,3,4 python -m torch.distributed.launch --nproc_per_node=4 train.py --resume --resume_checkpoint [checkpoint file path] --instance_seg --direction_pred --fusion_mode seg-masked-atten --align_fusion --dataroot ./dataset/nuScenes_trainval --prior_map_root ./satmap/prior_map_trainval --version v1.0-trainval --logdir [output dir] --nepochs=30
Note: replace [checkpoint file path] with the checkpoint file path of the previous epoch.
Note: replace [output dir] with your preferred directory for storing parameter checkpoints and training logs.
- Get the pretrained parameters of ''Pretrained Prior for Map Refinement'' module, and put it to
./modelfolder: mae_head_and_seg_head.pt
Note: you should name the checkpoint file mae_head_and_seg_head.pt, do not change the file name.
- Run:
NCCL_TIMEOUT=360 CUDA_VISIBLE_DEVICES=1,2,3,4 python -m torch.distributed.launch --nproc_per_node=4 train.py --instance_seg --direction_pred --fusion_mode seg-masked-atten --align_fusion --dataroot ./dataset/nuScenes_trainval --prior_map_root ./satmap/prior_map_trainval --version v1.0-trainval --logdir [output dir] --nepochs=45 --hdmap_finetune
- If you want to pretrain ''Pretrained Prior for Map Refinement'' module on your own, refer to project Drviver2Map-pretrain
-
Get the full checkpoint file (including all parameters and components) from: driver2map-fullparameter-model37.pt
-
For IoU metrics, run:
CUDA_VISIBLE_DEVICES=0 python evaluate.py --modelf [checkpoint file path] --instance_seg --direction_pred --fusion_mode seg-masked-atten --align_fusion --dataroot ./dataset/nuScenes_trainval --version v1.0-trainval --hdmap_finetune
Note: replace [checkpoint file path] with the checkpoint file path.
- For AP metrics:
First, export results to .json file:
CUDA_VISIBLE_DEVICES=0 python export_pred_to_json.py --modelf [checkpoint file path] --instance_seg --direction_pred --fusion_mode seg-masked-atten --align_fusion --dataroot ./dataset/nuScenes_trainval/ --version v1.0-trainval --output [output file path] --hdmap_finetune
Note: replace [checkpoint file path] and [output file path] with yours.
Next, run:
CUDA_VISIBLE_DEVICES=0 python evaluate_json.py --result_path `[json file path]`
Note: replace [json file path] with the .json file you had just generated.
evaluate_fps.py measures pure model-forward throughput on the GPU. Dataset
loading, host-to-device copies, checkpoint loading, and metric computation are
excluded from the timed region. The script performs a warm-up phase before
recording CUDA event timings, then reports minimum/maximum FPS, average FPS,
mean iteration FPS, and average latency.
For example, to benchmark the validation split with one GPU:
CUDA_VISIBLE_DEVICES=0 python evaluate_fps.py \
--modelf [checkpoint file path] \
--instance_seg --direction_pred --fusion_mode seg-masked-atten \
--align_fusion --dataroot ./dataset/nuScenes_trainval \
--prior_map_root ./satmap/prior_map_trainval \
--sd_map_root ./sdmap --version v1.0-trainval \
--hdmap_finetune --warmup 100 --iterations 100 --bsz 1
Replace [checkpoint file path] with the checkpoint to test. Useful options
include --warmup (untimed warm-up iterations), --iterations (timed
iterations), --bsz (batch size), and --split train|val. Add --no_sd_map
to measure the model without SD-map fusion. A CUDA-capable GPU is required.
The directional camera-weighted fusion is implemented in
model/homography.py, in IPM.compute_camera_weights. The two coefficients
alpha = 1.0 # distance-weight contribution
beta = 2.0 # viewing-direction contribution
control the relative influence of distance and camera viewing direction. Edit these values, run the same training/evaluation command, and compare the resulting metrics for an ablation study. Keep all other settings (checkpoint, data split, and random seed) fixed for a fair comparison.
First of all, download the global Road Map, SD Map, Satellite Map of Boston-Seaport, Singapore-OneNorth, Singapore-Queenstown, and Singapore-HollandVillage of NuScenes to /map_files directory.
Then following the instructions below to generate SD & satellite map tiles.
preprocess/prior_map_preprocess.py extracts a pose-aligned map patch for
every nuScenes sample and writes the images plus a prior_map.json index. Before
running it, set the source map image names in map_files, data_root, and the
nuScenes dataroot/prior_map_root arguments in the __main__ block to match
your machine. The generated directory should then be passed to
--prior_map_root (for example ./satmap/prior_map_trainval).
The script uses the ego pose and map location to crop and rotate each patch, so the corresponding nuScenes map assets must be available before execution.
data/make_data.py can generate aligned per-sample map patches (the current
implementation reads the four nuScenes SD-map images and writes sd_map.json).
Edit version, out_dir, the nuScenes dataroot, and the four input map image
paths near the top of the file before running:
The output directory contains one rotated/cropped map image per sample and an
sd_map.json mapping ego-pose tokens to image files. To generate satellite-map
patches instead, provide the satellite tile images in those input-map paths;
the same pose-based cropping pipeline can then be used to produce the
corresponding sat-map set.
- To visualize labels (Ground Truth):
python vis_label.py ./dataset/nuScenes_trainval --version v1.0-trainval
Note: change the input/output path settings in vis_label.py before running.
- To visualize predictions:
CUDA_VISIBLE_DEVICES=4 python vis_pred.py --modelf [checkpoint file path] --instance_seg --direction_pred --map_prior --hdmap_finetune --align_fusion
Note: replace [checkpoint file path] with the checkpoint file path.
If you found this paper or codebase useful, please cite our paper:
@misc{yin2026driver2map,
title={Driver2Map: Imitating Human Driving for Online High-Definition Map Construction},
author={Pan Yin and Runtian Xia and Weisong Kuang and Kaiyu Li and Cong Zhao and Xiangyong Cao},
year={2026},
eprint={2608.01338},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.01338},
}