Data Loading and Preprocessing Logic in the Lift-Splat-Shoot Architecture
Dataset Initialization and Configuration
The data pipeline begins by instantiating the NuScenes API to manage scene metadata. The SegmentationData class orchestrates the separation of training and validation splits. It invokes logic to generate lists of scene tokens, separating them based on predefined split configurations. Subsequently, a preprocessing step samples frames from video sequences at fixed intervals (e.g., 0.5 seconds) to create a list of sample indices (ixes).
A critical setup step involves defining the Bird's Eye View (BEV) grid parameters using the gen_dx_bx utility. This establishes the resolution (dx, typically 0.5 meters), the starting coordinate of the grid center (bx), and the grid dimensions (nx, e.g., 200x200 cells). These parameters define the spatial extent of the output feature map.
Camera Selection and Image Augmentation
During the data fetching phase, the system selects a subset of cameras from the available sensors—often choosing 5 out of 6 cameras—and retrieves the corresponding image data and calibration parameters. To improve model robustness, a stochastic augmentation strategy is applied. This process calculates random parameters for resizing, cropping, flipping, and rotation.
def compute_augmentation_params(config, is_train_mode):
src_h, src_w = config['orig_dim']
tgt_h, tgt_w = config['target_dim']
if is_train_mode:
# Random scaling within a defined limits
scale_factor = np.random.uniform(*config['resize_limits'])
new_w, new_h = int(src_w * scale_factor), int(src_h * scale_factor)
# Random crop calculation
crop_y_max = int((1 - np.random.uniform(*config['crop_bottom_limits'])) * new_h) - tgt_h
crop_x = int(np.random.uniform(0, max(0, new_w - tgt_w)))
crop_box = (crop_x, crop_y_max, crop_x + tgt_w, crop_y_max + tgt_h)
# Random flip and rotation
do_flip = config['flip_enabled'] and np.random.randint(2)
rot_angle = np.random.uniform(*config['rotation_limits'])
else:
# Deterministic resizing for validation
scale_factor = max(tgt_h / src_h, tgt_w / src_w)
new_w, new_h = int(src_w * scale_factor), int(src_h * scale_factor)
crop_y = int((1 - np.mean(config['crop_bottom_limits'])) * new_h) - tgt_h
crop_x = int(max(0, new_w - tgt_w) / 2)
crop_box = (crop_x, crop_y, crop_x + tgt_w, crop_y + tgt_h)
do_flip = False
rot_angle = 0.0
return scale_factor, (new_w, new_h), crop_box, do_flip, rot_angle
Geometric Transformation and Calibration Update
When images undergo geometric augmentation, the intrinsic and extrinsic calibration matrices must be adjusted accordingly to maintain geometric consistency. The img_transform function applies the physical modifications to the image pixel array while simultaneously updating a 'post-homography' transformation matrix. This matrix accounts for the changes in rotation and translation induced by resizing, cropping, and rotating, ensuring that the mapping from image space to 3D space remains valid.
def apply_image_transforms(img, rot_mat, tran_mat, scale, dims, crop, flip, angle):
# Apply visual augmentations
img = img.resize(dims)
img = img.crop(crop)
if flip:
img = img.transpose(method=Image.FLIP_LEFT_RIGHT)
img = img.rotate(angle)
# Update geometric transformation matrices
rot_mat *= scale
tran_mat -= torch.tensor(crop[:2])
if flip:
flip_mat = torch.tensor([[-1, 0], [0, 1]])
flip_tran = torch.tensor([crop[2] - crop[0], 0])
rot_mat = flip_mat.matmul(rot_mat)
tran_mat = flip_mat.matmul(tran_mat) + flip_tran
# Rotate coordinates around the crop center
rot_rad = angle / 180 * np.pi
rot_transform = get_rotation_matrix(rot_rad)
center_offset = torch.tensor([crop[2] - crop[0], crop[3] - crop[1]]) / 2
# Adjust translation to account for rotation around the center
tran_mat = rot_transform.matmul(tran_mat) + (center_offset - rot_transform.matmul(center_offset))
rot_mat = rot_transform.matmul(rot_mat)
return img, rot_mat, tran_mat
Ground Truth BEV Map Generation
The final step involves generating the ground truth binary segmentation map (binimg). This requires projecting 3D object annotations into the pre-defined BEV grid. The process reads the ego-pose translation and rotation to construct an inverse transformation matrix, mapping coordinates from the global frame to the ego-vehicle frame.
For each annotated instance, the algorithm retrieves the bounding box dimensions (width, height, length) and calculates the corner coordinates. The bottom corners are extracted, transformed into the ego coordinate system, and discretized into grid indices. The resulting tensors typically include the stacked multi-view images, rotation and translation parameters for each camera, intrinsic matrices, and the post-augmentation transforms, alongside the generated BEV ground truth map.
# Projecting bounding box corners to BEV grid coordinates
# pts: corners in world coordinates
# ego_inv: inverse transformation matrix of the ego vehicle
# grid_params: containing dx (resolution) and bx (bounds)
# Transform to ego frame
pts_ego = (ego_inv[:3, :3] @ pts.T).T + ego_inv[:3, 3]
# Convert metric coordinates to grid indices
grid_indices = np.round(
(pts_ego[:, :2] - grid_params['bx'][:2] + grid_params['dx'][:2] / 2.0) / grid_params['dx'][:2]
).astype(np.int32)
# Resulting data shapes
# imgs: [num_cams, 3, H, W]
# rots: [num_cams, 3, 3]
# trans: [num_cams, 3]
# intrins: [num_cams, 3, 3]
# post_rots: [num_cams, 3, 3]
# post_trans: [num_cams, 3]
# binimg: [1, grid_h, grid_w]