Hi, I've been trying to utilize your implementation of multiview LSS backbone for my thesis. I was wondering if the returned depth should be visually interpretable. I am aware that the depth tensor is spatially downsampled from original input images, but still I expect to get coarse depth image. Right now i just get a full noise like image(see image below) after decoding depth bins to meters (argmax(depths) * bin_size). I was expecting to see something like depth net output vis here

Hi, I've been trying to utilize your implementation of multiview LSS backbone for my thesis. I was wondering if the returned depth should be visually interpretable. I am aware that the depth tensor is spatially downsampled from original input images, but still I expect to get coarse depth image. Right now i just get a full noise like image(see image below) after decoding depth bins to meters (
argmax(depths) * bin_size). I was expecting to see something like depth net output vis here