Given:
--gpu_per_parallel 2 \
--parallel_per_task 1 \
--circular_eval False \
--score_target
BUG:
Shard 0 (4329 samples): 0%| | 0/4329 [00:00<?, ?it/s]You shouldn't move a model that is dispatched using accelerate hooks.
[Shard 0] ❌ Error: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cuda:1! (when checking argument for argument mat2 in method wrapper_CUDA_bmm)
[Shard 0] ⚠️ Max retries reached, skipping sample
Shard 0 (4329 samples): 0%| | 1/4329 [00:04<4:53:38, 4.07s/it]You shouldn't move a model that is dispatched using accelerate hooks.
[Shard 0] ❌ Error: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cuda:1! (when checking argument for argument mat2 in method wrapper_CUDA_bmm)
Given:
BUG:
Shard 0 (4329 samples): 0%| | 0/4329 [00:00<?, ?it/s]You shouldn't move a model that is dispatched using accelerate hooks.⚠️ Max retries reached, skipping sample
[Shard 0] ❌ Error: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cuda:1! (when checking argument for argument mat2 in method wrapper_CUDA_bmm)
[Shard 0]
Shard 0 (4329 samples): 0%| | 1/4329 [00:04<4:53:38, 4.07s/it]You shouldn't move a model that is dispatched using accelerate hooks.
[Shard 0] ❌ Error: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cuda:1! (when checking argument for argument mat2 in method wrapper_CUDA_bmm)