Skip to content
This repository was archived by the owner on Jun 1, 2025. It is now read-only.
This repository was archived by the owner on Jun 1, 2025. It is now read-only.

Some tests fails #16

Description

@zasdfgbnm

Failure 1

bash start_test.sh torch_allreduce_test.py --backend ucc

Wrong results, segfault at exit

Failure 2

bash start_test.sh torch_allreduce_test.py --backend gloo --use-cuda

Wrong results

Failure 3

If I change the torch_ucc_test_setup.py like this

- dist.init_process_group('ucc', rank=comm_rank, world_size=comm_size)
+ dist.init_process_group('gloo', rank=comm_rank, world_size=comm_size)

and run

bash start_test.sh torch_allreduce_test.py --backend ucc  # on CPU

I will get

Traceback (most recent call last):
  File "/home/gaoxiang/torch_ucc/test/torch_allreduce_test.py", line 24, in <module>
    dist.all_reduce(tensor_ucc)
  File "/home/gaoxiang/.local/lib/python3.9/site-packages/torch/distributed/distributed_c10d.py", line 1269, in all_reduce
    work.wait()
RuntimeError: [../third_party/gloo/gloo/transport/tcp/pair.cc:598] Connection closed by peer [fe80::3697:f6ff:fe32:8d08]:21658

Failure 4

If I change the torch_ucc_test_setup.py like this

- dist.init_process_group('ucc', rank=comm_rank, world_size=comm_size)
+ dist.init_process_group('nccl', rank=comm_rank, world_size=comm_size)

and run

bash start_test.sh torch_allreduce_test.py --backend ucc --use-cuda

I will get wrong results:

Allreduce test
count                result
failed on rank 0
failed on rank 1
1                    Failed

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions