I am seeing segmentation fault at exit. It can be reproduced with
bash ./test/start_test.sh ./test/torch_allreduce_test.py --backend=ucc
See also the test results of #14.
I looked deep into it, seems that the root causes are two issues:
Issue 1: The destruction of ProcessGroupUCC happens earlier than ~CommUCC. The latter will invoke ucc_context_destroy to destroy UCC's context, which uses a c10d::Store object as an out-of-band communicator for allgather. However, at the time ucc_context_destroy is called, the c10d::Store object is already destroyed, which causes access to an invalid pointer. I have a fix of this issue at #13.
Issue 2: The ~ProcessGroupUCC will do ucc_destroy_team(team), but at the time when this happens, the progress_loop thread can still be running ucc_context_progress, which triggers a segfault.
In my test, I modified the beginning of start_test.sh to size=2 instead of size=4, and the rank=0 process will fail due to issue 2, and the rank=1 process will fail due to issue 1.
cc: @Sergei-Lebedev
I am seeing segmentation fault at exit. It can be reproduced with
See also the test results of #14.
I looked deep into it, seems that the root causes are two issues:
Issue 1: The destruction of
ProcessGroupUCChappens earlier than~CommUCC. The latter will invokeucc_context_destroyto destroy UCC's context, which uses ac10d::Storeobject as an out-of-band communicator for allgather. However, at the timeucc_context_destroyis called, thec10d::Storeobject is already destroyed, which causes access to an invalid pointer. I have a fix of this issue at #13.Issue 2: The
~ProcessGroupUCCwill doucc_destroy_team(team), but at the time when this happens, theprogress_loopthread can still be runningucc_context_progress, which triggers a segfault.In my test, I modified the beginning of
start_test.shtosize=2instead ofsize=4, and the rank=0 process will fail due to issue 2, and the rank=1 process will fail due to issue 1.cc: @Sergei-Lebedev