Skip to content

Training Issue #8

Description

@Aries0217

I found that after enabling LoRA, the training process terminates abruptly during the backward pass (gradient backpropagation). Could you explain why this might be happening? (I haven't made any changes to the source code.)

  • LoRA enabled (Status: Successful execution) \ LoRA disabled (Status: Automatic termination/Aborted:
        if use_lora:
            self.aggregator = apply_lora(vggt.aggregator,layer_names=['qkv','proj','fc1','fc2'], dropout=0.05)
            verify_frozen_parameters(self.aggregator)
        else:
            self.aggregator = vggt.aggregator
            self.freeze_parameters_except_heads()
  • logs:
/root/miniconda3/envs/recondrive/lib/python3.10/site-packages/torch/functional.py:534: UserWarning: torch.meshgrid: in an upcoming release, it will be required to pass the indexing argument. (Triggered internally at ../aten/src/ATen/native/TensorShape.cpp:3595.)
  return _VF.meshgrid(tensors, **kwargs)  # type: ignore[attr-defined]
/root/miniconda3/envs/recondrive/lib/python3.10/site-packages/pytorch_lightning/utilities/data.py:79: Trying to infer the `batch_size` from an ambiguous collection. The batch size we found is 2. To avoid any miscalculations, use `self.log(..., batch_size=batch_size)`.
Sanity Checking DataLoader 0:  50%|████████████████████████████████████████████████████████████                                                            | 1/2 [06:15<06:15,  0.00it/s]torch.Size([2, 42, 3, 280, 518])
Epoch 0:   0%|                                                                                                                                                   | 0/26730[00:00<?, ?it/s][GPU 0] Sampling rendered ids: [0]
(recondrive) root@autodl-container-efac46bb3d-04b9a4d8:~/autodl-tmp/ReconDrive-main# 
  • The program exits automatically after calling self._backward_fn(step_output.closure_loss):
    @override
    @torch.enable_grad()
    def closure(self, *args: Any, **kwargs: Any) -> ClosureResult:
        step_output = self._step_fn()

        if step_output.closure_loss is None:
            self.warning_cache.warn("`training_step` returned `None`. If this was on purpose, ignore this warning...")

        if self._zero_grad_fn is not None:
            self._zero_grad_fn()

        if self._backward_fn is not None and step_output.closure_loss is not None:
            self._backward_fn(step_output.closure_loss)

        return step_output

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions