Skip to content

[Question] Rationale for timestep-pair sampling vs. LCM-style DDIM anchors #3

Description

@Aircraft-1111

Hi, thank you for releasing the Flash-WAM code and paper!

I am trying to understand the rationale behind the timestep-pair sampling strategy used for consistency distillation.

From the current implementation, my understanding is:

  1. The starting timestep index is sampled uniformly from all 1,000 training timesteps for each frame:
    distillation/data.py

  2. The stride is computed as:

    k = num_train_timesteps // num_ddim_timesteps

    distillation/trainer.py

  3. The endpoint is then selected by advancing k scheduler indices and clamping to the final index:

    end_ids = (start_ids + k).clamp(max=num_train_timesteps - 1)

    This is done for both video and action:
    distillation/step.py

With the released configuration,

num_train_timesteps = 1000
num_ddim_timesteps = 2

we have k = 500:

distillation/config.py

This produces pairs such as:

0   -> 500
100 -> 600
499 -> 999
500 -> 999
700 -> 999
900 -> 999
999 -> 999

Therefore, 501 out of the 1,000 possible starting indices map to the same terminal index 999. In addition, start_id = 999 produces a zero-length pair.

This seems different from the standard LCM training implementation. LCM first constructs a reduced DDIM solver grid—commonly 50 DDIM timesteps—and samples adjacent pairs from that grid. The number of teacher/DDIM discretization steps is therefore not necessarily the same as the final number of student inference steps:

I am particularly curious about the following points:

  1. Why are starting timesteps sampled from all 1,000 scheduler indices instead of from an LCM-style reduced grid, such as 50 DDIM anchors?

  2. Is num_ddim_timesteps=2 intentionally tied to the target inference NFE?
    In LCM, the DDIM solver resolution used to construct training pairs can be much larger than the final student sampling NFE.

  3. Is the terminal concentration intentional?
    With k=500, approximately 50.1% of all starting indices are clamped to end_id=999. Does this act as an intentional clean-boundary weighting, or is it mainly an implementation simplification?

  4. Why is the interval defined in scheduler-index space rather than sigma space?
    Flash-WAM uses different SNR shifts for video and action (shift=5 vs. shift=1), so the same 500-index stride corresponds to very different and highly non-uniform Δσ values.

    For example, approximately:

    Video, shift=5:
    index 0   -> 500: sigma 1.000 -> 0.833
    index 500 -> 999: sigma 0.833 -> 0.005
    
    Action, shift=1:
    index 0   -> 500: sigma 1.000 -> 0.500
    index 500 -> 999: sigma 0.500 -> 0.001
    
  5. Does a single teacher Euler step remain sufficiently accurate for the largest sigma intervals?
    Since the Euler target error generally grows with the interval size, I wonder whether a finer teacher grid, multiple teacher substeps, or a higher-order solver was considered.

  6. Did you compare this strategy with any of the following alternatives?

    • LCM-style 50-step DDIM anchor pairs;
    • sampling exact pairs used by the final 1-step/2-step inference schedule;
    • phase-aligned endpoints, as in Phased Consistency Models;
    • continuous or arbitrary timestep pairs, as in ECT or CTM;
    • sampling based on a fixed Δσ instead of a fixed scheduler-index stride.
  7. Could you share the theoretical or empirical basis for this timestep sampler?
    In particular, was it inspired by a previous consistency-distillation or flow-matching paper? If so, could you point to the relevant paper, section, or implementation?

Some potentially related works I have been comparing against are:

I may be misunderstanding the scheduler indexing direction or the intended relationship between the training stride and inference NFE, so please correct me if that is the case.

Thank you for your time and for open-sourcing the project!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions