Hello!
I modified the code to launch it with the VLM model: "Qwen3-VL-8B-Instruct-FP8".
I have results where the avatars perform actions. But when it comes to render them, they perform the actions in random locations or poses. For example, a simulated human wants to turn on the tv by sitting "inside" the fridge, or sometimes, it walks to a random place and does something weird, like painting in the air (image attached). Also, no objects appear in their hands. Nevertheless, the environment is loaded.
Is this because the simulation is not rendering the environment properly? Or is this an issue given by the limitations of the VLM?
Thanks!
Hello!
I modified the code to launch it with the VLM model: "Qwen3-VL-8B-Instruct-FP8".
I have results where the avatars perform actions. But when it comes to render them, they perform the actions in random locations or poses. For example, a simulated human wants to turn on the tv by sitting "inside" the fridge, or sometimes, it walks to a random place and does something weird, like painting in the air (image attached). Also, no objects appear in their hands. Nevertheless, the environment is loaded.
Is this because the simulation is not rendering the environment properly? Or is this an issue given by the limitations of the VLM?
Thanks!