Hello, author~ I tried testing Lyra 2.0 on the DL3DV test set. The test images and the prompts I used are shown below, but the results are quite poor. Does the model have any additional requirements for the input image resolution or the prompts? I noticed that the code mentions using scripts/gemini_caption.py to generate captions, but I could not find the corresponding script. I would appreciate any guidance or suggestions you could provide.
Input Image:

Using Prompts:
A traditional Chinese courtyard pavilion with wooden architecture, gray brick walls, gray tiled roof, dark brown wooden columns, a raised platform with wooden stairs and geometric wooden railings, intricate blue-green and turquoise painted beams and eaves. The camera slowly walks through the quiet courtyard, approaching the stairs, moving closer to the wooden columns and railings, exploring the architectural details and layered spatial composition. Soft natural daylight, overcast sky, calm atmosphere, ultra realistic, cinematic, highly detailed, smooth camera movement, realistic materials and perspective.
Final Result:
https://github.com/user-attachments/assets/c69cd105-add1-483c-adb8-92dd41115992
Hello, author~ I tried testing Lyra 2.0 on the DL3DV test set. The test images and the prompts I used are shown below, but the results are quite poor. Does the model have any additional requirements for the input image resolution or the prompts? I noticed that the code mentions using
scripts/gemini_caption.pyto generate captions, but I could not find the corresponding script. I would appreciate any guidance or suggestions you could provide.Input Image:

Using Prompts:
Final Result:
https://github.com/user-attachments/assets/c69cd105-add1-483c-adb8-92dd41115992