Skip to content

[Question] Recommended prompt & coordinate protocol for VLM grounding evaluation #109

Description

@Yjonben

Context

We are evaluating the VLM / language-generation path of wall-oss-0.5 (Wall-X MoE VLA; we use official WallxModelWrapper.infer_vqa() language path) (not the action policy rollout) on embodied grounding benchmarks via an internal evaluator (RefCOCO-family bbox, Where2Place point, RoboSpatial-Home point/yes-no, etc.).

We want to align prompts and coordinate interpretation with the officially intended format, but could not find a complete grounding evaluation protocol in the README / VLMEvalKit wrapper / paper.

Questions

  1. For VLM-only inference (eval_in_vqa=True or equivalent), what is the recommended user prompt template for:
    • referring expression bbox (RefCOCO-style)
    • point / free-space pointing (Where2Place, RoboSpatial-Home context)
  2. What output format should we expect and parse?
    • JSON (bbox_2d / point_2d), tags (<box>, <point>), or free-form coordinates?
  3. What coordinate system should predictions use?
    • normalized 0–1, normalized 0–1000, or absolute pixel (original or resized image)?
  4. Is there an official grounding eval script or VLMEvalKit config we should reuse instead of custom prompts?

Our current setup (for reference)

  • Backend: official repo wrapper, language generation only
  • Prompt: neutral wording like "Report the bounding box / point coordinates" (no explicit range), because official docs do not specify one
  • Scoring: we try norm01 / norm1000 / pixel based on smoke outputs

Smoke observations (may be wrong due to prompt mismatch)

  • Official grounding uses <box> / <point> tags; reverse_grounding_points() expects resize-aware pixel coords
  • RefCOCO/Where2Place grounding scores near 0 even with tag prompts + pixel scoring
  • Non-grounding MCQ (RoboBench, MindCube) works on smoke
  • Confirm: is infer_vqa() the correct entry for grounding eval, and what tag/coordinate convention should evaluators use?

Could you confirm the correct protocol, or point us to the official eval entry / prompt file? Thank you!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions