Context
We are evaluating the VLM / language-generation path of wall-oss-0.5 (Wall-X MoE VLA; we use official WallxModelWrapper.infer_vqa() language path) (not the action policy rollout) on embodied grounding benchmarks via an internal evaluator (RefCOCO-family bbox, Where2Place point, RoboSpatial-Home point/yes-no, etc.).
We want to align prompts and coordinate interpretation with the officially intended format, but could not find a complete grounding evaluation protocol in the README / VLMEvalKit wrapper / paper.
Questions
- For VLM-only inference (
eval_in_vqa=True or equivalent), what is the recommended user prompt template for:
- referring expression bbox (RefCOCO-style)
- point / free-space pointing (Where2Place, RoboSpatial-Home context)
- What output format should we expect and parse?
- JSON (
bbox_2d / point_2d), tags (<box>, <point>), or free-form coordinates?
- What coordinate system should predictions use?
- normalized 0–1, normalized 0–1000, or absolute pixel (original or resized image)?
- Is there an official grounding eval script or VLMEvalKit config we should reuse instead of custom prompts?
Our current setup (for reference)
- Backend: official repo wrapper, language generation only
- Prompt: neutral wording like "Report the bounding box / point coordinates" (no explicit range), because official docs do not specify one
- Scoring: we try norm01 / norm1000 / pixel based on smoke outputs
Smoke observations (may be wrong due to prompt mismatch)
- Official grounding uses
<box> / <point> tags; reverse_grounding_points() expects resize-aware pixel coords
- RefCOCO/Where2Place grounding scores near 0 even with tag prompts + pixel scoring
- Non-grounding MCQ (RoboBench, MindCube) works on smoke
- Confirm: is
infer_vqa() the correct entry for grounding eval, and what tag/coordinate convention should evaluators use?
Could you confirm the correct protocol, or point us to the official eval entry / prompt file? Thank you!
Context
We are evaluating the VLM / language-generation path of wall-oss-0.5 (Wall-X MoE VLA; we use official
WallxModelWrapper.infer_vqa()language path) (not the action policy rollout) on embodied grounding benchmarks via an internal evaluator (RefCOCO-family bbox, Where2Place point, RoboSpatial-Home point/yes-no, etc.).We want to align prompts and coordinate interpretation with the officially intended format, but could not find a complete grounding evaluation protocol in the README / VLMEvalKit wrapper / paper.
Questions
eval_in_vqa=Trueor equivalent), what is the recommended user prompt template for:bbox_2d/point_2d), tags (<box>,<point>), or free-form coordinates?Our current setup (for reference)
Smoke observations (may be wrong due to prompt mismatch)
<box>/<point>tags;reverse_grounding_points()expects resize-aware pixel coordsinfer_vqa()the correct entry for grounding eval, and what tag/coordinate convention should evaluators use?Could you confirm the correct protocol, or point us to the official eval entry / prompt file? Thank you!