DRS-GUI adds a training-free dynamic region search stage before coordinate prediction. Its UI Perceptor uses OmniParser V2 to parse GUI elements and INSTRUCTOR to match those elements to the user instruction. An MCTS Action Planner then schedules three reversible field-of-view actions:
- Focus contracts the view around instruction-relevant elements.
- Shift moves the view to a distant relevant region.
- Scatter expands the view to recover missing context.
The best region found by MCTS is passed to a base grounding model. This release supports the two model families used in the paper: Qwen2.5-VL-7B-Instruct and UGround-V1-7B.
- Python 3.10
- NVIDIA GPU with CUDA
Create the environment from the repository root:
conda create -n drsgui python=3.10 -y
conda activate drsgui
# Install a CUDA build of PyTorch that matches your driver first.
# The following is only an example for CUDA 12.1.
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txtVerify that PyTorch can see the GPU:
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"The second value must be True. If your cluster already provides a suitable PyTorch build, keep that build and install the remaining requirements normally.
DRS-GUI/
├── checkpoints/ # downloaded model weights (gitignored)
├── data/ # downloaded screenshots (gitignored)
├── outputs/ # predictions and metrics (gitignored)
└── drsgui/
├── scripts/
│ ├── run_screenspot_v1.sh
│ ├── run_screenspot_v2.sh
│ └── run_screenspot_pro.sh
└── src/
├── infer.py # single-screenshot inference
├── run.py # benchmark evaluation
├── screenspot_data.py # v1/v2/Pro schema normalization
├── ui_perceptor.py
├── models/ # Qwen2.5-VL and UGround adapters
└── policies/drsgui/ # MCTS planner and three actions
Model weights, screenshots, generated crops, logs, and outputs are intentionally excluded from Git.
Install the Hugging Face CLI through requirements.txt, then create the local model directory:
mkdir -p checkpoints
# Optional for gated/private repositories: hf auth loginDRS-GUI needs both the icon detector and the Florence-2 icon caption model from microsoft/OmniParser-v2.0:
hf download microsoft/OmniParser-v2.0 \
--local-dir checkpoints/OmniParser-v2.0The two paths supplied at runtime will be:
checkpoints/OmniParser-v2.0/icon_detect/model.pt
checkpoints/OmniParser-v2.0/icon_caption/
hf download hkunlp/instructor-large \
--local-dir checkpoints/instructor-largeYou may also pass the Hub ID hkunlp/instructor-large directly, but a local path makes offline runs reproducible.
At least one of the following is required:
# Qwen2.5-VL used with --model-type qwen2_5vl
hf download Qwen/Qwen2.5-VL-7B-Instruct \
--local-dir checkpoints/Qwen2.5-VL-7B-Instruct
# UGround-V1 used with --model-type ugroundv1
hf download osunlp/UGround-V1-7B \
--local-dir checkpoints/UGround-V1-7BAfter downloading all optional weights, the relevant structure is:
checkpoints/
├── OmniParser-v2.0/
│ ├── icon_detect/model.pt
│ └── icon_caption/
├── instructor-large/
├── Qwen2.5-VL-7B-Instruct/
└── UGround-V1-7B/
The --images argument must point to the directory relative to which every annotation's img_filename can be resolved. Pass the directory containing the benchmark JSON files through --annotations.
The canonical ScreenSpot v1 release is maintained in the SeeClick repository. The following Hugging Face mirror keeps the original images/ and annotations/ directory layout and is convenient for command-line download:
hf download benwiesel/ScreenSpot \
--repo-type dataset \
--local-dir data/screenspot_v1Expected paths:
data/screenspot_v1/images/<image files>
data/screenspot_v1/annotations/screenspot_desktop.json
data/screenspot_v1/annotations/screenspot_mobile.json
data/screenspot_v1/annotations/screenspot_web.json
If you use the original SeeClick download instead, pass its image and annotation directories directly.
Download the official OS-Copilot/ScreenSpot-v2 release and extract the image archive:
hf download OS-Copilot/ScreenSpot-v2 \
--repo-type dataset \
--local-dir data/screenspot_v2
unzip data/screenspot_v2/screenspotv2_image.zip -d data/screenspot_v2Expected paths:
data/screenspot_v2/screenspotv2_image/<image files>
data/screenspot_v2/screenspot_desktop_v2.json
data/screenspot_v2/screenspot_mobile_v2.json
data/screenspot_v2/screenspot_web_v2.json
This repository includes the ScreenSpot-Pro annotation JSON files. Download the official likaixin/ScreenSpot-Pro dataset to obtain the screenshots:
hf download likaixin/ScreenSpot-Pro \
--repo-type dataset \
--local-dir data/screenspot_proUse data/screenspot_pro/images as --images and data/screenspot_pro/annotations as --annotations.
The official v1 and v2 boxes are [x, y, width, height]; the loader converts them to the internal [x1, y1, x2, y2] representation. ScreenSpot-Pro already uses [x1, y1, x2, y2].
For custom preprocessed v1/v2 JSON in [x1, y1, x2, y2] format, add "bbox_format": "xyxy" to each record to prevent conversion. Missing IDs, image sizes, applications, platforms, and UI types are filled when possible.
The examples below use Qwen2.5-VL and locally downloaded weights.
CUDA_VISIBLE_DEVICES=0 bash drsgui/scripts/run_screenspot_v1.sh \
--model-type qwen2_5vl \
--model-path checkpoints/Qwen2.5-VL-7B-Instruct \
--images data/screenspot_v1/images \
--annotations data/screenspot_v1/annotations \
--detector-path checkpoints/OmniParser-v2.0/icon_detect/model.pt \
--caption-model checkpoints/OmniParser-v2.0/icon_caption \
--instructor-model checkpoints/instructor-large \
--output outputs/qwen2_5vl_screenspot_v1.jsonCUDA_VISIBLE_DEVICES=0 bash drsgui/scripts/run_screenspot_v2.sh \
--model-type qwen2_5vl \
--model-path checkpoints/Qwen2.5-VL-7B-Instruct \
--images data/screenspot_v2/screenspotv2_image \
--annotations data/screenspot_v2 \
--detector-path checkpoints/OmniParser-v2.0/icon_detect/model.pt \
--caption-model checkpoints/OmniParser-v2.0/icon_caption \
--instructor-model checkpoints/instructor-large \
--output outputs/qwen2_5vl_screenspot_v2.jsonCUDA_VISIBLE_DEVICES=0 bash drsgui/scripts/run_screenspot_pro.sh \
--model-type qwen2_5vl \
--model-path checkpoints/Qwen2.5-VL-7B-Instruct \
--images data/screenspot_pro/images \
--annotations data/screenspot_pro/annotations \
--detector-path checkpoints/OmniParser-v2.0/icon_detect/model.pt \
--caption-model checkpoints/OmniParser-v2.0/icon_caption \
--instructor-model checkpoints/instructor-large \
--output outputs/qwen2_5vl_screenspot_pro.jsonTo evaluate UGround-V1, change only these two options:
--model-type ugroundv1
--model-path checkpoints/UGround-V1-7B
Useful evaluation options:
--annotations /path/to/jsons: directory containing the benchmark JSON files.--task screenspot_web_v2: run one JSON file; use comma-separated stems for multiple files.--mcts-iterations 8: number of MCTS simulations for each sample.--max-depth 3: maximum action-search depth.--num-chunks N --chunk-idx K: evaluate chunkKofN; give every chunk a different--outputpath.
Each output JSON contains per-sample predictions plus overall, text/icon, platform, application, and group metrics. Grounding accuracy is point-in-bounding-box accuracy.
For a screenshot that has no benchmark annotation, provide the screenshot and the GUI element instruction directly:
CUDA_VISIBLE_DEVICES=0 python drsgui/src/infer.py \
--image examples/example.png \
--instruction "click the Settings button" \
--platform windows \
--application unknown \
--model-type qwen2_5vl \
--model-path checkpoints/Qwen2.5-VL-7B-Instruct \
--detector-path checkpoints/OmniParser-v2.0/icon_detect/model.pt \
--caption-model checkpoints/OmniParser-v2.0/icon_caption \
--instructor-model checkpoints/instructor-large \
--output outputs/example_prediction.jsonplatform and application are optional hints used to choose a domain-specific semantic prompt. Supported platform examples include windows, macos, linux, web, ios, and android. The returned pred field is the predicted [x, y] point in the original screenshot's pixel coordinate system.
DRS-GUI is released under the Apache License 2.0. The retained OmniParser utility code is covered by its own license in drsgui/src/OmniParser/LICENSE. Dataset and model weights remain subject to their respective upstream licenses.