A Model Context Protocol (MCP) server that provides autonomous GUI automation capabilities to LLM clients (like Claude Desktop). This tool allows an AI to interact directly with your computer's screen, mouse, keyboard, and browser.
- Computer Vision & OCR: Extract text from screen regions and find UI elements visually.
- Mouse & Keyboard Control: Perform clicks, typing, hotkeys, and precise movements.
- Window Management: Focus, minimize, maximize, and query open windows.
- System Automation: Manage processes, execute system commands, adjust volume, and check system health.
- Browser Automation: Full web browsing capabilities (Selenium) for navigating, finding elements, extracting text, and clicking.
- File System: Read/write files, list directories, download files.
- Python 3.12 or higher
- Windows (recommended for PyAutoGUI compatibility)
You can install the package directly from PyPI (https://pypi.org/project/echo-mcp/):
pip install echo-mcpYou can automatically configure your AI apps and IDEs (Google Antigravity, Claude Desktop, Cursor, Windsurf, VS Code / Cline) to use the MCP server by running:
python -m echo_mcp.configTo configure all supported apps automatically without interactive prompts:
python -m echo_mcp.config --allThis setup wizard configures your client settings files (e.g. ~/.gemini/config/mcp_config.json for Antigravity). Restart your app/IDE after configuration for the new tools to take effect.
The server provides a comprehensive suite of tools spanning multiple categories:
computer: Unified action tool supportingscreenshot,mouse_move,left_click,right_click,middle_click,double_click,triple_click,left_click_drag,mouse_down,mouse_up,type(Unicode clipboard paste),key,hotkey,scroll, andcursor_position. Includes an automatic action-observation feedback loop returning updated screenshots.- Multimodal Screen Perception: High-visibility pointer crosshair rendered on captured screenshots, Windows DPI auto-scaling, and native MCP
Imageoutput format.
mouse_click,mouse_double_click,move_mouse,mouse_drag,mouse_scroll,get_mouse_positiontype_text,press_key,hotkey
take_screenshot(native FastMCP Image output with cursor overlay),get_screen_size,get_text_at_coords,find_text_on_screen,click_element
list_all_windows,focus_window,maximize_window,minimize_window,get_window_geometry,close_all_windows_by_title,get_active_window_infolist_processes,kill_process,get_process_stats,wait_for_process
system_power,set_volume,get_system_health,get_disk_usage,ping,list_network_interfaceslaunch_application,get_environment_variable,get_clipboard,set_clipboard,wait
list_directory,read_file_content,write_to_file,delete_file,create_directory,download_file
browser_open,browser_close,browser_navigate,browser_get_url,browser_get_title,browser_get_page_sourcebrowser_find_element,browser_click,browser_fill_form,browser_extract_textbrowser_screenshot,browser_scroll,browser_execute_script,browser_wait_for_element
Giving an LLM control over your mouse and keyboard is powerful but risky. Only run this server with prompts you trust, and never leave the automation unattended while active.