Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    desktop vision controller

    1

    Automate Windows UI interactions using visual capture, OCR, and coordinate-based clicks with safety confirmations.

    Free

    0 installsSecurity scanned
    desktop vision controller

    desktop vision controller

    Example session with this skill installed

    Capture the window titled 'Calculator', find the number '8' using OCR ( no incluyed ) , and click it. I confirm this action.

    • Read your context and instructions
    • Compiled the desktop vision controller

    I have successfully captured the Calculator window, identified the coordinates for the '8' button, and performed the click as requested.

    Desktop Action Log

    StepActionTargetResultCoordinates/Path
    1capturewindow:CalculatorSuccesscaptures/calc_snap_17158200.png
    2ocrcaptures/calc_snap_17158200.pngText Found"8" detected at 120, 450
    3click_at120,450SuccessUser Confirmed: True

    Next steps

    • Verify if the Calculator display now shows the number '8'.
    • Specify the next digit or operator you want to click.
    • Use type_text if you need to input a full equation string instead of individual clicks.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    The problem

    Automating desktop applications without accessible APIs or DOM elements is slow and error-prone. Developers struggle to bridge the gap between AI reasoning and physical UI interaction on Windows systems.

    What it does

    • Captures high-resolution screenshots of the full desktop, specific windows, or defined coordinate regions.
    • Performs OCR on captured images to identify text labels and UI elements for interaction.
    • Executes precise mouse clicks and text input based on pixel coordinates or text labels.
    • Provides a list of active window titles to target specific applications for automation.

    Frameworks & tools

    Windows Desktop Environment, MCP (Model Context Protocol), desktop_vision_controller MCP server.

    Why this beats prompting it yourself

    Generic prompts often hallucinate UI coordinates or fail to handle the safety implications of OS-level actions. This skill enforces a structured schema, mandatory user confirmation for clicks/typing, and standardizes error handling for offline desktop automation.

    Use cases

    • Automating data entry into legacy Windows software that lacks an API.
    • Performing visual regression testing or UI audits on desktop applications.
    • Building AI-driven tutorials that interact with real software in real-time.
    • Extracting text from images or non-selectable UI components via OCR.

    Known limitations

    Requires the desktop_vision_controller MCP server to be installed and running. Text input must not be used for sensitive passwords.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 14 days ago

    • Free to download with an account

    Listed14 days ago

    Frequently Asked Questions