- Home
- Skills
- Workflow Automation
- desktop vision controller
Works with the AI tools you already use
desktop vision controller
Automate Windows UI interactions using visual capture, OCR, and coordinate-based clicks with safety confirmations.
Free
desktop vision controller
Example session with this skill installed
Capture the window titled 'Calculator', find the number '8' using OCR ( no incluyed ) , and click it. I confirm this action.
- Read your context and instructions
- Compiled the desktop vision controller
I have successfully captured the Calculator window, identified the coordinates for the '8' button, and performed the click as requested.
Desktop Action Log
| Step | Action | Target | Result | Coordinates/Path |
|---|---|---|---|---|
| 1 | capture | window:Calculator | Success | captures/calc_snap_17158200.png |
| 2 | ocr | captures/calc_snap_17158200.png | Text Found | "8" detected at 120, 450 |
| 3 | click_at | 120,450 | Success | User Confirmed: True |
Next steps
- Verify if the Calculator display now shows the number '8'.
- Specify the next digit or operator you want to click.
- Use
type_textif you need to input a full equation string instead of individual clicks.
Connects securely to your tools. The creator never sees your data.
About this skill
The problem
Automating desktop applications without accessible APIs or DOM elements is slow and error-prone. Developers struggle to bridge the gap between AI reasoning and physical UI interaction on Windows systems.
What it does
- Captures high-resolution screenshots of the full desktop, specific windows, or defined coordinate regions.
- Performs OCR on captured images to identify text labels and UI elements for interaction.
- Executes precise mouse clicks and text input based on pixel coordinates or text labels.
- Provides a list of active window titles to target specific applications for automation.
Frameworks & tools
Windows Desktop Environment, MCP (Model Context Protocol), desktop_vision_controller MCP server.
Why this beats prompting it yourself
Generic prompts often hallucinate UI coordinates or fail to handle the safety implications of OS-level actions. This skill enforces a structured schema, mandatory user confirmation for clicks/typing, and standardizes error handling for offline desktop automation.
Use cases
- Automating data entry into legacy Windows software that lacks an API.
- Performing visual regression testing or UI audits on desktop applications.
- Building AI-driven tutorials that interact with real software in real-time.
- Extracting text from images or non-selectable UI components via OCR.
Known limitations
Requires the desktop_vision_controller MCP server to be installed and running. Text input must not be used for sensitive passwords.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 14 days ago
- Free to download with an account