- Home
- Skills
- Agents & Orchestration
- video-analyzer
Works with the AI tools you already use
Video Analyzer
Transform raw video files into structured JSON and Markdown reports with local character detection and transcription.
Free
video-analyzer
Example session with this skill installed
Analyze latest_demo.mp4. I need to know who appears in it, what they said, and a summary of the visual timeline for my project report.
- Read your context and instructions
- Compiled the video-analyzer
Analysis complete. I've generated latest_demo_report.md and latest_demo_report.json.
Key Highlights
- Characters: 2 (Presenter, Interviewee)
- Transcript: Discussing the new API v2 features.
- Visuals: Starts with screen share of terminal, transitions to 2-person split view.
Connects securely to your tools. The creator never sees your data.
About this skill
Converts video files into structured Markdown and JSON reports using local preprocessing and Claude Vision.
This Claude Skill acts as a bridge for LLM environments (such as Claude Projects) that do not natively support direct video uploads. It extracts audio transcripts, identifies key visual transitions, groups frames by color composition, and structures this data into AI-friendly and human-readable formats.
Architectural Overview
The skill operates through a decoupled multi-stage pipeline designed to minimize API token usage while capturing critical structural and textual details:
- Input Video Processing: The source file is ingested by the preprocessor script to decouple audio and visual data.
- Audio Analysis: The extracted track is converted and transcribed locally using a high-efficiency Whisper engine.
- Frame Selection: Visual frames are extracted at a scaled resolution and clustered based on color metrics to remove redundancy.
- Multimodal Synthesis: Claude Vision analyzes the optimized frame selection to track character data and visual state.
- Report Generation: A final compilation stage bundles metadata, timestamps, transcriptions, and visual observations into unified JSON and Markdown outputs.
Technical Specifications & Pipeline Stages
Stage 1: Extraction & Preprocessing (process_video.py)
The pipeline begins by isolating the audio track and extracting downscaled video frames.
- Frame Extraction: Frames are extracted at a configurable rate (default: 0.5 fps, or one frame every 2 seconds) and capped at a maximum count (default: 30 frames) to keep processing efficient.
- Resolution Scaling: Frames are scaled so that their longest edge is at most 1568px while preserving the original aspect ratio. This keeps the visual token budget predictable (approximately 1,400 tokens per frame).
- Audio Extraction & Transcription: The audio track is resampled to mono 16kHz. It is then transcribed locally using OpenAI Whisper (defaulting to the highly efficient
large-v3-turbomodel). The engine supports both
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
6 installs
Downloaded by developers to date
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 3 months ago
- Free to download with an account