- Home
- Skills
- Social Media
- YouTube Video to Structured Notes Agent
More screenshots
Works with the AI tools you already use
YouTube Video to Structured Notes Agent
It is built for users who need substantially more than a generic AI summary.
$9.99
YouTube Video to Structured Notes Agent
Example session with this skill installed
YOUTUBE VIDEO
URL
https://www.youtube.com/watch?v=EXAMPLE123
Output Language
English
Note Depth
Deep
Video Type
Founder Podcast
Expected Duration
Approximately 2 hours
REQUIRED OUTPUT
Metadata
Yes
TL;DR:
Yes
Clickable Key Timestamps
Yes
Topic Breakdown
Yes
Official Chapters
Use if available
Inferred Topics
Create if official chapters are unavailable or insufficient
People Mentioned
Yes
Companies Mentioned
Yes
Tools / Products
Yes
Books / Papers / Resources
Yes
Frameworks
Yes
Data Points
Yes
Examples / Case Studies
Yes
Mistakes
Yes
Decisions
Yes
Actionable Takeaways
Yes
Contradictions
Yes
Questions / Uncertainties
Yes
Caption Quality
Yes
Visual Dependencies
Yes
JSON
No
TRANSCRIPT RULES
Prefer manually authored captions when available.
If only auto-generated captions exist
Use them but flag likely errors.
Pay special attention to
Names
Companies
Dollar Amounts
Percentages
Years
Technical Terms
ATTRIBUTION
For business metrics, market statistics, scientific claims, financial numbers, predictions, or strong factual assertions:
Attribute the statement to the speaker.
Do not convert speaker claims into verified facts.
TIMESTAMPS
Use clickable timestamps for
Major Arguments
Important Data Points
Examples
Frameworks
Recommendations
Contradictions
Important Stories
Use the earliest relevant timestamp that clearly supports the note.
LONG VIDEO PROCESSING
Do not create a shallow summary.
Chunk the transcript into manageable semantic sections with overlap.
Analyze each section separately.
Then reconcile all sections globally.
Remove duplicate ideas.
Merge repeated entities.
Preserve contradictions.
Preserve chronology.
SPONSOR SEGMENTS
Identify sponsor segments when reasonably clear.
Do not treat sponsor claims as independent recommendations.
VISUAL COVERAGE
Flag moments where the speaker refers to
Charts
Slides
Code
Diagrams
Screen Demonstrations
UI
Visual Comparisons
Do not invent visual details from transcript text.
FINAL STRUCTURE
Title
Channel
Duration
Published Date
URL
Transcript Type
Transcript Language
Chapter Status
Source Quality
TL;DR
Key Timestamped Points
Topic Breakdown
People Mentioned
Companies
Tools / Products
Books / Papers / Resources
Frameworks
Quantitative Claims
Examples / Case Studies
Mistakes & Lessons
Decisions
Contradictions / Tensions
Actionable Takeaways
Questions / Uncertainties
Caption Quality Notes
Visual Dependencies
Final Source-Quality Assessment
IMPORTANT
Do not reproduce the entire transcript.
Prefer concise paraphrase.
Clearly identify
VIDEO CLAIM
VERIFIED EXTERNALLY
UNCERTAIN CAPTION
where appropriate.
- Read your context and instructions
- Compiled the youtube video to
Building a $100M Software Company Without a Traditional Sales Team
Channel
Founder Conversations
Duration
2:03:41
Published
August 12, 2026
Transcript
Auto-Generated Captions
Transcript Language
English
Official Chapters
Available
Visual Coverage
Transcript-based analysis. Charts, slides, code, UI demonstrations, and other visual-only information were not independently analyzed.
=== SOURCE QUALITY ===
Transcript Completeness
HIGH
Caption Quality
MEDIUM-HIGH
Speaker Attribution
HIGH
Number Reliability
MEDIUM
Visual Dependency
LOW-MEDIUM
Overall Note Confidence
MEDIUM-HIGH
=== TL;DR ===
- The guest argues that product-led growth worked because the company reduced the time between a user's first interaction and first measurable result.
- Early growth came primarily from small technical teams rather than enterprise buyers.
- The company initially avoided building a large sales organization and instead invested heavily in onboarding, documentation, and product instrumentation.
- The guest says one major strategic mistake was adding enterprise complexity before product-market fit was strong enough.
- Hiring decisions were based more heavily on ownership and execution speed than traditional credentials.
- Several important growth figures are presented as the guest's own recollection and are not independently verified in this note set.
- The strongest recurring operating principle is to shorten feedback cycles across product, customer support, and engineering.
=== KEY TIMESTAMPED POINTS ===
[03:42] The guest explains why the company originally rejected a traditional sales-first growth model.
Why It Matters
This frames the operating philosophy discussed throughout the episode.
Attribution
Guest statement.
[11:18] The guest says the company focused on reducing time-to-value before investing heavily in acquisition.
Key Idea
Acquisition efficiency mattered less if new users could not reach a meaningful product outcome quickly.
[24:51] The guest describes the first major product-market-fit signal.
The company reportedly saw a sharp increase in users inviting colleagues into existing workspaces.
Status
VIDEO CLAIM
[37:09] The guest says approximately 70% of early growth came from referrals and internal team invitations.
Status
VIDEO CLAIM
Caption Confidence
MEDIUM
Reason
The percentage is clearly spoken once but is not repeated elsewhere.
[51:22] The guest describes the decision to build enterprise features earlier than the company was operationally ready to support.
Lesson
Revenue opportunity can create roadmap pressure before the product has sufficient organizational maturity.
[1:08:44] The discussion shifts to hiring.
The guest says the strongest early hires frequently had broad ownership rather than narrowly defined responsibilities.
[1:34:06] The guest describes shortening the distance between customer complaints and engineering decisions.
Key Principle
Customer-support data should become product input rather than remain isolated inside a support function.
[1:51:38] The guest explains why the company eventually added a formal enterprise sales organization.
This is important because it qualifies the earlier criticism of traditional sales.
The guest's position is not
"Sales teams are unnecessary."
It is closer to
"Sales infrastructure should match the product's maturity and customer segment."
=== TOPIC BREAKDOWN ===
[00:00–16:10] Company Origins and Product-Led Growth
The guest explains how the company's initial go-to-market model emerged.
Major Themes
Product-Led Growth
Fast User Activation
Low-Friction Onboarding
Team Invitations
Minimal Early Sales Infrastructure
[16:10–35:40] Product-Market Fit
The conversation focuses on behavioral indicators the team used instead of relying only on revenue.
Important indicators included
Repeat Usage
Invitations
Workspace Expansion
Support Questions Becoming More Advanced
[35:40–57:15] Early Scaling Mistakes
The guest describes several mistakes
Premature Enterprise Features
Too Many Product Exceptions
Weak Internal Documentation
Insufficient Instrumentation
[57:15–1:18:30] Hiring and Team Design
The guest emphasizes
High Ownership
Cross-Functional Capability
Execution Speed
Short Decision Chains
[1:18:30–1:42:50] Customer Feedback and Product Operations
Major theme
Reduce the time required for user feedback to reach the people making product decisions.
[1:42:50–2:03:41] Enterprise Growth and Final Lessons
The guest explains why the company later built more conventional enterprise systems.
The episode closes with operating principles around:
Timing
Organizational Complexity
Hiring
Customer Feedback
Iteration Speed
=== PEOPLE MENTIONED ===
Paul Graham
Context
Mentioned while discussing startup advice and early product validation.
First Mention
[18:04]
Confidence
HIGH
Peter Thiel
Context
Mentioned in a discussion about distribution.
First Mention
[44:29]
Confidence
HIGH
Possible Name
Transcript
"patrick collinson"
Likely
Patrick Collison
Context
Stripe is discussed immediately afterward.
Confidence
MEDIUM
Caption Note
Possible auto-caption normalization.
=== COMPANIES MENTIONED ===
Stripe
Linear
Slack
Notion
Atlassian
These companies are primarily used as comparison examples.
They should not be interpreted as endorsements unless explicitly described that way.
=== TOOLS / PRODUCTS / RESOURCES ===
Notion
Context
The guest says the early team used Notion for internal documentation.
Timestamp
[1:02:11]
Sentiment
POSITIVE / PERSONAL USE
Linear
Context
Discussed as an example of opinionated product design.
Timestamp
[1:16:09]
Sentiment
POSITIVE / EXAMPLE
Slack
Context
Used in an example about communication overload.
Timestamp
[1:27:42]
Sentiment
CRITICAL IN SPECIFIC CONTEXT
=== QUANTITATIVE CLAIMS ===
Claim
Approximately 70% of early growth came from referrals and internal invitations.
Timestamp
[37:09]
Speaker
Guest
Status
VIDEO CLAIM
Confidence
MEDIUM
Claim
The company reportedly reached its first 10,000 active users in under one year.
Timestamp
[29:54]
Speaker
Guest
Status
VIDEO CLAIM
Confidence
MEDIUM-HIGH
Claim
Enterprise contracts eventually represented more than half of revenue.
Timestamp
[1:55:20]
Speaker
Guest
Status
VIDEO CLAIM
Confidence
MEDIUM
=== MISTAKES & LESSONS ===
- Building Enterprise Complexity Too Early
Timestamp
[51:22]
Mistake
Adding customer-specific enterprise requirements before the organization could support them efficiently.
Lesson
Revenue opportunity can introduce complexity faster than operational maturity develops.
- Weak Product Instrumentation
Timestamp
[56:48]
Mistake
The company could see total usage but struggled to understand why specific cohorts retained.
Lesson
Track behavior that explains value realization, not only aggregate activity.
- Documentation Lag
Timestamp
[1:04:37]
Mistake
Internal knowledge existed primarily in conversations.
Lesson
Rapid hiring magnifies undocumented operational knowledge.
=== IMPORTANT DECISION ===
Decision
Delay building a large sales organization.
Reason
The company believed the product could acquire and activate small teams directly.
Outcome
The model reportedly worked during early growth but became insufficient when larger enterprise customers became strategically important.
Relevant Timestamps
[03:42]
[1:51:38]
=== POTENTIAL TENSION ===
Earlier
[03:42]
The guest strongly criticizes sales-first growth for the company's early stage.
Later
[1:51:38]
The guest describes building a substantial enterprise sales function.
Interpretation
The two statements are not necessarily contradictory.
The guest appears to argue that organizational systems should match company stage and buyer segment.
=== ACTIONABLE TAKEAWAYS ===
START
Measure time-to-value explicitly.
Source
[11:18]
START
Track behavioral product-market-fit indicators such as collaboration and repeat use rather than relying only on acquisition.
Source
[24:51]
TEST
Create a direct mechanism that regularly exposes product teams to recurring customer-support patterns.
Source
[1:34:06]
AVOID
Adding customer-specific complexity before the organization can support it operationally.
Source
[51:22]
MEASURE
Determine whether documentation quality deteriorates as hiring velocity increases.
Source
[1:04:37]
=== CAPTION QUALITY NOTES ===
Potential Proper-Name Error
[44:31]
Caption
"patrick collinson"
Likely
Patrick Collison
Confidence
MEDIUM
Potential Number Sensitivity
[37:09]
The guest appears to say
70%
The figure is not repeated.
Treat as
VIDEO CLAIM / MEDIUM CONFIDENCE
=== VISUAL DEPENDENCIES ===
[48:13]
The speaker says
"you can see the retention curve here."
The transcript does not describe the underlying values.
Result
VISUAL DEPENDENCY
The chart itself was not analyzed.
[1:20:48]
The guest refers to
"this diagram"
The transcript does not provide enough information to reconstruct its structure.
=== FINAL SOURCE ASSESSMENT ===
The transcript provides strong coverage of the conversation's major arguments and chronology.
The primary uncertainties are
Several Numerical Claims
One or More Proper Names
Two Visual References
The notes reliably represent what the speakers discuss, but company metrics and external factual claims remain VIDEO CLAIMS unless separately verified.
These notes are primarily transcript-based. Visual-only information may not be fully represented.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
YouTube Video to Structured Notes Agent is a premium long-form video research, transcript-analysis, and knowledge-extraction agent designed to transform YouTube videos into structured, navigable, source-aware notes.
It is built for users who need substantially more than a generic AI summary.
Instead of reducing a 90-minute interview or 2-hour podcast to a few vague bullets, the agent creates a reusable knowledge artifact that preserves:
Video Metadata Main Thesis TL;DR Clickable Timestamps Key Arguments Topic Sections Official Chapters Inferred Topic Boundaries Speaker Attribution People Mentioned Companies Mentioned Tools and Products Books and Resources Studies and Papers Frameworks Examples Case Studies Quantitative Claims Recommendations Predictions Contradictions Actionable Takeaways Uncertainties Caption Errors Visual Dependencies
The canonical workflow is:
YouTube URL → Validate Video → Retrieve Metadata → Retrieve Available Captions / Transcript → Normalize Transcript → Detect Chapters → Detect Topic Boundaries → Chunk Long Content → Analyze Each Chunk → Extract Claims / Entities / Data → Reconcile Across Chunks → Build Timestamped Notes → Add Quality Warnings → Add Visual Coverage Disclosure → Produce Final Knowledge Artifact
The agent is designed for:
YouTube Videos Podcasts Interviews Lectures Webinars Tutorials Coding Videos Business Videos Founder Interviews Educational Content Research Discussions Panel Discussions Debates Long-Form Conversations Archived Livestreams
The core differentiator is specificity.
A weak summarizer may output:
"The video discusses productivity and focus."
This agent instead aims for findings such as:
[12:48] The guest argues that the first 90 minutes of the workday should be protected from meetings and reactive communication because this is when the team's highest-value work is usually completed.
Why It Matters: This becomes the foundation for the workflow proposed later in the conversation.
Attribution: Guest statement.
Verification: Not independently verified.
The agent therefore optimizes for:
Structure Navigation Attribution Evidence Awareness Compression Traceability
VIDEO METADATA
When available, the agent can collect:
Title Video ID Canonical URL Channel Creator Upload Date Duration Description Language Caption Availability Caption Type Official Chapters
Optional metadata can include:
View Count Like Count Tags Categories
Dynamic metrics such as views and likes should not dominate evergreen research notes unless explicitly requested.
The agent distinguishes metadata provenance.
Possible classifications include:
PLATFORM METADATA USER-SUPPLIED DERIVED UNKNOWN
TRANSCRIPT SOURCE
The preferred transcript hierarchy is:
- User-Supplied Human Transcript
- Manually Authored Captions
- Creator-Uploaded Subtitles
- Auto-Generated YouTube Captions
- Authorized Speech-to-Text Fallback
The agent should always disclose which transcript source was used.
Possible classifications:
MANUAL AUTO-GENERATED USER-SUPPLIED ASR-FALLBACK UNKNOWN
This matters because auto-generated captions are significantly more error-prone for:
Names Numbers Acronyms Technical Terms Product Names Book Titles Companies Code Foreign Words Overlapping Speakers
TRANSCRIPT NORMALIZATION
Before analysis, the agent can normalize:
Subtitle Line Breaks Duplicated Caption Segments Overlapping Subtitle Text Repeated Words Timestamp Structure Speaker Labels Basic Punctuation
Normalization must not materially rewrite what the speaker said.
The transcript remains the evidentiary basis for the notes.
AUTO-CAPTION ERROR DETECTION
The skill actively looks for probable caption errors.
Possible classifications include:
CAPTION UNCERTAINTY POSSIBLE NAME ERROR POSSIBLE NUMBER ERROR POSSIBLE TECHNICAL TERM ERROR POSSIBLE SPEAKER ATTRIBUTION ERROR
Example:
Transcript: "sam alt men"
Likely Interpretation: Sam Altman
Confidence: Medium
Possible Auto-Caption Correction: Yes
The agent should not silently correct ambiguous transcript text.
NUMERICAL CLAIMS
Numbers receive additional scrutiny because speech recognition can distort:
Percentages Years Dollar Amounts Dates Measurements User Counts Revenue Growth Rates Model Numbers Quantities
Example:
[23:10] Growth figure mentioned by the guest.
Transcript Quality: The captions could represent either 15% or 50%.
Result: NUMBER UNCERTAIN
The agent should not choose one arbitrarily.
SPEAKER ATTRIBUTION
When speaker identities are available, important claims should be attributed.
Example:
[34:12] The guest says the company reduced onboarding time by 40%.
Attribution: Guest statement.
Verification: VIDEO CLAIM — not independently verified.
The skill distinguishes:
The speaker says X.
from:
X is objectively true.
This is especially important for:
Scientific Claims Financial Claims Medical Claims Political Claims Legal Claims Business Metrics Market Statistics Predictions
CLAIM CLASSIFICATION
Useful statements can be classified as:
FACTUAL CLAIM OPINION PERSONAL EXPERIENCE PREDICTION RECOMMENDATION ANALOGY EXAMPLE DEFINITION DATA POINT QUOTED THIRD-PARTY CLAIM
This creates substantially better research notes than treating every sentence as equivalent.
TL;DR
The TL;DR should answer:
What is the video fundamentally about?
What is the main thesis?
What are the most important conclusions?
Typical output:
Short Video: 2–4 bullets
Medium Video: 4–7 bullets
Long Podcast: 5–10 bullets
The TL;DR should not simply repeat metadata or chapter titles.
KEY POINTS
Important points should include:
Clickable Timestamp Specific Idea Context Why It Matters Optional Speaker Attribution Optional Verification Note Optional Confidence
Example:
[18:26] Constraints can improve creative execution.
The guest argues that deliberately reducing available options early in a project reduces decision fatigue and accelerates implementation.
Why It Matters: This principle becomes one of the operating rules proposed later in the episode.
TOPIC BREAKDOWN
The agent can produce hierarchical topic structure.
Example:
- Why the current workflow fails
- The proposed operating model 2.1 Deep-work blocks 2.2 Communication windows 2.3 Weekly review
- Team implementation
- Failure cases
- Final recommendations
When official YouTube chapters exist:
Use Them
but do not simply copy chapter titles.
Summarize the actual content inside each chapter.
If no official chapters exist:
Create Inferred Topic Sections
and clearly label them:
INFERRED TOPIC SECTIONS
Never present AI-generated topic boundaries as official creator chapters.
CLICKABLE TIMESTAMPS
Important notes should use clickable YouTube timestamps.
The timestamp should lead directly to the relevant moment.
Possible format:
[12:48]
linked to the canonical video URL with the appropriate timestamp.
For videos under one hour:
MM:SS
For longer videos:
H:MM:SS
The timestamp should correspond to the earliest transcript segment that clearly supports the note.
PEOPLE MENTIONED
The agent can extract:
Name Role or Context First Mention Timestamp Reason Mentioned Caption Confidence
Example:
Andrew Huberman
Context: Mentioned as an example of long-form educational content.
First Mention: [41:09]
Confidence: High
Ambiguous names should remain flagged.
TOOLS, SOFTWARE, PRODUCTS, AND RESOURCES
The skill can extract:
Software Apps Hardware Platforms Frameworks Products Books Papers Studies Datasets Methodologies
For each resource it can record:
Name Category Context Timestamp Speaker Sentiment Confidence
Speaker sentiment may be:
RECOMMENDED POSITIVE NEUTRAL CRITICAL COMPARED MERELY MENTIONED
A casual mention should not automatically become a recommendation.
SPONSORED TOOLS
When a tool appears inside a clearly sponsored segment, the agent should identify that context.
Example:
Sponsor Segment: [12:20–13:25]
The promoted product should not automatically be categorized as an independent recommendation.
DATA POINTS
The agent can extract:
Percentages Revenue Costs Durations Dates Growth Figures User Counts Benchmarks Experimental Results Quantities Measurements
Recommended structure:
Value Context Speaker Timestamp Attribution Verification Status Confidence
Possible verification statuses:
VIDEO CLAIM VERIFIED EXTERNALLY UNCERTAIN CAPTION
Without external research, the default is:
VIDEO CLAIM
EXAMPLES AND CASE STUDIES
Useful examples can be structured as:
Case Context Result Lesson Timestamp
Personal anecdotes should remain clearly identified as personal experience rather than universal evidence.
FRAMEWORK EXTRACTION
When a speaker introduces a framework, the agent can extract:
Framework Name Purpose Components Sequence Example Timestamp
The agent should not invent a formal framework name unless the speaker provides one.
ACTIONABLE TAKEAWAYS
Potential action categories include:
START STOP CONTINUE TEST MEASURE RESEARCH
Actions should be specific.
Weak:
"Work harder."
Strong:
"Batch reactive communication into two scheduled windows so the first uninterrupted work block remains protected."
The agent should not convert every descriptive observation into advice.
CONTRADICTIONS AND TENSIONS
Long interviews frequently contain nuanced or apparently contradictory statements.
The agent can surface them.
Example:
[18:02] The guest argues that speed matters more than polish.
[1:22:19] The guest later says that shipping unfinished work can damage trust.
Potential Interpretation: The statements may apply to different stages of product development, but the distinction is not explicitly stated.
The agent should preserve contradictions rather than forcing artificial consistency.
VISUAL DEPENDENCY
Transcript-based analysis cannot fully reconstruct information shown visually.
Potential visual dependencies include:
Charts Graphs Slides Diagrams Code Screen Recordings Software Demonstrations Product Interfaces Gestures Objects Before/After Comparisons Visual Measurements
If the speaker says:
"As you can see here..."
or:
"Look at this chart..."
the agent should flag:
VISUAL DEPENDENCY
Example:
[27:41] VISUAL DEPENDENCY
The speaker refers to a chart, but its actual values are not sufficiently described by the transcript.
The agent should never pretend it analyzed a chart when only transcript text was available.
The final notes should normally contain a clear disclosure such as:
Visual Coverage: These notes are based primarily on transcript/caption text and metadata. On-screen diagrams, demonstrations, charts, code, slides, UI interactions, gestures, and other visual-only information may not be fully represented.
LONG-FORM VIDEO PROCESSING
The skill is specifically designed for long videos.
A 2-hour podcast should not be reduced to a shallow summary simply because the transcript is large.
The preferred processing model is:
Transcript → Semantic Chunks → Per-Chunk Notes → Entity Extraction → Claim Extraction → Data Extraction → Global Reconciliation
Example conceptual chunking:
00:00–20:00 18:00–38:00 36:00–56:00 54:00–1:14:00
Short overlap helps prevent important ideas from being lost at chunk boundaries.
Each chunk can contain:
Chunk ID Start Time End Time Summary Topics People Tools Data Points Claims Takeaways Uncertainties
After all chunks are analyzed, the agent performs global reconciliation.
This includes:
Merge Duplicate Topics Resolve Repeated Names Combine Tool Mentions Deduplicate Takeaways Preserve Chronology Preserve Contradictions Merge Repeated Claims Normalize Entities Rank Important Themes
This step is essential.
The agent should never simply concatenate independent chunk summaries.
PODCAST MODE
For podcasts and interviews, the agent can emphasize:
Host Questions Guest Thesis Recurring Themes Stories Mistakes Recommendations Personal Experience Disagreements Named Resources Business Lessons
LECTURE MODE
For academic or educational videos, it can emphasize:
Definitions Concept Hierarchy Formulas Examples Dependencies Misconceptions Likely Exam-Worthy Material
TUTORIAL MODE
For tutorials, it can produce:
Prerequisites Step-by-Step Process Tools Commands Expected Results Warnings Troubleshooting
If exact code is only visible on screen and is not present in the transcript:
The agent must state:
ON-SCREEN CODE NOT FULLY AVAILABLE FROM TRANSCRIPT
and must not invent the code.
BUSINESS / FOUNDER MODE
Can emphasize:
Company History Strategy Metrics Mistakes Hiring Operations Tools Processes Market Claims Decision-Making Operating Principles
RESEARCH MODE
Can emphasize:
Hypotheses Evidence Referenced Studies Quantitative Claims Evidence Strength Open Questions Claims Requiring Verification
DEBATE / PANEL MODE
The skill keeps viewpoints separate.
Possible structure:
Speaker A Position Speaker B Position Moderator Questions Shared Ground Primary Disagreements
It must not blend opposing claims into one synthetic position.
EXECUTIVE BRIEF MODE
For users who need fast review:
Main Thesis 5–10 Critical Points Important Numbers Decisions Actions Risks
STUDY MODE
Can additionally generate:
Definitions Concept Map Study Questions Flashcards Memory Cues
KNOWLEDGE-BASE MODE
Can generate reusable Markdown metadata such as:
Source Video ID Title Channel Date Duration Topics People Tools
followed by structured notes.
JSON OUTPUT
When requested, the agent can produce structured machine-readable output containing:
metadata tldr key_points topics people tools data_points takeaways warnings
Key point objects can contain:
timestamp_seconds timestamp_label timestamp_url summary speaker importance confidence
Data-point objects can contain:
value unit context speaker timestamp_seconds status confidence
TRANSCRIPT COMPLETENESS
The agent can inspect the transcript for:
Missing Beginning Missing Ending Long Timestamp Gaps Abrupt Cutoffs Repeated Blocks
Example:
Transcript Gap: 42:18–48:05
This should be disclosed.
OVERALL SOURCE QUALITY
The agent can summarize source quality using:
Transcript Type Caption Quality Visual Dependency Transcript Completeness Speaker Attribution Quality Overall Note Confidence
Possible overall confidence levels:
HIGH MEDIUM-HIGH MEDIUM LOW
NO TRANSCRIPT AVAILABLE
If a video has no usable captions and no authorized transcription path exists:
The correct output is:
Transcript: Unavailable
Spoken-Content Analysis: Not Performed
Metadata-only notes can still be provided.
The agent must never fabricate spoken content.
AUTHORIZED SPEECH-TO-TEXT FALLBACK
Where the environment and access rights permit, a speech-to-text workflow can be used.
Possible pipeline:
Authorized Media → Audio Extraction → Speech-to-Text → Timestamp Normalization → Notes
The transcript source must then be labeled:
ASR-FALLBACK
ASR may struggle with:
Accents Crosstalk Jargon Names Numbers Code Music
PROCESSING RELIABILITY
For long videos, the agent can maintain processing states such as:
CREATED METADATA_READY TRANSCRIPT_READY NORMALIZED CHUNKING ANALYZING RECONCILING COMPLETE PARTIAL FAILED
Checkpointing can preserve:
Metadata Transcript Chunk Boundaries Completed Chunk Results
This allows long analyses to resume without reprocessing everything.
ERROR TAXONOMY
The skill supports a structured error model:
YTN-001 INVALID_URL YTN-002 VIDEO_UNAVAILABLE YTN-003 PRIVATE_OR_RESTRICTED YTN-004 METADATA_UNAVAILABLE YTN-005 TRANSCRIPT_UNAVAILABLE YTN-006 CAPTION_LANGUAGE_UNSUPPORTED YTN-007 TRANSCRIPT_INCOMPLETE YTN-008 CAPTION_QUALITY_LOW YTN-009 TIMESTAMP_GAP YTN-010 CHUNK_ANALYSIS_FAILED YTN-011 ENTITY_AMBIGUOUS YTN-012 NUMBER_UNCERTAIN YTN-013 SPEAKER_UNCERTAIN YTN-014 VISUAL_DEPENDENCY YTN-015 TOOLING_UNAVAILABLE YTN-016 NETWORK_UNAVAILABLE YTN-017 RATE_LIMITED YTN-018 PROCESS_TIMEOUT YTN-019 SOURCE_CHANGED YTN-020 INTERNAL_ERROR
Possible severity levels:
INFO WARNING HIGH BLOCKING
BASH + NETWORK WORKFLOW
The preferred execution environment includes:
Bash Network Access Python
When available, command-line utilities can retrieve public metadata and available subtitle files.
The workflow can then use Python for:
Subtitle Parsing Timestamp Normalization Chunking Entity Deduplication Timestamp Link Construction Markdown Rendering JSON Rendering Quality Validation
Command-line tools should always be checked for availability rather than assumed.
Temporary working directories, bounded timeouts, exit-code checks, quoted paths, and safe filenames should be used.
Transcript text must never be executed as shell commands.
COPYRIGHT-AWARE SUMMARIZATION
The purpose of the skill is knowledge extraction.
It should not reproduce entire transcripts.
Prefer paraphrased summaries.
When a direct quote adds value:
Keep It Short Attribute It Include Timestamp
The final artifact should remain substantially more compressed than the source transcript.
ACCESS CONTROL
The agent processes legitimate public or authorized content.
It should not bypass:
Private Video Restrictions Membership Restrictions Paywalls Authentication Geographic Controls Creator Dashboard Access Protected Systems
QUALITY STANDARD
A high-quality final result should allow a user who did not watch the entire video to:
Understand the Main Thesis Navigate to Important Moments See Major Topics Identify Important People and Tools Capture Quantitative Claims Understand What the Speaker Actually Said Distinguish Claims from Verified Facts Recognize Caption Uncertainty Recognize Missing Visual Information Reuse the Notes for Research, Learning, or Decision-Making
That is the core standard of the skill.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 17 days ago
- Passed all security checks, Safe to install