Describe Rename Sound Files

    by Fredrik Akerstrom

    4

    SoundTag AI: Automatically describe and batch-rename audio files based on their actual sound using local ML or Gemini AI.

    Secure checkout via Stripe

    0 installsSecurity scanned

    Works with the AI tools you already use

    CClaude CodeCCursorCCodex CLIGGitHub CopilotGGemini CLIVVS CodeWWindsurfMManus+14 more

    See it in action

    You say

    I have a folder of audio files with messy names like 'ElevenLabs' timestamps and 'Epic_trumpet' IDs. Can you analyze what they actually sound like and rename them to be descriptive?

    Your agent does

    AST Model (local, instant): ElevenLabs_2024-10-14T02_58_09_Eric_pre_s23_sb19_se69_b_m2.mp3 → Elevenlabs_Eric_Voice_Speech.mp3 Epic_trumpet_birthda_#1-1772723461703.wav → Bright_Trumpet_Horn_Fanfare.wav AST + Gemini two-step (best quality): Lots_Snow_From_Knock.mp3 → Heavy_Wooden_Knocking.mp3 Rover_Car_Wheele_Vehicle_Aircraft.mp3 → Jet_Turbine_Spin_Up.mp3 Electronic_Glitch_Owl.mp3 → Digital_Radio_Static.mp3

    What you get

    Identify specific instruments to rename generic music project tracksConvert cryptic field recording names into descriptive environmental labelsOrganize voiceover exports by speaker name and performance styleBatch-process sound effect libraries using AI-generated content tags

    About this skill

    SoundTag AI -Listens & Renames Your Sound Files

    What it does

    This skill solves the problem of messy, auto-generated audio filenames like audio_track_v2_final_99.wav. It analyzes the actual content of sound files and renames them with human-readable, descriptive titles such as ElevenLabs_2024-07-21T15_43_56_George_pre_s50_sb75_se0_b_m2.mp3 → Elevenlabs_George_Voice_Speech.mp3 or Bright_Trumpet_Fanfare.wav or Large_Crowd_Cheering.mp3.

    Supported tools

    • Local ML (AST): Uses the MIT Audio Spectrogram Transformer to classify sounds into 527 categories (Speech, Music, Explosion, etc.) entirely offline.
    • Google Gemini API: Leverages advanced multimodal AI for nuanced descriptions of cinematic SFX, moods, and complex textures.
    • Batch Processing: Supports .wav, .mp3, .ogg, .flac, .aac, .m4a, and more.

    Why use this skill

    Unlike simple prompting, this skill implements a sophisticated two-step workflow. It first attempts a high-speed local classification to save on API costs and privacy. For ambiguous sounds, it provides a structured "improvement pass" using Gemini. It intelligently combines ML labels with hidden hints from the original filename to ensure context is never lost. It handles environment constraints automatically, including specific dependency versions (Torch/Transformers) to fit within sandboxed resource limits.

    Output

    The result is a clean, organized directory where every sound file follows a consistent Title_Case_With_Underscores naming convention, making your sample libraries and field recordings instantly searchable.

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 4 months ago

    Listed4 months ago

    Creator

    F
    Fredrik Akerstrom

    Claude Skills

    15 skills on Agensi

    Freelancer, internet entrepreneur and music producer from Sweden.

    Frequently Asked Questions

    Popular in Productivity