- Home
- Skills
- Productivity
- describe-rename-sound-files
Describe Rename Sound Files
SoundTag AI: Automatically describe and batch-rename audio files based on their actual sound using local ML or Gemini AI.
$10
Works with the AI tools you already use
describe-rename-sound-files
Example session with this skill installed
I have a folder of audio files with messy names like 'ElevenLabs' timestamps and 'Epic_trumpet' IDs. Can you analyze what they actually sound like and rename them to be descriptive?
- Read your context and instructions
- Compiled the describe-rename-sound-files
AST Model (local, instant):
ElevenLabs_2024-10-14T02_58_09_Eric_pre_s23_sb19_se69_b_m2.mp3 → Elevenlabs_Eric_Voice_Speech.mp3
Epic_trumpet_birthda_#1-1772723461703.wav → Bright_Trumpet_Horn_Fanfare.wav
AST + Gemini two-step (best quality):
Lots_Snow_From_Knock.mp3 → Heavy_Wooden_Knocking.mp3
Rover_Car_Wheele_Vehicle_Aircraft.mp3 → Jet_Turbine_Spin_Up.mp3
Electronic_Glitch_Owl.mp3 → Digital_Radio_Static.mp3
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
SoundTag AI -Listens & Renames Your Sound Files
What it does
This skill solves the problem of messy, auto-generated audio filenames like audio_track_v2_final_99.wav. It analyzes the actual content of sound files and renames them with human-readable, descriptive titles such as
ElevenLabs_2024-07-21T15_43_56_George_pre_s50_sb75_se0_b_m2.mp3 → Elevenlabs_George_Voice_Speech.mp3 or
Bright_Trumpet_Fanfare.wav or Large_Crowd_Cheering.mp3.
Supported tools
- Local ML (AST): Uses the MIT Audio Spectrogram Transformer to classify sounds into 527 categories (Speech, Music, Explosion, etc.) entirely offline.
- Google Gemini API: Leverages advanced multimodal AI for nuanced descriptions of cinematic SFX, moods, and complex textures.
- Batch Processing: Supports .wav, .mp3, .ogg, .flac, .aac, .m4a, and more.
Why use this skill
Unlike simple prompting, this skill implements a sophisticated two-step workflow. It first attempts a high-speed local classification to save on API costs and privacy. For ambiguous sounds, it provides a structured "improvement pass" using Gemini. It intelligently combines ML labels with hidden hints from the original filename to ensure context is never lost. It handles environment constraints automatically, including specific dependency versions (Torch/Transformers) to fit within sandboxed resource limits.
Output
The result is a clean, organized directory where every sound file follows a consistent Title_Case_With_Underscores naming convention, making your sample libraries and field recordings instantly searchable.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 6 months ago
- Passed all security checks, Safe to install