spotify-downloader under the hood: spotify to mp3 via youtube

at 17 i built spotify-downloader, a website for downloading songs and playlists from spotify. i posted it on a subreddit and in three days it had over 400,000 views and 2,000 active users. vercel's free tier served more than 100 gb in hours. i switched to serving it from home, within another 48 hours the traffic had passed 1 tb, and i shut the public service down. since then, anyone who wants to use it has to host their own instance.
finding the spotify song on youtube
spotify doesn't give you the audio, only the data. the site asked the spotify api for it with app credentials, without the user having to log in: title, artists, album, cover art and duration in milliseconds. with that it searched for the song on youtube.
the 2024 version searched for ${track.name} by ${track.artists[0].name} official, took the first five videos and picked the one with the same duration as the spotify song or, if there wasn't one, the closest.
rereading the code, the first check almost never passed. the search library worked out the duration from the text youtube shows, "3:33", and converted it to milliseconds, so it always ended in three zeros. spotify gives real milliseconds. they only matched when the spotify song lasted an exact number of seconds, so in practice the rule was "the closest of the five".
today's code runs a different search and allows a margin:
query := fmt.Sprintf("ytsearch5:%s - %s lyrics", artist, title)
// ...
for _, video := range candidates {
diff := math.Abs(video.Duration - targetSeconds)
if diff < 1.5 {
return video.ID, nil
}
if diff < shortestDiff {
shortestDiff = diff
bestVideoID = video.ID
}
}
it now looks for the lyrics video instead of the official one. of the five results, in youtube's order, the first one within a second and a half of the spotify duration wins. if none falls inside that, the closest wins. a music video with a twenty-second intro never gets through the filter, and only gets picked if the other four are even further off.
the mp3 that wasn't an mp3
what came down from youtube was audio inside an mp4 container. the first versions just gave it a .mp3 extension, and i wrote that down myself in an issue on the repository, #2. to write the title, artist and cover art into it, it had to be properly converted with ffmpeg.
on 1 march 2024 the server did the conversion, with fluent-ffmpeg. the next day i moved it into the browser with ffmpeg.wasm, and browser-id3-writer wrote the tags. on 6 march i added a fast mode that handed over the .m4a as it came, without converting or tagging it, because converting inside the browser was slow.
that version had another bug, which shows up when you reread the code. the track number was read from track.disk_number, a field the spotify api doesn't have, because it's called disc_number. every mp3 came out with the text "undefined" as its track number.
today's code lets yt-dlp extract the mp3 and then makes a second pass with ffmpeg and -c:a copy, which rewrites the tags without re-encoding the audio. before that it wipes youtube's metadata with -map_metadata -1, so everything in the file comes from spotify: title, artists, album, date, track number out of the total, disc, isrc, label and copyright. the cover art gets the type Cover (front), because windows explorer doesn't show covers left with the default type. and high quality doesn't promise 320 kbps, because youtube's original audio rarely goes above 160.
why it went down
in the public version, every song was a request to the server. the vercel function downloaded the whole audio from youtube into memory and sent it back as the response. for a playlist, the browser asked for ten at once, and when a batch finished it asked for ten more.
in other words, every byte a user downloaded came into my server and went back out again. outgoing traffic was the sum of every song everyone downloaded. moving ffmpeg into the browser on 2 march took compute work off the server, but not a single byte of traffic.
in issue #3 i wrote that vercel was complaining about bandwidth, with about 141 gb sent to users in half a day. the fix i suggested was to send the stream straight to the user instead of downloading it to the server first, and at the end i added "not even sure if you can send a stream through an api...". i closed it three hours later.
at home there was no free-tier limit, but the maths was the same. every song still went in and out through my server, and now the server was my home connection. in two days the traffic passed 1 tb.
the kill switch
i shut the public service down for two reasons. my home network was saturated and, more importantly, i realised what it meant to be handling that volume of copyrighted music downloads. from then on, the project became a tool that each person installs on their own machine.
the change of model fixes the traffic without touching a single line of the search code. if every user has their own instance, the audio goes from youtube to the machine of whoever asked for it, and my server is no longer in the middle. according to the product document in the repository, a small group of friends and family uses it.
the rewrite in go
on 21 september 2026 i replaced all the old code with a rewrite i'd started in a separate repository on 5 december 2025. it's now a go backend with a next.js frontend, behind nginx, in three containers.
each download is a job with an id. it goes into a queue, progress reaches the browser over a websocket, and at the end the server packs a zip, even when it's a single song. jobs survive a page reload, because the browser stores them and reconnects to them.

an instance needs spotify api credentials, and according to the readme, spotify now only lets you create them from a premium account. the 30-second preview for each song is almost always disabled, because spotify withdrew preview_url for most apps at the end of 2024.
and there's a global limit. MAX_CONCURRENT_DOWNLOADS sets how many songs can be searched, downloaded and tagged at the same time across the whole server, whether they come from one job or fifty, and the default is 5. in the code it's a go channel with five slots shared by every job. in the version that moved 1 tb, each browser asked for ten per playlist and the server had no ceiling at all.