ks2048 · 87 points · 23 comments · 昨日 · Open original
Comments
4 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
TOtopwalktown昨日
"Overall, we attempted to download 130M videos and achieved a link success rate of approximately 60%, resulting in 80M successfully retrieved videos with a total duration of 10M hours."
I am astonished that the success rate is so high. How Youtube didn't block them, I don't know. But I think that this URL list won't age well because youtube will very quickly block any researcher trying to download these videos themselves.
VIvivzkestrel昨日
- can someone with expertise give us an overview of the architecture involved doing this
- let us say you ran yt-dlp inside python aiohttp
- surely your ll run a limit soon as your ip address will be flagged
- what solutions do we have to auto rotate proxies in python
- are there better, faster and more reliable ways to go about downloading a 100 million videos without getting your ip address blocked?
VOvoidUpdate昨日
Damn, that's a lot of videos for them to contact the creators and ask for permission to use their content as machine learning training content. Unless of course, they didn't, and just went ahead with it anyway...
Comments
4 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
"Overall, we attempted to download 130M videos and achieved a link success rate of approximately 60%, resulting in 80M successfully retrieved videos with a total duration of 10M hours." I am astonished that the success rate is so high. How Youtube didn't block them, I don't know. But I think that this URL list won't age well because youtube will very quickly block any researcher trying to download these videos themselves.
- can someone with expertise give us an overview of the architecture involved doing this - let us say you ran yt-dlp inside python aiohttp - surely your ll run a limit soon as your ip address will be flagged - what solutions do we have to auto rotate proxies in python - are there better, faster and more reliable ways to go about downloading a 100 million videos without getting your ip address blocked?
Damn, that's a lot of videos for them to contact the creators and ask for permission to use their content as machine learning training content. Unless of course, they didn't, and just went ahead with it anyway...
[deleted]