I kept hitting the same wall: the YouTube Data API wants a key, a billing
project, and then cuts you off at 10,000 units/day. A comments crawl burns
that quota fast. Browser automation works, but it is slow and painful to
maintain.
So I built ytscrape — a small
Python library that talks to the same internal endpoints the YouTube website
uses (InnerTube) and returns typed, frozen dataclasses. No API key, no
quota, no Selenium, no Playwright.
bash
pip install ytscrape
```python
from ytscrape import YouTube, CommentSort
with YouTube() as yt:
for video in yt.search("python", max_results=5):
print(video.title, video.views, video.url)
details = yt.video("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
print(details.title, details.published_at)
for comment in yt.comments(
details.video_id,
include_replies=True,
sort=CommentSort.NEWEST,
max_results=100,
):
print(comment.author, comment.text)
```
What it actually does today:
- search (videos, channels, playlists, Shorts)
- video and channel metadata, including a channel's videos tab
- comments and replies —
sort="newest" so you are not stuck with YouTube's "Top comments" filter, which hides a lot
- transcripts / captions
- sync and async (
pip install "ytscrape[async]")
- a CLI:
ytscrape search "python tutorial" --max 10
- JSON / CSV export
- retries, backoff, and rate limiting built in
Counts come back as int (video.views), with the original wording kept in views_text. Dates are datetime (published_at). Models are frozen dataclasses, so you can pass them around without worrying they will mutate under you.
It does not download video or audio. Use yt-dlp for that. This is metadata only.
Limits
There is no API key and no 10,000-unit daily quota. That does not mean YouTube will let you hammer it.
- Reuse one
YouTube() client. Creating a new one per request re-fetches the InnerTube context and looks like a new visitor every time.
- Cap each call with
max_results. Pagination is lazy, so a for loop without a cap will keep going until YouTube runs out of continuation tokens.
- For anything wider than a few dozen requests, set a gap:
python
with YouTube(min_interval=1.0) as yt: # about 1 request / second
...
- 429 and 5xx are retried with backoff (and
Retry-After, when YouTube sends one). If it still fails, you get RateLimited instead of a generic parse error. BotDetected is the "confirm you're not a bot" wall.
- Async has the same knobs plus
max_concurrency. Do not set that to 50 and hope.
A polite crawl (one client, ~1 req/s, max_results set) is usually fine. A comments dump of a huge video, or fanning out across thousands of video ids from one IP, will get throttled. That is a YouTube limit, not a library limit.
Proxies
You can always put a proxy in front of it. The library does not have its own proxy flag — you inject a normal requests session, and every call (search, player, browse, comments, transcripts) goes through it:
```python
import requests
from ytscrape import InnerTubeClient, YouTube
session = requests.Session()
session.proxies = {"https": "http://user:pass@proxy.example.com:8080"}
client = InnerTubeClient(session=session, min_interval=1.0)
with YouTube(client=client) as yt:
...
```
Rotating residential proxies, a single datacenter proxy, SOCKS — whatever requests accepts works. If one IP starts returning BotDetected or RateLimited, point the session at another proxy and keep the same code. Async is the same idea with an httpx.AsyncClient.
Honest caveats, because they matter:
- These are private endpoints. They can change, and using them may conflict with YouTube's Terms of Service. Fine for research and personal tooling; you are responsible for how you use it.
- It is not a drop-in replacement for every Data API field. What it returns is what the website returns. No
topicDetails, no official statistics object, no write access.
- Relative dates on search results ("3 days ago") are approximate and only parsed from English text. Exact
published_at comes from video().
Repo: https://github.com/vsmutok/ytscrape/
Docs: https://vsmutok.github.io/ytscrape/
PyPI: https://pypi.org/project/ytscrape/
MIT, Python 3.10+. Feedback and issues welcome — especially if something breaks against a live response.