← back to blog

Telegram bot frameworks that survive production

Most Telegram bot frameworks look fine in a demo. You clone a repo, drop in a token, send /start, get a reply back. The gap shows up later, when the bot has been running for three weeks, the host rebooted twice, a user sent 40 messages in ten seconds, and you need to know whether the thing is still holding state correctly. That’s the part nobody tests before launch, and it’s the part that actually determines whether a bot survives production.

We host Telegram bots and userbots for a living, so this isn’t theoretical. Here’s what actually separates frameworks that hold up from ones that fall over quietly.

What “surviving production” actually means

A bot in production has to do four things reliably: stay connected to Telegram, keep track of what it already told each user, respond within Telegram’s rate limits without getting flood-controlled, and come back cleanly after a crash or restart. None of these are exotic requirements, but a lot of frameworks handle one or two of them well and quietly skip the rest, usually the state and restart parts, because those don’t show up in a quick demo.

If you’re picking a framework, the question isn’t “can it send a message.” Every framework can send a message. The question is what happens to an in-flight conversation when the process dies mid-handler, or when your host rotates the container and the bot comes back up thirty seconds later.

Long polling vs webhooks, and why it matters for hosting

Telegram gives you two ways to receive updates: long polling with getUpdates, or a webhook where Telegram pushes updates to a public HTTPS endpoint you control.

Long polling is the simpler mode. Your bot opens a connection to api.telegram.org, holds it open for up to the timeout you set, and gets a batch of updates back (or an empty response, and it asks again). This is easy to run on a single VPS because it needs no inbound port and no certificate. The catch is that Telegram only allows one active getUpdates consumer per bot token at a time. Run two instances of the same bot by accident, say during a rolling deploy where the old process hasn’t exited yet, and you’ll get 409 Conflict errors and duplicate or dropped updates. This is one of the most common production bugs we see, and it’s a deployment problem, not a framework bug.

Webhooks flip that around. Telegram pushes each update to your endpoint as it happens, which means lower latency and no polling loop, but now you need a valid TLS certificate, a stable public address, and a process that can actually receive inbound traffic (which rules out running behind certain restrictive NAT or CGNAT setups without a reverse proxy). Webhooks also mean you’re now running a web server, with everything that implies for error handling on malformed requests.

Neither mode is objectively better. Polling is simpler to host and easier to debug locally. Webhooks scale better once you’re handling meaningful volume, because you’re not holding an open connection per bot. Most of the frameworks below support both, but the ecosystem and tooling around each mode varies more than people expect.

The framework landscape, and what it’s actually built on

For bots that talk to the Bot API (not full MTProto user accounts), the common choices are python-telegram-bot and aiogram on the Python side, and Telegraf or grammY on the Node side. All four are wrappers around the same underlying HTTP API, so functionally they can all do the same things. The differences are in concurrency model and how much they do for you around state and middleware.

aiogram and grammY are async-first and were designed around Python’s asyncio and JavaScript’s event loop respectively from day one. python-telegram-bot moved to full async support a few versions back, but plenty of tutorials and forks still float around from its synchronous, threaded era, which is worth checking before you build on top of an old example. Telegraf has been the default choice in the Node ecosystem for years and has a large middleware library, though grammY has picked up ground because it was built with TypeScript typing in mind from the start.

If you’re building a userbot, i.e. a bot that logs in as a real Telegram account and speaks the full MTProto protocol instead of the Bot API, you’re in a different category entirely: Telethon and Pyrogram, both Python. These aren’t wrappers around api.telegram.org; they implement the same protocol Telegram’s own client apps use, which means they can do things bot accounts can’t, like reading full chat history or joining channels the way a normal user does. That power comes with real account-safety implications, which is a separate topic from framework choice.

Async vs sync, and why it matters once you have real traffic

A synchronous framework processes one update at a time per worker thread. That’s fine for a single bot answering a handful of messages an hour. It becomes a problem the moment one handler makes a slow outbound call, like hitting a third-party API or writing to a database that’s under load, because every other user’s message queues up behind it. An async framework can be doing that slow I/O for one user while handling another user’s message in the same event loop, which is the difference between a bot that feels responsive and one that goes silent for a few seconds under any real load.

This matters more the more bots you’re running on one host. We run many bots on shared infrastructure, and the ones built on sync frameworks with blocking calls buried in handlers are reliably the ones that need a restart when something upstream gets slow.

State and session persistence

Every framework gives you some notion of “state” for a conversation, whether it’s called a finite state machine, a context object, or a conversation handler. The default in most of them is in-memory: a Python dict or an object living in process memory, keyed by chat ID. That’s fast and simple, and it’s also the single biggest cause of bots that “lose” users. Restart the process, rebuild the container, or run more than one instance, and every in-memory state disappears with it.

The fix is the same across frameworks: back the state store with something that survives a restart, typically Redis for speed or a small database table if you want it queryable. aiogram, python-telegram-bot, and grammY all support pluggable storage backends for exactly this reason, but the plugin isn’t wired up by default, so it’s easy to ship a bot that works fine in testing and then silently forgets every user the first time the host reboots it.

Rate limits and flood control

Telegram enforces limits on how fast a bot can send messages: roughly one message per second to a given chat, and a higher aggregate ceiling across different chats, with the exact numbers subject to change and enforced server-side regardless of what any framework tells you. Push past it and you get a 429 Too Many Requests response with a retry_after value telling you how long to back off.

What separates frameworks in practice is whether they handle that response for you or just let the exception bubble up. Most modern libraries have some built-in retry-after handling, but broadcast-style code, where a bot loops over a list of chat IDs and fires messages as fast as possible, will hit the limit regardless of framework if you don’t throttle it yourself. This is worth building explicitly rather than trusting a library default, because getting flood-controlled repeatedly is also a factor Telegram weighs when deciding whether an account or bot looks abusive.

Proxy configuration, and why it’s not one setting

Bot API frameworks talk to api.telegram.org over plain HTTPS. If that endpoint is reachable, you don’t need a proxy at all. If it isn’t, from a specific hosting region for instance, you route the framework’s HTTP client through a standard SOCKS5 or HTTP proxy, which every framework mentioned above supports as a client-level setting.

MTProto userbot libraries like Telethon and Pyrogram are a different case, because they’re not making ordinary HTTPS calls. They connect directly to Telegram’s data centers over the MTProto transport, and they can be configured to route through a SOCKS5 proxy or a Telegram-native MTProto proxy. The proxy choice there interacts with account safety in ways a Bot API proxy doesn’t, since a user account’s connection pattern is part of what Telegram’s abuse systems look at. That’s a decision worth making deliberately, not copying from a script you found somewhere.

What actually holds up

There’s no single “best” framework here. The ones that survive production are the ones where you’ve deliberately closed the three gaps: persistent state instead of an in-memory dict, a deployment process that guarantees only one polling instance is ever running, and rate-limit handling that doesn’t rely on the library catching every case for you. Pick a framework that fits your language and concurrency needs, then spend the time on those three things instead of assuming the framework handles them by default.

If you’re running bots or userbots on infrastructure you’d rather not babysit, that’s exactly what we build for at telegramvault.org.

Get new guides and videos first — join the Telegram channel.

need infra for this today?