Serialization¶
A queued call is stored as {'function': ..., **kwargs}. The default format is
JSON.
Why not pickle¶
Whatever can write to the queue decides what the bot container executes.
With pickle that means arbitrary code. JSON is narrower, but not merely "a
different Telegram call": a payload naming an FSInputFile picks a path, so a
queue writer can make the bot upload any file its container can read. Treat the
queue as a trust boundary either way — see
SECURITY.md.
1.0.4 moved to pickle because keyboards would not survive as plain dicts. That is no longer true: aiogram 3 models are pydantic v2 and round-trip cleanly.
What survives the queue¶
| Type | Notes |
|---|---|
| Keyboards | all four reply_markup types, nested buttons intact |
| aiogram models | InputMedia*, MessageEntity, LinkPreviewOptions, ReplyParameters, … |
datetime, date |
ISO format |
Decimal |
exact, as a string |
bytes |
base64 |
| Enums | by value |
URLInputFile, FSInputFile, BufferedInputFile |
see below |
| Plain data | strings, numbers, booleans, lists, dicts, None |
Files¶
URLInputFile and BufferedInputFile carry everything they need. FSInputFile
carries only a path, so the file must also exist in the bot container — share a
volume, or send the bytes instead.
Anything else that is not JSON-representable raises SerializationError when
queued, naming the alternative:
FooInputFile cannot be queued. Send a file_id or a URL instead,
or set TELEGRAM_BOT['SERIALIZER'] to 'pickle' together with
ALLOW_PICKLE = True, or the reader will refuse what it writes.
Falling back to pickle takes both keys, and a queue nothing untrusted can write to — the reader refuses pickled payloads unless told otherwise:
The method name is checked against an allowlist¶
A payload names the method to call, so that name is validated before anything is
looked up on the bot. Only the Telegram API methods aiogram exposes are
accepted: 185 names match aiogram.methods at the time of writing, and 181
are allowed once four are denied. Anything else is refused with a ValueError.
Those four of aiogram's own are denied on purpose: set_webhook and delete_webhook
reconfigure where Telegram delivers, log_out invalidates the token, and close
tears down the session the consumer is using. None is a message, and a queue is
not the place to reach them from — manage.py tgbot_webhook is.
That closes off the other public attributes a Bot carries: download_file
would write to the container's filesystem, token would hand out the
credential. Neither is reachable from the queue.
This narrows what a queue writer can do; it does not make the queue safe. Redis remains a trust boundary — see SECURITY.md.
Two details that make it work¶
Every model is tagged with its class name. Decoding looks the class up
rather than inferring it from a union. Without this, InputMediaPhoto comes
back as InputMediaAudio whenever the discriminator is missing.
Default sentinels are tagged too. aiogram fills unset fields with a
Default marker that pydantic cannot serialize. The obvious fix —
exclude_unset=True — also strips discriminators, which is exactly how the
InputMediaPhoto corruption happens. So the sentinels are preserved by name
and rebuilt on the way out.
Class lookup is limited to aiogram.types members that subclass
TelegramObject, so a payload cannot name an arbitrary import path.
Why not a faster JSON library¶
Asked often enough to be worth answering with a number rather than a preference.
The bench, runnable as it stands from a Django shell:
import json
import statistics
import timeit
import uuid
from django_aiogram.wire.envelope import pack
from django_aiogram.wire.serializers import get_serializer
payload = pack(
'send_message',
{'chat_id': 12345, 'text': 'hello there, a realistic message'},
uuid.UUID('11111111-1111-1111-1111-111111111111'),
1700000000.0,
)
serializer = get_serializer() # bound once, the way the queueing path binds it
calls = 200_000
rounds = 5 # the median of five, so one slow round cannot set the number
def per_call(what):
"""Microseconds per call, median of `rounds` runs of `calls` each."""
return statistics.median(timeit.timeit(what, number=calls) / calls * 1e6 for _ in range(rounds))
print(per_call(lambda: serializer.dumps(payload)))
print(per_call(lambda: json.dumps(payload)))
print(per_call(lambda: get_serializer().dumps(payload)))
print(per_call(lambda: json.dumps(payload, separators=(',', ':'))))
A fixed correlation id and timestamp, so the payload is byte-stable between runs: a 32-character body, 189 bytes encoded. Per-call means over 200 000 calls, CPython 3.13.14 on arm64 macOS.
serializer.dumps(payload), serializer bound |
0.90 µs |
json.dumps(payload) — same bytes |
0.80 µs |
get_serializer().dumps(payload) — lookup included |
0.98 µs |
json.dumps(payload, separators=(',', ':')) — 177 bytes, different output |
0.96 µs |
Every row from one run, median of five, so the differences below are subtractions of these numbers rather than separate measurements — which is how they came to disagree.
So the tagging costs about 0.10 µs over a bare json.dumps producing the same
bytes: the price of default being available to encode aiogram models. The third row
is a separate 0.08 µs for resolving the serializer, which the queueing path pays once
per write rather than once per message — worth separating, because it is the same size
as the overhead and easy to attribute to the wrong thing.
A faster library has to beat 0.10 µs plus the 0.80 µs underneath it — roughly a
microsecond in total, against a Redis round trip measured at 14 µs on Linux (105 µs on
macOS, so treat it as an order of magnitude) and a Telegram call
in tens of milliseconds. orjson would also change what is representable, since it
has its own rules about dict keys and subclasses while the tagging here depends on
default being called for exactly the types it registers.
The last row is worth knowing for a different reason: this package encodes with Python's default separators, so a payload carries about 6% more bytes than it needs to. Cheap to change and deliberately not changed here — it would rewrite every queued payload's bytes, which is a decision for a release thinking about storage rather than one closing out its checks.
Pickle, the escape hatch¶
JSON is the format. Pickle is what is left when a payload has no JSON form at
all — UnsupportedInputFileError names it for exactly that reason when you try
to queue an open file. It is not a migration aid: 3.0 removed the shim and the
1.x queue is long drained, and the setting is still here because the escape
hatch is still needed.
It is off by default because unpickling queue data is code execution. Whoever can write to the queue can run code in the bot container — a Redis list or stream, an AMQP queue, a Kafka topic — so this is a trust boundary, not a preference.
Reading pickled payloads takes one key:
Writing them takes both — writing a format the reader refuses would discard
every message, which is what E022 reports before deployment:
Two behaviors make the mixed case work:
- Reads sniff the format per message. A queue holding both formats drains
without being stopped, so switching
SERIALIZERneeds no downtime. - A refused pickle stays in flight, rather than being acknowledged — on
Redis 6.2 and newer. It is one of the cases where
dispatch()returnsFalse, which is what withholds the acknowledgement; see Delivery. TurningALLOW_PICKLEoff while a producer is still writing pickled payloads leaves them in the worker's processing list with a log line saying so; set it back and restart the worker under the same worker identity —WORKER_NAME, or the fixedhostname:it falls back to — and they are delivered. Start it under a different one and the backlog sits in a list the new worker never opens; see Delivery.
Turning ALLOW_PICKLE off is only recoverable where LMOVE exists
Without it there is no in-flight list — the consumer has already popped the message when
it refuses it, so a refused pickle is gone. The consumer says which mode it is in at
startup, as tg_crash_safe on the delivery started line. On such a server, stop or
upgrade every pickle producer before turning the flag off. See
Delivery.
Only worth it if you must queue objects JSON cannot represent, and only with a queue nothing untrusted can write to.
Failures¶
Both serializers raise SerializationError — never a bare TypeError,
ValueError or RecursionError — so callers have one exception to catch. On
the consumer side an undecodable message is logged and dropped rather than
stopping the worker.