Part of the repository modernization effort.
Background
Commit 72d2aaf ("fix: Minor performance improvements in hot paths") precompiled struct.Struct packers for the request side. See bmemcached/protocol.py:61-78. That work covers the send path only.
The receive path still builds a format string on every call.
bmemcached/protocol.py:479 in get:
struct.unpack('!L%ds' % (bodylen - 4), ...)
bmemcached/protocol.py:556-558 in get_multi:
struct.unpack('!L%ds%ds' % (keylen, bodylen - keylen - 4), ...)
Each call builds a new format string. CPython caches compiled formats in an LRU cache. The default size is 100 entries. A workload with more than 100 distinct value lengths evicts entries and recompiles on every call.
This is the same pattern that 72d2aaf removed on the send side.
Warning: do not precompile a Struct per length
A Struct object for '!L%ds' % n is specific to one length n. It is not reusable across different value sizes. Precompiling does not help here.
Use a fixed-format read and a slice instead:
flags, = FLAGS_UNPACKER.unpack_from(buf) # a module-level Struct('!L')
value = buf[4:]
A slice needs no format compilation at all.
Second item: one extra copy per response
bmemcached/protocol.py:214-232 in _read_socket builds a bytearray, appends each recv() chunk, then returns bytes(value). The bytes() call copies the whole buffer once more on every response.
After the six removal lands, return the bytearray and let the caller use unpack_from and slicing on it. Neither operation needs a bytes object.
Measure this change. Do not merge it on theory alone.
Acceptance criteria
Files to change
bmemcached/protocol.py.
Order
Do the six removal issue first. That issue changes the bytes handling in the same functions.
Read the diff of 72d2aaf before you start. Do not redo work that commit already did.
Part of the repository modernization effort.
Background
Commit 72d2aaf ("fix: Minor performance improvements in hot paths") precompiled
struct.Structpackers for the request side. Seebmemcached/protocol.py:61-78. That work covers the send path only.The receive path still builds a format string on every call.
bmemcached/protocol.py:479inget:bmemcached/protocol.py:556-558inget_multi:Each call builds a new format string. CPython caches compiled formats in an LRU cache. The default size is 100 entries. A workload with more than 100 distinct value lengths evicts entries and recompiles on every call.
This is the same pattern that 72d2aaf removed on the send side.
Warning: do not precompile a
Structper lengthA
Structobject for'!L%ds' % nis specific to one lengthn. It is not reusable across different value sizes. Precompiling does not help here.Use a fixed-format read and a slice instead:
A slice needs no format compilation at all.
Second item: one extra copy per response
bmemcached/protocol.py:214-232in_read_socketbuilds abytearray, appends eachrecv()chunk, then returnsbytes(value). Thebytes()call copies the whole buffer once more on every response.After the
sixremoval lands, return thebytearrayand let the caller useunpack_fromand slicing on it. Neither operation needs abytesobject.Measure this change. Do not merge it on theory alone.
Acceptance criteria
bmemcached/protocol.py:479uses a module-levelstruct.Struct('!L')and a slice. It builds no format string.bmemcached/protocol.py:556-558uses the same approach._read_socketavoids the extrabytes()copy, or the pull request states why it stays.getandget_multi.getandget_multireturn the same types as before.Files to change
bmemcached/protocol.py.Order
Do the
sixremoval issue first. That issue changes the bytes handling in the same functions.Read the diff of 72d2aaf before you start. Do not redo work that commit already did.