Skip to content

gh-158585: Use PyBytesWriter in io.BufferedReader.readline() - #158615

Open
vstinner wants to merge 1 commit into
python:mainfrom
vstinner:buffered
Open

vstinner wants to merge 1 commit into
python:mainfrom
vstinner:buffered

Conversation

@vstinner

@vstinner vstinner commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Replace a list of bytes object with PyBytesWriter.

Replace a list of bytes object with PyBytesWriter.
@vstinner

vstinner commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

cc @cmaloney

I ran a benchmark testing different line lengths (around 80 bytes, BUFFER_SIZE / 4, around BUFFER_SIZE, BUFFER_SIZE * 10) and different number of lines (100, below 200, 3000+).

In short, there is no impact on performance, or it's a little bit faster (1.01x faster).

Common text files have short lines: they fit into PyBytesWriter small buffer (255 bytes). So a single bytes object is directly created with the exact final size in PyBytesWriter_Finish().

If the line can be read from the "readahead" io.BufferedRead buffer, PyBytes_FromStringAndSize(buffer, size) is used: the PR doesn't change this code path. The benchmark uses a buffer size (4 kiB) smaller than the default buffer size (128 kB) to test less often this unchanged code path.

There are mostly differences on the benchmark when readline() needs to combine multiple read() outputs: "medium lines" and "long lines".

Results:

Benchmark ref writer
medium lines 271 us 269 us: 1.01x faster
long lines 6.48 ms 6.43 ms: 1.01x faster
Geometric mean (ref) 1.00x slower

Benchmark hidden because not significant (6): Lib/struct.py, Lib/io.py, Lib/colorsys.py, Modules/_io/bufferedio.c, Lib/typing.py, short lines.

Benchmark run on Fedora 44 with CPU isolation. Python built with gcc -O3.

Benchmark code:

Details
import pyperf
import io

NLINE = 100
# Python 3.16 uses 128 kiB by default
BUFFER_SIZE = 4 * 1024

runner = pyperf.Runner()

def readlines(file):
    file.seek(0)
    for line in file:
        pass

    file.seek(0)
    while True:
        line = file.readline()
        if not line:
            break

def bench(name, data):
    raw = io.BytesIO(data)
    buffered = io.BufferedReader(raw, buffer_size=BUFFER_SIZE)
    runner.bench_func(name, readlines, buffered)


filenames = (
    # under 200 lines of ~80 bytes
    'Lib/struct.py',
    'Lib/io.py',
    'Lib/colorsys.py',

    # 3000+ lines of ~80 bytes
    'Modules/_io/bufferedio.c',
    'Lib/typing.py',
)

for filename in filenames:
    with open(filename, "rb") as fp:
        data = fp.read()

    bench(filename, data)

# Test short lines (BUFFER_SIZE / 4): a single readline()
# call only needs to call read() once
short_line = b'x' * (BUFFER_SIZE // 4 - 4) + b'abc\n'
data = [short_line for _ in range(NLINE)]
data = b''.join(data)
bench('short lines', data)

# Test medium lines around BUFFER_SIZE: a single readline()
# call needs to call read() once or twice
medium_line_a = b'x' * (BUFFER_SIZE - 16) + b'abc\n'
medium_line_b = b'x' * (BUFFER_SIZE + 16) + b'abc\n'
combined = medium_line_a + medium_line_b
data = [combined for _ in range(NLINE // 2)]
data = b''.join(data)
bench('medium lines', data)

# Test long lines longer than BUFFER_SIZE: a single readline()
# needs multiple read()
long_line = b'x' * (BUFFER_SIZE * 10) + b'trailer\n'
data = [long_line for _ in range(NLINE)]
data = b''.join(data)
bench('long lines', data)

@vstinner vstinner changed the title gh-158585: Use PyBytesWriter in _io._Buffered.readline() gh-158585: Use PyBytesWriter in io.io.BufferedReader.readline() Oct 2, 2026
@vstinner vstinner changed the title gh-158585: Use PyBytesWriter in io.io.BufferedReader.readline() gh-158585: Use PyBytesWriter in io.BufferedReader.readline() Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant