These two loops build the same 1.25 MB string with the same statement. One takes 0.269 milliseconds and the other takes 218.836.

<?php

declare(strict_types=1);

$piece = str_repeat('a', 50);

// fast
$s = '';
for ($i = 0; $i < 25_000; $i++) {
    $s .= $piece;
}

// 814 times slower
$s = '';
$keep = '';
for ($i = 0; $i < 25_000; $i++) {
    $keep = $s;
    $s .= $piece;
}

The appending statement is identical — the same .= the manual defines in one line. What changed is that something else is holding the buffer while it grows, and that is the whole subject.

How these numbers were taken

PHP 8.5.10, NTS, arm64, Homebrew build, on an Apple M4 Pro laptop with 24 GB and 12 logical cores, running macOS and not quiesced. OPcache and the JIT are off except where a measurement is specifically about OPcache, which is labeled where it appears. memory_limit is 4G.

Timings are the median of 15 interleaved runs with three warmup rounds discarded, taken with hrtime() inside the process, with the range given. Two exceptions are labeled: the four-way comparison at 25,000 parts is the median of five interleaved runs, and the scaling table for $s = $s . $piece is the median of three, because a single run of its largest size takes eleven seconds. Memory figures are exact and were identical across every round. Read the timings relative to each other rather than as absolutes, for the reasons the method post sets out.

Opcode dumps come from opcache.opt_debug_level=0x20000. One section measures the same script under three configurations — OPcache off, OPcache’s shared memory on in the CLI, and OPcache on under the built-in server — because the answer differs between them.

A string is a header, the bytes, and rounding

A zend_string is a 24-byte header — the refcount and flags, a cached hash, and the length — followed by the bytes and a terminating null. The allocator then rounds the total up to one of its size bins, which is what a userland measurement actually sees:

len 0       0 bytes
len 1      32 bytes
len 7      32 bytes
len 8      40 bytes
len 15     40 bytes
len 16     48 bytes
len 24     56 bytes
len 32     64 bytes
len 100   128 bytes

Twenty-four plus one plus the length, rounded to a multiple of eight, with a 32-byte floor. Length 7 fits 24 + 7 + 1 = 32 exactly; length 8 needs 33 and gets 40. The empty string costs nothing at all, because it is a shared singleton the engine hands out rather than an allocation.

That is the cost of the container. What the variable holding it costs is sixteen bytes and does not change, whatever the string’s length.

A literal is not an allocation

Two strings with identical contents, one written in the source and one produced by a function call:

<?php

declare(strict_types=1);

$baseline = memory_get_usage();
$literal = 'elephantphp';
printf("literal        %+d bytes\n", memory_get_usage() - $baseline);

$baseline = memory_get_usage();
$built = strrev('phptnahpele');
printf("built at run   %+d bytes\n", memory_get_usage() - $baseline);

debug_zval_dump($literal);
debug_zval_dump($built);
literal        +0 bytes
built at run   +40 bytes
string(11) "elephantphp" interned
string(11) "elephantphp" refcount(3)

The literal is interned: stored once, shared by every place in the compiled code that mentions it, and not reference counted, because nothing will ever release it. The runtime string is an ordinary allocation with a count on it — 40 bytes, which is the arithmetic from the previous section for eleven characters.

=== between them is still true. Interning is about storage, not identity.

Where interned strings live, and why the CLI lies about it

With OPcache on, interned strings from compiled code go into a pool of their own, sized by opcache.interned_strings_buffer and separate from the memory the opcode cache itself uses. It is not empty when your code starts:

interned buffer: used 2,599,584 of 8,388,608 bytes, 10,507 strings

Eight megabytes by default, and a quarter of it already spent on the engine’s own strings and the script just compiled. It is a distinct knob from opcache.memory_consumption, and it can fill on its own — a codebase with a very large number of distinct class, method and literal strings exhausts the interned buffer while the opcode cache still has room, and once it is full, later strings are simply not interned. That is a fourth pool alongside the three this archive has already separated.

Now the part that will cost you an afternoon if nobody warns you. The same literal, measured four ways:

OPcache off, CLI                      interned
OPcache shared memory on, CLI         refcount(3)
opcache.file_cache_only=1, CLI        interned
OPcache on, built-in web server       interned

Three of those four agree and the second does not. In the CLI SAPI with shared memory enabled, the literal in the script being executed was a refcounted string rather than an interned one, reproducibly, on every run and whether or not a file cache had been written on a previous invocation.

I could not establish why. What I can say is what it means for measurement: the CLI with OPcache enabled is the one configuration where interning does not behave as the other three do, and it is also the configuration most people would reach for to test this. Measure it under a SAPI that keeps its cache warm.

Building a big string

Ten megabytes, as 200,000 pieces of 50 bytes, by the two routes people argue about:

Building 10 MB Time Peak memory
$s .= $piece 2.569 ms (2.431–2.637) 10,496,312
implode('', $parts) 2.673 ms (2.572–2.862) 17,918,368

The advice to replace concatenation with implode() is a tie on time and a loss on memory. implode() cannot start until every part exists, so the peak holds the parts array and the finished string at once — 1.71 times what the append route needed. The append route holds one buffer and makes it longer.

Neither of those is the case that ruins an export job. This is:

25,000 parts, 1.25 MB Time Peak memory
$s .= $piece 0.269 ms (0.255–0.277) 1,739,040
implode('', $parts) 0.326 ms (0.316–0.354) 2,668,888
$s = $s . $piece 228.456 ms (221.886–235.782) 2,992,424
$s .= $piece, with a second holder 218.836 ms (215.966–230.497) 2,992,424

Eight hundred and forty-nine times, for a statement most people would read as a spelling variant. And the fourth row is the same operator as the first, made slow by a line that does not touch it.

One opcode against two

The dump separates the first two rows before any measurement does:

<?php

function appendInPlace(string $piece): string { $s = ''; for ($i=0;$i<3;$i++) { $s .= $piece; } return $s; }
function reassign(string $piece): string { $s = ''; for ($i=0;$i<3;$i++) { $s = $s . $piece; } return $s; }
appendInPlace:
0004 ASSIGN_OP (CONCAT) CV1($s) CV0($piece)

reassign:
0004 T3 = FAST_CONCAT CV1($s) CV0($piece)
0005 ASSIGN CV1($s) T3

ASSIGN_OP names its destination as the thing being operated on. FAST_CONCAT builds a new string in a temporary and ASSIGN then replaces the variable with it, which means allocating the full new length and copying every byte already written, on every iteration.

concat_function() in zend_operators.c is where the first form gets its advantage, and the comment says what it is for:

if (result == op1) {
    /* Destroy the old result first to drop the refcount, such that $x .= ...; may happen in-place. */
    ...
    result_str = zend_string_extend(op1_string, result_len, 0);

When the destination and the left operand are the same zval, the engine drops its own reference and extends the existing allocation. zend_string_extend() can hand back the same pointer, grown, when nothing else is holding it.

Which is why the second holder costs so much

That condition — nothing else is holding it — is what the fourth row of the table breaks. $keep = $s; raises the count on the buffer, so the append can no longer extend in place and has to allocate a full copy instead. One statement, elsewhere in the loop, converts the cheap path into the expensive one, and the cost profile becomes the same as $s = $s . $piece: 218.836 ms against 228.456, and identical peak memory.

This is exactly the rule the first post in this series measured on arrays — the write pays, and it pays when more than one name is holding the value. Strings are the same mechanism, with the difference that an array separates once and a growing string separates on every append, so the same rule that costs an array one copy costs a string a quadratic number of them.

What that grows into:

Parts appended with $s = $s . $piece Time
12,500 57.352 ms
25,000 235.630 ms
50,000 1,582.328 ms
100,000 11,249.585 ms

Median of three runs each. Doubling the work multiplied the time by 4.11, then 6.72, then 7.11 — worse than the four a purely quadratic copy would predict, because the copies themselves get less cache-friendly as the buffer grows past what fits.

Finding the second holder

The count is visible from userland, with one condition: it has to be read inside the loop. After the loop both versions read the same, because the extra holder is left pointing at a buffer the last append already replaced.

<?php

declare(strict_types=1);

function countOf(string &$var): string
{
    ob_start();
    debug_zval_dump($var);
    $dump = ob_get_clean();

    return trim(substr($dump, strrpos($dump, '"') + 1)) ?: 'interned';
}

$piece = str_repeat('a', 10);

$s = '';
for ($i = 0; $i < 3; $i++) {
    printf("sole owner    iteration %d: %s\n", $i, countOf($s));
    $s .= $piece;
}
sole owner    iteration 2: refcount(2)
second holder iteration 2: refcount(3)

One apart, every iteration. The absolute values are inflated — the helper holds a reference of its own and so does the output buffer — so read the difference between the two loops rather than either number, which is the same caution debug_zval_dump() needed the first time this series used it. The sole-owner case is the smaller one, and it is the one that extends in place.

What to do with a loop that builds a string

Append with .= and leave the buffer alone. That is the fast path, it is the lowest-memory path, and it is the default that folklore has been steering people away from.

Then check who else is holding it. The failure is not the operator; it is a second name on the same buffer while the loop runs. Look for an assignment that snapshots the buffer, a value passed to a function that stores it, a property that keeps a copy, and a reference bound to it — the ampersand this series measured at 32 bytes is a second holder too. Any of them turns one statement into a copy per iteration, and none of them is on the line the profiler blames.

Reach for implode() when the parts already exist for another reason — you built an array of rows because something else needed the array. Building the array solely to implode it costs a second copy of the data at peak and buys nothing measurable in time.

Write literals rather than assembling constant strings at run time. A literal is interned and free; the same content produced by a call is an allocation with a count on it, and in a loop it is an allocation per iteration.