These two loops build the same 1.25 MB string with the same statement. One takes 0.269 milliseconds and the other takes 218.836.
<?php
declare(strict_types=1);
$piece = str_repeat('a', 50);
// fast
$s = '';
for ($i = 0; $i < 25_000; $i++) {
$s .= $piece;
}
// 814 times slower
$s = '';
$keep = '';
for ($i = 0; $i < 25_000; $i++) {
$keep = $s;
$s .= $piece;
}
The appending statement is identical — the same
.= the manual defines in one line.
What changed is that something else is holding the buffer while it grows, and
that is the whole subject.
How these numbers were taken
PHP 8.5.10, NTS, arm64, Homebrew build, on an Apple M4 Pro laptop with 24 GB and
12 logical cores, running macOS and not quiesced. OPcache and the JIT are off
except where a measurement is specifically about OPcache, which is labeled where
it appears. memory_limit is 4G.
Timings are the median of 15 interleaved runs with three warmup rounds
discarded, taken with hrtime() inside the process, with the range given. Two
exceptions are labeled: the four-way comparison at 25,000 parts is the median of
five interleaved runs, and the scaling table for $s = $s . $piece is the median
of three, because a single run of its largest size takes eleven seconds. Memory
figures are exact and were identical across every round. Read the timings
relative to each other rather than as absolutes, for
the reasons the method post sets out.
Opcode dumps come from opcache.opt_debug_level=0x20000. One section measures
the same script under three configurations — OPcache off, OPcache’s shared memory
on in the CLI, and OPcache on under the built-in server — because the answer
differs between them.
A string is a header, the bytes, and rounding
A zend_string is a 24-byte header — the refcount and flags, a cached hash, and
the length — followed by the bytes and a terminating null. The allocator then
rounds the total up to one of its size bins, which is what a userland measurement
actually sees:
len 0 0 bytes
len 1 32 bytes
len 7 32 bytes
len 8 40 bytes
len 15 40 bytes
len 16 48 bytes
len 24 56 bytes
len 32 64 bytes
len 100 128 bytes
Twenty-four plus one plus the length, rounded to a multiple of eight, with a 32-byte floor. Length 7 fits 24 + 7 + 1 = 32 exactly; length 8 needs 33 and gets 40. The empty string costs nothing at all, because it is a shared singleton the engine hands out rather than an allocation.
That is the cost of the container. What the variable holding it costs is sixteen bytes and does not change, whatever the string’s length.
A literal is not an allocation
Two strings with identical contents, one written in the source and one produced by a function call:
<?php
declare(strict_types=1);
$baseline = memory_get_usage();
$literal = 'elephantphp';
printf("literal %+d bytes\n", memory_get_usage() - $baseline);
$baseline = memory_get_usage();
$built = strrev('phptnahpele');
printf("built at run %+d bytes\n", memory_get_usage() - $baseline);
debug_zval_dump($literal);
debug_zval_dump($built);
literal +0 bytes
built at run +40 bytes
string(11) "elephantphp" interned
string(11) "elephantphp" refcount(3)
The literal is interned: stored once, shared by every place in the compiled code that mentions it, and not reference counted, because nothing will ever release it. The runtime string is an ordinary allocation with a count on it — 40 bytes, which is the arithmetic from the previous section for eleven characters.
=== between them is still true. Interning is about storage, not identity.
Where interned strings live, and why the CLI lies about it
With OPcache on, interned strings from compiled code go into a pool of their own,
sized by opcache.interned_strings_buffer and separate from the memory the
opcode cache itself uses. It is not empty when your code starts:
interned buffer: used 2,599,584 of 8,388,608 bytes, 10,507 strings
Eight megabytes by default, and a quarter of it already spent on the engine’s own
strings and the script just compiled. It is a distinct knob from
opcache.memory_consumption, and it can fill on its own — a codebase with a very
large number of distinct class, method and literal strings exhausts the interned
buffer while the opcode cache still has room, and once it is full, later strings
are simply not interned. That is
a fourth pool alongside the three this archive has already separated.
Now the part that will cost you an afternoon if nobody warns you. The same literal, measured four ways:
OPcache off, CLI interned
OPcache shared memory on, CLI refcount(3)
opcache.file_cache_only=1, CLI interned
OPcache on, built-in web server interned
Three of those four agree and the second does not. In the CLI SAPI with shared memory enabled, the literal in the script being executed was a refcounted string rather than an interned one, reproducibly, on every run and whether or not a file cache had been written on a previous invocation.
I could not establish why. What I can say is what it means for measurement: the CLI with OPcache enabled is the one configuration where interning does not behave as the other three do, and it is also the configuration most people would reach for to test this. Measure it under a SAPI that keeps its cache warm.
Building a big string
Ten megabytes, as 200,000 pieces of 50 bytes, by the two routes people argue about:
| Building 10 MB | Time | Peak memory |
|---|---|---|
$s .= $piece |
2.569 ms (2.431–2.637) | 10,496,312 |
implode('', $parts) |
2.673 ms (2.572–2.862) | 17,918,368 |
The advice to replace concatenation with implode() is a tie on time and a loss
on memory. implode() cannot start until every part exists, so the peak holds
the parts array and the finished string at once — 1.71 times what the append
route needed. The append route holds one buffer and makes it longer.
Neither of those is the case that ruins an export job. This is:
| 25,000 parts, 1.25 MB | Time | Peak memory |
|---|---|---|
$s .= $piece |
0.269 ms (0.255–0.277) | 1,739,040 |
implode('', $parts) |
0.326 ms (0.316–0.354) | 2,668,888 |
$s = $s . $piece |
228.456 ms (221.886–235.782) | 2,992,424 |
$s .= $piece, with a second holder |
218.836 ms (215.966–230.497) | 2,992,424 |
Eight hundred and forty-nine times, for a statement most people would read as a spelling variant. And the fourth row is the same operator as the first, made slow by a line that does not touch it.
One opcode against two
The dump separates the first two rows before any measurement does:
<?php
function appendInPlace(string $piece): string { $s = ''; for ($i=0;$i<3;$i++) { $s .= $piece; } return $s; }
function reassign(string $piece): string { $s = ''; for ($i=0;$i<3;$i++) { $s = $s . $piece; } return $s; }
appendInPlace:
0004 ASSIGN_OP (CONCAT) CV1($s) CV0($piece)
reassign:
0004 T3 = FAST_CONCAT CV1($s) CV0($piece)
0005 ASSIGN CV1($s) T3
ASSIGN_OP names its destination as the thing being operated on. FAST_CONCAT
builds a new string in a temporary and ASSIGN then replaces the variable with
it, which means allocating the full new length and copying every byte already
written, on every iteration.
concat_function() in zend_operators.c is where the
first form gets its advantage, and the comment says what it is for:
if (result == op1) {
/* Destroy the old result first to drop the refcount, such that $x .= ...; may happen in-place. */
...
result_str = zend_string_extend(op1_string, result_len, 0);
When the destination and the left operand are the same zval, the engine drops its
own reference and extends the existing allocation. zend_string_extend() can
hand back the same pointer, grown, when nothing else is holding it.
Which is why the second holder costs so much
That condition — nothing else is holding it — is what the fourth row of the
table breaks. $keep = $s; raises the count on the buffer, so the append can no
longer extend in place and has to allocate a full copy instead. One statement,
elsewhere in the loop, converts the cheap path into the expensive one, and the
cost profile becomes the same as $s = $s . $piece: 218.836 ms against 228.456,
and identical peak memory.
This is exactly the rule the first post in this series measured on arrays — the write pays, and it pays when more than one name is holding the value. Strings are the same mechanism, with the difference that an array separates once and a growing string separates on every append, so the same rule that costs an array one copy costs a string a quadratic number of them.
What that grows into:
Parts appended with $s = $s . $piece |
Time |
|---|---|
| 12,500 | 57.352 ms |
| 25,000 | 235.630 ms |
| 50,000 | 1,582.328 ms |
| 100,000 | 11,249.585 ms |
Median of three runs each. Doubling the work multiplied the time by 4.11, then 6.72, then 7.11 — worse than the four a purely quadratic copy would predict, because the copies themselves get less cache-friendly as the buffer grows past what fits.
Finding the second holder
The count is visible from userland, with one condition: it has to be read inside the loop. After the loop both versions read the same, because the extra holder is left pointing at a buffer the last append already replaced.
<?php
declare(strict_types=1);
function countOf(string &$var): string
{
ob_start();
debug_zval_dump($var);
$dump = ob_get_clean();
return trim(substr($dump, strrpos($dump, '"') + 1)) ?: 'interned';
}
$piece = str_repeat('a', 10);
$s = '';
for ($i = 0; $i < 3; $i++) {
printf("sole owner iteration %d: %s\n", $i, countOf($s));
$s .= $piece;
}
sole owner iteration 2: refcount(2)
second holder iteration 2: refcount(3)
One apart, every iteration. The absolute values are inflated — the helper holds a
reference of its own and so does the output buffer — so read the difference
between the two loops rather than either number, which is the same caution
debug_zval_dump() needed
the first time this series used it. The
sole-owner case is the smaller one, and it is the one that extends in place.
What to do with a loop that builds a string
Append with .= and leave the buffer alone. That is the fast path, it is the
lowest-memory path, and it is the default that folklore has been steering people
away from.
Then check who else is holding it. The failure is not the operator; it is a second name on the same buffer while the loop runs. Look for an assignment that snapshots the buffer, a value passed to a function that stores it, a property that keeps a copy, and a reference bound to it — the ampersand this series measured at 32 bytes is a second holder too. Any of them turns one statement into a copy per iteration, and none of them is on the line the profiler blames.
Reach for implode() when the parts already exist for another reason — you built
an array of rows because something else needed the array. Building the array
solely to implode it costs a second copy of the data at peak and buys nothing
measurable in time.
Write literals rather than assembling constant strings at run time. A literal is interned and free; the same content produced by a call is an allocation with a count on it, and in a loop it is an allocation per iteration.
Комментарии (0)
Пока нет комментариев — будьте первым.