Skip to content

Benchmarks

"I only trust a benchmark I've falsified myself." — not Churchill

So don't take ours. Numbers you can trust come from a method you can see — and every figure below ships with the harness that made it, so you can run it on the hardware, libc, and allocator you actually deploy on. Everything runs over the same fixtures the parity tests prove byte-identical, so each benchmark times exactly the behaviour a test certifies.

All numbers here are measured on Linux via the repo's Docker image (benchmarks/docker) — not macOS, which distorts filesystem- and allocation-bound work (the why is below).

End to end, per request (incl. SQL)

In-memory SQLite, real queries, controller-shaped workloads. Vanilla Eloquent vs. the same models with HasGrease. Output is byte-identical.

Endpoint — one request, incl. SQL (p50)vanilla+ GreaseΔ
list 100 users → JSON3.21 ms0.47 ms−86%
100 posts with author → JSON6.63 ms1.28 ms−81%
100 posts with tags (m2m) → JSON17.92 ms5.04 ms−72%
show one post (with author)0.11 ms0.05 ms−51%
load 150, mutate, save7.62 ms6.59 ms−14%

Live from benchmarks/realworld.php · linux-docker · sqlite :memory: · JIT · PHP 8.4.22 · 0915130 · parity pass · generated 2026-06-25

These numbers are live — rendered straight from the JSON the parity-gated harness emits, so the table can never drift from what the benchmark actually measures (and never publishes at all if parity fails). Run it yourself, or regenerate the figures above:

bash
php benchmarks/realworld.php          # human-readable table
bash benchmarks/export-metrics.sh     # regenerate the live JSON the docs read

The gain scales with how much your request hydrates: wide selects, eager loads, and serialization-heavy API responses benefit most. A request that does almost no Eloquent work has almost nothing for Grease to speed up — and that's the honest shape of it. These are :memory: numbers — read the methodology note below before mapping them onto a networked database.

Per operation (CastBench, A/B)

In-memory, paired *Vanilla / *Greased subjects, so you read the per-operation delta directly — live from CastBench + DateSerializationBench:

Operationvanilla+ GreaseΔ
hydrate a row7.47 µs3.20 µs−57%
read all casts50.8 µs36.9 µs−27%
read an enum cast2.62 µs1.27 µs−51%
set + dirty-check32.3 µs13.7 µs−58%
toArray() (serialize)111.5 µs51.8 µs−54%
date serialization (timestamps)21.9 µs3.29 µs−85%
date serialization (datetime casts)31.7 µs3.83 µs−88%
bash
composer bench

The standout is date serialization: skipping the Carbon parse-and-reformat round-trip saves roughly 27 µs per date column per row. On an API response with a few timestamps across a hundred rows, that single tier is most of the win.

Building your array by hand — Scout's toSearchableArray, a JsonResource, an export — bypasses toArray() and so this tier. The serialization helpers hand it back: greaseSerializeDate() (−86% on a hand-picked date) and greaseSerializeOnly() (−91% on a curated subset of a wide model). Validate both with php benchmarks/serialize_helpers.php.

Beyond Eloquent: the dispatcher

Measured separately (DispatcherBench, EventStormBench), because it's app-wide, not per-model:

Dispatchvanilla+ GreaseΔ
dispatch, no listener0.39 µs0.19 µs−51%
dispatch, with listeners0.73 µs0.57 µs−21%
event-dense request, warm15.6 µs7.09 µs−55%
event-dense request, cold (wildcards)77.8 µs36.0 µs−54%

On an event-dense request it roughly halves the event overhead. There's a further win on a Blade- or Livewire-heavy page: the framework fires creating:/composing: through a hasListeners() guard (callCreator/callComposer), not a bare dispatch(), and the greased dispatcher memoizes that presence check — so re-rendering the same components stops re-scanning wildcards every time. How much that's worth depends on how many observer/wildcard listeners you've registered. See The Event Dispatcher.

Beyond Eloquent: Blade components

A third axis again — the render path, not the model. The provider swaps two singletons (blade.compiler and view) for greased, byte-identical drop-ins. The macro (benchmarks/blade.php) now runs nine parity-gated variants, each asserting the HTML is identical before it times anything:

Render — 1,000 components (p50)vanilla+ GreaseΔ
simple avatar (initials + one merge)16.60 ms11.72 ms−29%
simple avatar, @foreach (realistic loop)18.59 ms12.92 ms−31%
rich avatar (5 props, @php, conditionals, slots)24.15 ms18.92 ms−22%
rich avatar, @foreach (realistic loop)24.36 ms18.49 ms−24%
app page (class components, slots, @include/@each, composer)46.63 ms38.23 ms−18%
data table (nested @foreach, heavy $loop use)0.44 ms0.32 ms−28%
layout inheritance (@extends/@section/@yield/@push)0.46 ms0.37 ms−20%
asset stacks (@push/@prepend per row, @stack)1.31 ms1.17 ms−11%
full page (extends layout, 5 sections, 100-row @foreach table, components)0.17 ms0.15 ms−10%

The full page is the realistic composite — every tier firing at once. It lands lower than any single-axis variant (−11.5%) because on a normal page genuine work dominates (~53% compiled template bodies, ~24% e() escaping — both off-limits); the single-axis rows show what each tier is worth where it does dominate.

bash
php benchmarks/blade.php

That's seven byte-identical wins compounded: @props resolution, the $attributes->merge() pipeline, a greased bag for class components, getCompiledPath memoization, @foreach's $loop bookkeeping, @yield's content stitching, and the @push/@prepend stack assembly — not a halving, but a real, byte-identical cut. The split is by page shape: component greasing wins on component-dense pages, loop greasing on cheap-bodied loops (tables, lists) — and the two compose with zero regression. The honest scope, the dead ends we measured and rejected, and how to profile it are in Blade Components.

The whole stack, compounding

The tiers are independent opt-ins, and they stack. This is a real request through the HTTP kernel — boot, route, query the DB (the four shapes above, plus a filtered/paginated API listing), serialize to JSON and render to Blade — measured with each tier layered in, in order of least-invasive opt-in: +models+events+blade+container+request. Every cell is byte-identical to vanilla (a parity test asserts all five shapes × six levels before any timing), and the last row is the cumulative retained-memory cost of all those caches.

Route — one request through the kernelvanilla+ models+ events+ blade+ container+ request
index_users.json3.54 ms−84%−84%−84%−84%−84%
index_users.blade2.85 ms−18%−18%−38%−39%−37%
posts_with_author.json6.35 ms−81%−81%−82%−81%−81%
posts_with_author.blade3.00 ms−18%−20%−37%−36%−37%
show_post.json205 µs−28%−28%−29%−29%−29%
show_post.blade205 µs−8%−6%−11%−13%−10%
bulk_update.json12.82 ms−48%−49%−49%−49%−47%
bulk_update.blade11.54 ms−18%−21%−27%−27%−26%
filtered_users.json808 µs−67%−67%−67%−67%−67%
filtered_users.blade733 µs−10%−10%−27%−26%−26%
retained memory Δ10.7 MB+1.4% +2.0% +2.7% +2.9% +2.9%

Read it by column: +models does the heavy lifting on data-bound routes (JSON especially); the Blade tier shows up where a page actually renders (*.blade rows roughly double their delta when it lands); +container and +request add the final, smaller compounding slices — resolution and input-reading are thin slices of a full request, so each moves it a few percent, honestly. The headline: the full mixed page-load suite runs ~−47% end to end, and the entire six-tier cache footprint costs ~+2% retained memory — the caches are nearly free against a request's working set. The isolated per-tier numbers (container −38.8% per resolve, request −41% per input-heavy request) are in The Container and The Request.

bash
php benchmarks/stack_pipeline.php     # the table above, human-readable
composer bench -- benchmarks/Bench/StackPipelineBench.php   # phpbench, per-level

How to read these honestly

This package was built measure-first, and the docs hold the same line. A few things worth knowing so you can map these numbers onto your deploy:

In-memory SQLite inflates the percentage — read the absolute time

The macro runs on :memory: SQLite, where database I/O is near-zero. That makes the ORM layer (and therefore Grease's slice) a larger fraction of total request time than it would be against a networked Postgres/MySQL.

The portable figure is the absolute time Grease removes from the ORM layer — that stays roughly the same regardless of your database. The percentage shrinks as network and I/O take a bigger share of the request. So treat the per-request percentages as "Grease's share of the Eloquent-bound work," not "your p99 will drop 17%." If your endpoint is I/O-bound, you'll see the same milliseconds saved against a larger denominator.

  • Per-op vs per-request. A −54% on hydrate is a per-operation figure; it becomes a per-request figure only multiplied by how many rows you hydrate. The end-to-end table is the one that includes your SQL.
  • The win is workload-shaped. Grease accelerates hydration, casting, and serialization. It does nothing for query building, validation, or your business logic. Profile first — if Eloquent isn't your hot path, this isn't your package.
  • Marginal numbers stay marginal numbers. Where a tier benchmarked inside the noise, it was parked, not shipped, and that's recorded openly in the repo's notes. The figures here are the ones that cleared the bar. How a candidate gets from "looks faster" to "shipped" (or "measured dead end") is written up in The Method.

A benchmark is a property of the build

Why Linux, not the Mac these were first written on: macOS's /var/private/var symlink confuses opcache's realpath keying and CLI opcache behaves unlike production. It inflated the per-op microbench wins and understated Blade — the is_file() it ranked at ~8% of a render is ~3% on Linux. (Xdebug lied too: it ranked extract at ~14% of a render when a micro proved it ~0.6% — it disables JIT and mis-attributes internal-op cost. Use a sampling profiler; the repo ships an Excimer harness.)

But even "Linux" isn't one number. The same CastBench, same machine, only the libc and allocator changed:

Opglibcmusl
read all casts−26.5%−33%
toArray()−46%−52%

Grease's wins are allocation wins, and musl's allocator makes the vanilla arm pay more — so the same optimization reads bigger on musl. (jemalloc via LD_PRELOAD didn't even run — it crashed PHP's JIT; "drop-in allocator" is a myth. And run-to-run under load swung glibc setDirty −39%↔−27%, wider than the libc gap.) Treat the figures as representative, not a promise. Reproduce on your build — it's one command.

Reproduce everything

bash
# Linux, the canonical environment (glibc; swap Dockerfile.alpine for musl):
docker build -t grease-bench benchmarks/docker
docker run --rm -v "$PWD":/app -w /app grease-bench php benchmarks/realworld.php

# or directly, on whatever PHP you have:
composer test                       # parity tests — the byte-identical contract
composer bench                      # phpbench: CastBench (per-op A/B) + SuiteBench (SQL)
php benchmarks/realworld.php         # the end-to-end macro above
php benchmarks/blade.php             # the 1,000-component render benchmark
php benchmarks/serialize_helpers.php # greaseSerializeDate / greaseSerializeOnly, A/B
php benchmarks/blade_excimer.php     # honest sampling profile of the render path

Same fixtures, both sides. A bench runs exactly what a test proves identical.

Byte-identical to vanilla, or it's a failing test. · MIT Licensed