Keep the Gems
Tobi Lütke ported ONCE Campfire to Ruby 4 without Rails: Falcon serving from worker Ractors, plain modules over SQLite, and message passing between Ractors instead of Redis. It serves byte-identical HTML and passes the once-campfire-rust parity suite. On the same four CPU threads it serves dynamic pages 15–27× faster than the Rails app it came from. His explainer walks through where the speed comes from, and the repository is public.
He has also been running roundhouse check on Shopify core and sending fixes for what got in his way: #171 first, and now #503, a second batch of about a hundred commits covering Sorbet and RBS signatures, typed gem boundaries from Tapioca RBIs, constant resolution, narrowing and ActiveRecord modeling.
So I asked him what he wants from Roundhouse. His answer: performance of Rails, transpiled deterministically to Ruby with the magic removed. The goal is parity with his port, or better. Getting close may be good enough. The techniques in his explainer are candidates to evaluate and measure, not requirements.
That priority makes sense to me. Other contributors have their own priorities, and that's fine. But I now know what his are, and I see no reason to keep that private.
Why Ruby to Ruby
I haven't seen Shopify's code. The largest Rails codebase I've explored is Mastodon, so that's what I use to picture the problem. Its Gemfile.lock lists 155 direct dependencies and 346 locked gems. Now picture a codebase many times larger.
In As If I compared three ways to keep a fast version of an app current: compile it, port it again for every release, or maintain the port. For something the size of Campfire, all three are plausible. For a codebase like this, two of them aren't. You are not going to vibe code a replacement, once or on every release, and you are not going to carry every upstream change into it by hand. You are not going to compile it with Spinel any time soon either: every gem would have to come along, and many are C extensions. But if you could transpile it to Ruby, it keeps its gems and still runs much faster, because the cost being removed is Rails' per-request dispatch, callbacks and rendering, not Ruby.
Mastodon's lockfile shows how much of a real app that leaves alone. Setting aside test and development tooling and rails itself, 113 runtime gems remain:
| gems | examples | |
|---|---|---|
| no dependency on Rails or Active Support | 90 | pg, redis, sidekiq, nokogiri, faraday, aws-sdk, mail, sanitize, ruby-vips |
| Active Support or Active Model only | 6 | pundit, chewy, kt-paperclip, inline_svg |
| depend on Rails | 17 | devise, doorkeeper, kaminari, simple_form, haml-rails, active_model_serializers, discard |
About 80% are plain libraries that transpiled Ruby can keep calling as it does today. The lockfile doesn't show coupling that happens at runtime (cocoon's view helpers, rack-attack's middleware, Sidekiq's Active Job adapter), so the true count of Rails-coupled gems is somewhat higher than 17. Those are where the design work is: either expand their DSLs at build time, as Roundhouse already does for Rails' own, or give the transpiled app a Rails-shaped surface they can attach to. The typed gem boundaries in #503 are the typing half of keeping the rest.
Where Roundhouse is today
At the end of As If I estimated, from published numbers taken on different machines, that a port serves pages perhaps 2–4× faster than compiled Rails, and said the way to settle it was to run a port on the same machine, under the same harness. This is that run, with Tobi's port rather than the Rust one, and with Roundhouse's CRuby output rather than Spinel's.
I ran Roundhouse's CRuby output for Campfire through the benchmark from the once-campfire-rust repository, the same one Tobi's numbers come from, next to the Rails app and his port. All three are built from the same Campfire commit, seeded with the same data, and served on the same machine: four cores of a Ryzen 7 5700U for the server, four more for the load generator, clocks capped at 2.0 GHz so every run gets the same speed. Each figure is the median of three interleaved runs, in requests per second with 16 concurrent clients and gzip on:
| Rails | Roundhouse | Tobi's port | |
|---|---|---|---|
| room page | 57 | 266 | 1,167 |
| messages page | 102 | 1,854 | 1,815 |
| sidebar | 136 | 890 | 1,866 |
| search | 98 | 634 | 1,650 |
| post a message | 77 | 523 | 776 |
| avatar | 14,472 | 2,691 | 24,607 |
Roundhouse here is Puma with one worker per core and five threads each. The absolute numbers are lower than Tobi's published ones because this machine is slower; the ratios are what carry over.
Against Rails, the transpiled app is 4.7× faster on the room page and 18× on the messages page. Against Tobi's port the picture is uneven, and the estimate from As If holds for one page and not the other. The messages page is already at parity, and posting a message is two-thirds of the way there. The room page is 4.4× behind, and the gap isn't Ruby: our messages page renders the same 40 messages at almost the same size in 2.2 ms, while the room page takes 13.9 ms. Something only the room page does costs 11 ms, and I haven't profiled it yet. The avatar route is slower than Rails, which points at something specific on that path too.
What the run found
Measuring against someone else's harness finds things your own doesn't:
- The transpiled server never raised its open-file limit. Under Docker's default of 1,024 it stopped at 991 WebSocket connections and delivered nothing. Go's runtime raises the limit at startup, and so does Tobi's server; Roundhouse's now does too.
- The emitted asset task copies only a generic scaffold's files, so Campfire's own stylesheets returned 404 until I copied them into the image by hand. That goes on the list of things that don't transpile yet.
- Broadcasts don't cross Puma workers, which is already known. For the WebSocket tests only the one-worker configuration is valid, and it delivers about a quarter of the messages per second Tobi's port does at 100 clients.
Most of the techniques in Tobi's explainer belong to what As If called the invisible row: changes nobody can observe except by timing them, which is exactly what a compiler is allowed to make. Two are already in: one write permit instead of racing SQLite's lock, and WAL checkpoints moved off the request path. As If noted the checkpoint thread as the one invisible optimization Roundhouse lacked; it doesn't anymore. On my benchmark machine the two together made posting a message 21% faster and cut its 99th-percentile latency from 82 ms to 58 ms. The third, one read transaction per GET, didn't help on CRuby: the GET pages came out flat to 1.6% slower, probably because the extra BEGIN and COMMIT cost about what the snapshot saved. That's what "evaluate and measure" is for.
The plan
Ordered by the size of the gap, each step measured on the same harness:
- The room page's extra 11 ms. Profile it first; it's the largest single gap.
- The avatar route, which costs 1.5 ms even with a single client and barely improves with more cores.
- Gzip spliced from pre-compressed fragments. Tobi found gzip took 1.9 ms of his 2.4 ms room-page request. His room page also compresses to about 27 KB where Rails' and ours compress to 43–44 KB. As If observed that a fresh CSRF token in every form is what keeps Rails pages from being cached whole; the same per-form tokens are a likely reason the Rails page compresses worse. That's worth confirming before copying anything, because it would put the difference in the gray area rather than the invisible row.
- Allocation in rendering. This decides both the request cost and how well a Ractor version would scale, since Ruby's garbage collector stops every Ractor at once.
- Falcon and Ractors last, once each request is cheap enough that the server model is what matters.
Each technique lands once, in the shared lowering or runtime, so every app that goes through Roundhouse gets it, not just Campfire. The longer-term ledger is unchanged: a per-target account of how much of Rails transpiles, with the unsupported list driven down. What changes is where I'll spend my time first: making the Ruby that comes out fast enough that the comparison with a rewrite is close.
As If ended with John Backus, whose team feared that a FORTRAN program "only half as fast as its hand coded counterpart" would doom the whole idea. "Getting close may be good enough" is the same bar, and this time it comes from someone who has already written the fast version.