What Survived Contact
I'm back, and I said I'd report once I found out how much of what I brought survived contact with a room that knew the subject better than I do. First, the easy and entirely true things. The venue was gorgeous. The organizers were professional in the way that only looks effortless — the unconference ran so smoothly you forgot how much arranging that takes, and when more people wanted to lead sessions than the schedule could hold, proposals were capped only after they'd first made sure everyone who wanted a slot got at least one. I was made welcome, which I'd expected. What I'd wondered about going in was whether the ideas would be — a sample of one is exactly the kind of thing a room like this exists to stress-test — and they were: sought out, argued with, taken seriously.
I also brought my laptop, and Claude was quietly coding in the background throughout the sessions. This proved useful on multiple occasions: when an incredulous participant probed how I was using Claude, I could show live examples straight from the work in progress on my screen.
At one point a participant suggested that it was important to build a model before beginning. I countered that my approach was to flip this around: I asked the LLM to propose a model and explain it to me in terms I could understand.
A note on length before I start: this one runs long. The event was a firehose that ended way too soon — there was no way to attend every session, and the ones I did kept opening topics I wanted to go deeper on. And inevitably some of the ideas below only came to me afterward, once I'd had time to digest the week. So I've let this post run long rather than cut it; I'd rather leave nothing out than keep it short.
Twenty-one weeks: what changed, and what didn't
Before I left I'd held my one project against the findings of the retreat that preceded this one; coming home, I did the same thing at the level of the whole event — laid the public summary of that earlier retreat next to the fuller record of this one, twenty-one weeks and a continent apart, and looked for what had moved. I'll keep the same discipline I keep about my own project: I sat in this room and only read the summary of the other, so I can't always tell what is genuinely new from what simply didn't survive the compression into a public report. Take what follows on face value, with that seam showing.
The spine didn't move, and it's the spine I care most about because it's the one I keep arguing: rigor doesn't vanish when the agent writes the code — it migrates. Upstream into specifications, down into test suites treated as first-class artifacts, into type systems and constraints, into tiering code by how much damage a mistake could do. Both retreats land there, months and venues apart, which is about as close to a settled finding as a field this young produces.
What changed, on face value, is the scope. The earlier summary stays almost entirely inside the engineering organization. This retreat spilled outward — into economics, into ethics, into geopolitics: sovereignty and whom you trust to host a model, regulation and antitrust, environmental and human costs, the maintainers holding up infrastructure nobody pays for. Some of that is surely the room. This retreat was in Europe and the earlier one was not, and the sovereignty-and-regulation thread carries a more international crowd's fingerprints. But I can't cleanly separate the conversation grew from the public summary of the other one left this part out, and I won't pretend I can.
Where the two overlap, the engineering has aged fast. Things the earlier summary placed a year or three out — the supervisory layer between writing code and shipping it, agents as first-class participants in an org chart, decades-old semantic machinery pressed back into service as grounding for domain-aware agents — showed up here carrying production numbers instead of speculation. Even the furthest-out bet, systems that heal themselves, had started to move — but only halfway: the diagnosis is automated while the actual remediation is still held back from the machine, and stuck for precisely the reason the earlier summary had named: diagnosing a fault is safe, but letting a machine act on the fix unsupervised hands it the blast radius, and no one is yet ready to cede that. The forecast compressed unevenly, fastest where the earlier room had been most cautious.
And the biggest change is epistemic. Last time the mood was speculative in both directions; this time both the optimism and the worry arrive backed by data. But not equally, and the asymmetry is worth naming. The optimism rests mostly on uncontrolled, single-team, self-reported metrics — this harness cut our tokens fourfold, this pipeline shipped in hours — the kind of number I publish without conclusions myself, because a sample of one is exactly what it is. The worries lean on sturdier ground: outside studies, real surveys, measured trends. So it is "backed by data" on both sides, but the doubt is currently better evidenced than the hope — which, if it holds, is the uncomfortable thing to watch as more numbers come in.
One more thing about that scope, and it's the part I want to be careful with, because I'm about to spend two sections on how small my own sample is. The room is a sample too, and not a random one. A gathering like this — invited, senior, convened by a consultancy — leans heavily toward large, established organizations and the people who advise them, and lightly toward the startups whose reason for existing is to unsettle those organizations. That isn't a knock on the organizers; it's who comes to a room like this anywhere. But it tilts what gets written down. Big institutions are superb at the thing big institutions do — optimizing what they already have — and a room made mostly of them and their advisors will reach for how do we govern this, verify this, tier it by risk long before how do we use this to make someone else's advantage evaporate. The disruption was in the room; it was just discussed from behind the walls rather than outside them. Which is one more reason the risk-and-rigor register came through as loudly as it did — and one more reason to wonder what a room with twenty founders in it would have put on the board instead.
A hire, not a tool
Above is the room's ledger. Here is mine, and I'll be honest about what it weighs: one project, one retired developer, one agent — an anecdote, and I've marked every place it can't bear weight. What I won't pretend is that it left me anything other than firmly optimistic.
The session I'd put on the board was bring me a rock — exploration by elimination, the management dysfunction that turns into a method once an iteration costs minutes instead of days. The room pulled it somewhere narrower than I'd framed, and the narrower place was the more interesting one: not how to explore by elimination but who should even be allowed to. Product managers, increasingly people managers, are reaching for these models directly, and seasoned engineers get measurably better results from them than untrained people do — so the worry followed. If expertise is what separates a good outcome from slop, should non-engineers be steering the model at all?
It's a fair question, and I think it's the wrong one, because it mistakes the act. When a manager reaches for an LLM instead of routing the work to the team that reports to them, they didn't pick up a tool — they made a hire. And you don't ask permission to manage your own team; a manager who decides a piece of work is better given to a new participant than to the existing one is doing the most ordinary thing a manager does. Framed that way, the permission question dissolves into an older, better-understood one — the one Drucker named in 1959: when the worker knows more about the specifics than the manager does, you manage by objective, not by method. The non-engineer steering an agent is exactly that manager, out-known by the thing they're directing, and the slop the room feared is the old danger of managing by method when you should be managing by objective. The question isn't may they hire? It's do they know how to manage by objective? — which you can teach, hire for, and hold people to without anyone first becoming an engineer.
The two objections the room raised were both fair. We're regulated; we can't afford the risk — real, and also the exact sentence I heard about open source nearly thirty years ago, from someone whose call was defensible on what he knew and who was still guarding the wrong artifact; the durable governance never lived in inspecting every line, it lived in governing the thing the code answered to, and I'd bet the same relocation here. The model is non-deterministic — also true, but equally true about people, particularly once you exceed Dunbar's number — and it's the one that finally made me close a seam I'd carried since I first called the agent a peer.
Because this is the thing I actually changed my mind about. I'd worried, out loud and more than once, that a colleague has stake and the model doesn't, and that a peer without one might be a diminished peer. I no longer think that's the right worry — or rather, it's the right worry about the wrong mode of use. Hand the model a task — do this specific thing — and the missing stake is exactly what you feel: it does what you said without caring whether what you said was worth doing, and the caring was the part you wanted. But a task is the wrong thing to hand it. The model is at its best given a goal — an objective it can test itself against — and a thing driving toward a goal you authored behaves the way a stakeholder behaves: it tries, checks itself, throws out its own bad rocks, and keeps going until the goal is met. The stake I thought was missing was only ever missing from the worker. It lives in the objective. Manage by task and its absence is real; manage by objective and the objective supplies it — in a form you can actually check, which is the only form of it you could ever have relied on anyway.
How to evaluate an objective
Everything above rests on a load-bearing if: manage by objective only works when the objective is stated so it can be checked. An objective you can't evaluate is a wish, and a wish is exactly the stakeless task I just said to avoid. So the real question — the one the room and I circled from opposite sides — isn't what is the objective, it's how is it evaluated. That's the whole game, and it's where I want to be careful about a term the room reaches for and I've had to bend.
The room's comfortable word for this is BDD — behavior-driven development, specification by example — and I'll gladly adopt it, because it names the right move: state the objective as its own evaluation. A Given/When/Then scenario isn't a description of the behavior sitting next to a test of the behavior; it is the behavior, written so that it runs. That's what sets it above the other places the room agreed rigor goes. Prose specs can't run, so they can never be their own oracle, and ambiguity is their default. Types are inferable — the compiler closes them for free, so they're not where a human's rigor lands. Constraints and risk tiers are downstream; they presuppose you already know the behavior you're protecting. Of the lot, BDD is the only one that is the objective rather than pointing at it.
Where I've had to bend the term is in how the evaluation gets its answer. In classic BDD you author the expected result — you write down by hand what "done" looks like for each example. That's fine, and for genuinely new behavior it's unavoidable, because the intent lives only in your head and someone has to put it somewhere. But it's example-based, and examples cover only the cases you thought of — which is exactly why the room kept reaching past them, to property-based testing and to characterization suites mined from what a system actually did in production.
On Roundhouse I got to skip the authoring, and the reason is worth stating precisely. I already had a running reference — the Rails application itself — so I didn't have to write down what "done" looked like; I could ask the thing that already knew. The acceptance gate fetches the same URL from Rails and from each generated target and requires the responses to be byte-for-byte identical. That's still Given/When/Then in its bones — a seeded state, a request, an observed response — but the then isn't authored, it's referenced: the expected answer is whatever the reference produces, computed fresh for every input. Which frees it from the cases I happened to imagine. I can throw any request at it — enumerate them, replay real ones — and the oracle still knows the answer, because it isn't a list of answers, it's a way of generating them.
So this is just BDD pointed at a system that already exists — the standard black-box move for rebuilding something you can't fully read. The one twist worth naming is the then. The textbook version records the reference's answers once and freezes them into the examples — a characterization test in Gherkin's clothing, covering only the cases you recorded. I never froze them. The gate queries Rails live, per request, so "matches the reference" holds over inputs I never wrote down — the difference between checking forty scenarios and checking every request there is.
Which is the one thing, in the end, that the person doing the hiring can't hire out. If the building is delegated, the targets compiled, the types inferred, the implementation generated and discarded a dozen rocks at a time, what's left to author is the objective and the choice of how it's evaluated — the oracle. It's the fitness function the keep-the-candidate-that-survives game runs on, and the manager who hires the model can no more hand it back to the model than a manager can hand the objective to the worker being managed by it. It's also the room's permission question answered from the far end: you don't forbid the non-engineer from hiring the model, you require them to own the objective and its evaluation — because that's what makes the hire a hire and not a prayer.
The unfinished map
A sample of one, then, and I've marked its edges. But the principles are the part I'm firmly optimistic about. A collapsed cost doesn't end the inquiry, though; it moves it. So let me leave four questions on the board instead of pretending to answer them.
The first: once building a thing costs less than deciding whether to build it, why iterate serially at all? You could stand up several candidates at once and keep the one that survives — a form of software darwinism, perhaps — and I don't yet know what that does to how we plan, or what it demands of the oracle left to judge the survivors.
The second: when do you evolve the implementation you have, and when do you start over? Roundhouse was my third run at the same problem, and each time I chose to start over — cheap to choose because the oracle stayed constant across all three. What was new this time is that starting over no longer meant starting blind: at each decision point I could send the LLM back into the earlier implementations and ask it to recommend whether an approach I'd taken before was worth emulating or worth avoiding. The prior code stopped being sunk cost and became a corpus to consult — and I don't yet know where the line between evolving and restarting falls once a restart can carry that much forward.
The third: watch how fast the register flipped. A few months ago the sentence in rooms like this was LLMs produce slop, and we can't afford the security risk; today it's the audit the model ran caught things no human reviewer would have. Code review and security audit were supposed to be the last duties you'd hand a machine — the judgment of last resort — and instead they went early. So which of the other best practices follow: testing, documentation, dependency triage, the postmortem after an incident — each a place the non-engineer principal currently leans on engineers, each a candidate to hand to the model instead. The list of things only the human can do keeps getting shorter, and I don't yet know how short it gets — my bet is that when it stops the oracle is the only thing left on it.
The fourth comes from outside the room, and it's about scale. For a few months now, three of us — matz, Ori Pekelman, and me — have been building on three continents, each carrying a substantial piece, coordinating through nothing more exotic than issues and pull requests. If one person really can hold a whole subsystem now, maybe that plain old machinery is enough to compose large systems out of small ones, held by small teams that never have to merge into one. Maybe not. Either way it's a line I'd add to a map that room came to redraw, not to finish — offered in the spirit the map asks for, which is an admission of how much none of us yet knows.
One last thing, and it isn't a formality: my thanks to Thoughtworks, who hosted the retreat and ran it under the rule that lets me carry the ideas out while leaving the names behind — the reason I can write about any of this at all. Making room at a gathering of senior practitioners for a retired outsider with a sample of one was a generosity I didn't take for granted, and the welcome held even when I was disagreeing. Whatever I've pushed back on here, I pushed back as a guest who was glad to be in the room.
Roundhouse is open source: dual-licensed MIT / Apache-2.0. Issues and discussion welcome.