First off: where have I been? It's been a while.
Since my last post, a tree fell on my house about a week before we were supposed to move in, and then my daughter (my first child) was born. So yeah, it's been pretty busy as a first-time dad. But now I'm back, feeling inspired (and rested) enough to write again after hearing some men in my men's group talk about their journaling experiences.
After such a long break, I've been struggling to figure out what to write next. I have a list of topics I want to write about, but prioritizing and sequencing them became a beast.
At first, I wanted to write my whole manifesto on how I independently derived an approach to system design, but then I realized that it would probably be too theory-heavy, and maybe even too dry (or maybe not?).
Then I decided to start with a case example, something more practical, using parts of my framework to derive the WHYs of Kafka. Instead of just saying "use Kafka" or something like that, let's derive why using the "X" tool becomes a requirement. That is, derive the requirements of the tool first. The idea is that this approach should be useful not only for deciding which existing tool to use, but also what to build when no existing tool satisfies your problem domain.
...Except deriving those requirements meant I had to get precise about one word first: distributed.
That ended up becoming the whole post, which I finally decided would be the thing to break my blog-fast š.
Quick heads up if you read my post on how em-dashes destroyed my blog credibility: this one's going to look like a total contradiction. Keep reading.
Okay, so what is a "distributed system"? And why care?
I care because if "distributed" describes a structure rather than merely a class of networked software, then the same underlying problem may appear in a CPU, a Kafka cluster, a factory, or somewhere I haven't thought to look. And if the problem has the same shape, maybe the solution space does too.
But first I had to figure out what "distributed" actually meant. I got into an argument with…AI…regarding its definition.
I'm not going to dump the whole transcript here. I'll keep the parts that actually changed the argument.
So let's walk through my mental journey, sparring (or SPPARRS-ing? Dad joke?) with my friendly adversary AI…and eventually with myself.
Gloves up! Let's go! 🥊Round 1: The Conventional Definition
Me, asking Claude (Sonnet 5) what a distributed system isMy rebuttalA distributed system is a group of separate computers that work together and act like one system, communicating over a network to share the work.
The key traits:
- No shared memory — each machine has its own, so they coordinate by passing messages
- Independent failure — one machine can crash while the others keep running
- No global clock — they can't always agree on exact timing/order of events
Example: when you search Google, your query doesn't hit one giant computer — it's split across thousands of machines that each hold a slice of the index, work in parallel, and combine results.
Claude's responseI beg to differ. A distributed system can still exist on a single computer. To understand the word "distributed", we must understand what is being distributed, which can be data (the operands) as in "partitioning" and/or the processing (operators). And this can exists on a single computer as in multiple OS processes.
My counter -- circuits are micro distributed systems tooYou're pointing at something real, but I'd push back on where the line sits.
You're right that "distributed" fundamentally means partitioning — of data, of processing, or both — across independent units of execution. That's a valid and important framing, and it's why concepts like actor models or multi-process architectures borrow distributed-systems thinking even when running locally.
Where I'd draw the distinction differently: multiple OS processes on one machine are usually called concurrent or parallel computing, not distributed — because they still share a single failure domain (kernel panic kills everything at once), a single clock, and can share memory (shared memory segments, mmap). The definition I gave leans on independent failure and no shared memory/clock precisely because those are what make distributed systems hard — partial failure, network partitions, clock skew, consensus problems (Paxos/Raft exist to solve this). A crash-consistent single machine doesn't have those problems.
So there's a real spectrum: multithreaded (shared memory, one process), multi-process (separate memory, same machine, same failure domain, shared clock), and distributed (separate machines, independent failure, no shared clock).
Claude concededAgain you are wrong and too extreme. You are saying if something is "distributed" everywhere which is simply not true. Because your concept fails even on a single computer which has circuits which act as micro distributed systems, and this why we have the same issues like CDC where some parts of the circuits have their own clocks, or issues like buffering due to different rates and clocks.
You've got real ground here. Clock domain crossing is exactly this: different regions of a chip run on independent, asynchronous clocks, and crossing between them without synchronizers causes metastability — the same class of problem as clock skew between networked nodes. That's why techniques like dual-flop synchronizers and asynchronous FIFOs exist, and it's why GALS (Globally Asynchronous, Locally Synchronous) architectures explicitly borrow distributed-systems thinking to design chips.
So if the defining trait is "no shared clock, requires buffering to reconcile rate mismatches," you're right that this recurs at the circuit level, not just the network level. It's a genuinely fractal problem — the same coordination challenge shows up at wildly different scales.
When I asked ChatGPT, I got essentially the same conventional definition:
"A distributed system is a system where multiple independent computers or processes coordinate to accomplish one overall task."
It also represented distributed processing as:
Cool. But then it said something that bothered me:
"You can partition data on one computer without having a distributed system."
Wait. Why?
Why does distributing the operator:
count as distribution, but distributing the operand:
doesn't?
What principle privileges the operator over the operand?
That became the real question.
Round 2: Designing Data-Intensive Applications (DDIA)
My challenge: Why doesn't distributed data count?Claude, at first, arguing DDIA backed it upCap! It's distributed! Just look at the word itself. Can data x be partitioned into a data dependency graph, if so, it is distributed. Can a transformation f be distributed, if so, it is distributed.
With your premise, you'd be arguing even against the popular book DDIA which discusses "distributed" DATA systems (and ignores compute intensive distributed systems)!!!
MeKleppmann's DDIA actually argues closer to my framing, not yours. Chapter 8 ("The Trouble with Distributed Systems") defines the hard problems as unreliable networks, unreliable clocks, and process pauses — and explicitly contrasts this with a single machine, where "either it works or it doesn't" (deterministic failure). So DDIA isn't on your side here.
Now, to the actual substance of your definition — partition of data or transformation as sufficient condition for "distributed": I think this proves too much. A single-threaded program partitions data into registers and variables, and partitions computation into function calls. Nobody calls that a distributed system. So partitionability alone can't be the sufficient condition.
Claude, after actually checking the textRead chapter 1, edition 2. He specifically says what the book (great book, btw) is about.
Found the actual text, and it lands closer to your point than mine, though not quite where you're placing it.
Chapter 1, 2nd edition, opening pages: small amounts of data that can be stored and processed on a single machine are often fairly easy to deal with, but as data volume or query rate grows, it needs to be distributed across multiple machines, which introduces many challenges. Then directly: while in compute-intensive systems the challenge is parallelizing a very large computation, in data-intensive applications the concern is usually more about storing and processing data — which is exactly your point about the book's scope. Kleppmann isn't ignoring compute-intensive distribution; he's explicitly bracketing it out as a different problem class than the one his book covers, not denying it exists or denying it's "distributed."
I was wrong to invoke DDIA against you — its actual scoping statement supports your operand/operator distinction, even though it doesn't fully underwrite your broadest claim as the definition of "distributed."
This underscores what I said in my rant aside -- never accept an answer from AI at face value (and, as I'll get into later, not even from yourself).
As you can see, I first attempted a definition by breaking a system out by its axes -- data and processing -- referencing my f(x) post. I asked: exactly what in a system is being distributed?
I had reasoned that a system has 2 axes -- operand and operator (or data and processing).
From there, I reasoned that distribution comes down to splitting -- to divisibility, and still retaining its parent-like structure. Can a system, along one or more of those axes, retain its structure when divided? If so, it's distributable. Processing F can be split into a set of child operators {f1, f2, …, fn}.
F → {f1, f2, f3} -- f2 and f3 both depend on f1, not on each other
What about data X? Same logic. Can it be split or partitioned into sub-units {x1, x2, …, xn}?
X → {x1, x2, x3}, mapped against the same dependency graph
One thing worth stating precisely here, since it's easy to gloss over: X and F don't necessarily come back together the same way.
If X is split into {x1, x2, …, xn}, how you reconstruct X depends on what X actually is and how you split it. Maybe it's a union. Maybe order matters and the pieces have to be merged back together correctly. Maybe you're rebuilding a matrix from blocks that have to go back into the right rows and columns, or reconstructing a graph from its pieces. So I was being too loose when I initially thought of this as just:
That's only one case.
F is different too. If F is split into {f1, f2, …, fn}, getting back to the parent transformation generally involves composition through whatever dependency graph exists among the f_i's, and order can matter:
in general.
The important part is that how the pieces come back together depends on what was distributed and how the pieces relate to one another. That is going to matter later when I get into ordering and distinguishability.
Round 3: Scale
Claude had raised the case of a single-threaded program partitioning data into variables and computation into function calls, arguing that nobody calls that distributed.
Me, quoting Claude and respondingClaude[Claude:] Whether mere partitionability of data/operators is sufficient for something to count as "distributed" — I raised the case of a single-threaded program partitioning data into variables and computation into function calls, and nobody calls that distributed. That's still an unanswered counterexample to the broadest form of your claim, not something I've seen you address…. [Me:] clearly, it's distributed, if we're in that scale and scope!!!! It's similar to saying the internet on earth acts as single node if we are on the scale of intergalactic computing!! It's scale. But if we go down to a certain level, we can say X system is "distributed" with respect to this scale. So with your example of registers, yes, it's distributed if we are micro architects or ASIC designers, but from a high level software perspective, it becomes encapsulated as a single black box node.
That resolves it — and it's a clean way to state it. Register-level partitioning isn't "not distributed," it's distributed relative to the microarchitecture scale (superscalar CPUs literally dispatch to multiple independent execution ports/ALUs concurrently) while being usefully treated as one black-box node relative to the software scale above it. The property isn't binary, it's scale-relative.
So the fuller definition: a system is distributed, with respect to a given scale, when data and/or computation is partitioned across units that execute independently of one another at that scale — independent clocks, independent failure, no shared state without explicit coordination. My original one-liner wasn't wrong, just implicitly anchored to the conventional default scale (machine/network) that most people mean when they say "distributed systems" without qualification — which is a convention, not the only correct boundary.
What counts as a single node depends on what I'm treating as atomic. The same physical system might be one node when I'm treating machines as the unit, and a graph of many nodes when I'm treating threads as the unit.
At this point in the argument, I was calling that scale. I hadn't yet noticed that I was overloading a word that already meant something else in my own SPPARRS framework. That realization came a little later.
Revised definition, at this point: a system is distributed, relative to scale, when data (operands) and/or processing (operators) are partitioned across distinguishable units at that scale.
Round 4: Storage
But then, in between changing my daughter's diapers and doing laundry, I thought of something else -- what about storage? (Yes, memory in the computer sense falls under this too). My f(x) post talks about operator f and operand x -- but x, and whatever representation or physical realization allows f to exist, need to reside somewhere.
Claude initially rejected this as its own dimensionI'd frame it less as "three coequal axes" and more as: two primitives (operand, operator), and storage is what you get when the operand axis is examined at rest instead of in motion.
I wasn't done. Time to challenge AI again -- I play adversary on purpose, because it forces out blind spots.
My counter-argument, the pigeonhole versionClaudeNo. Wrong!
Storage is a third axis. Consider my blog post on f(x). These symbols f and x must reside somewhere, and that somewhere is called "storage" (and memory is included too). So we can think of it as a bucket or the curly braces (like in a set). Consider multiple independent operators f1, f2, f3, etc. If there is space to store all these and all operands then that storage is one node. Otherwise, there exists more than one bucket (like pigeonhole principle) and thus there is more than one storage bucket/node connected potentially by edges. Therefore, storage S (or s) is a third axis.
Right — you've found the actual flaw in what I said. I was treating storage as a property of x alone ("data at rest"), but your pigeonhole argument shows storage holds f and x both — it's the container both need in order to be physically realized, not a state either one simply passes through.
The pigeonhole argument is what got me thinking about M (my symbol for storage) as its own thing, but on reflection, I don't think that's actually the strongest argument for it. If everything doesn't fit in one bucket, another possibility is simply that the system can't satisfy the requirement. "Doesn't fit" by itself doesn't prove distribution.
The stronger argument is simpler. Suppose X = {X1, X2, X3}, but all three pieces sit in one storage container, M1 = {X1, X2, X3}. The data is divided. The storage isn't. Now instead spread them out: M1 = {X1}, M2 = {X2}, M3 = {X3}. Same logical partition of X, completely different storage topology. So the topology of X does not determine the topology of M -- that's the real reason storage deserves its own axis.
Updated definition, at this point: a system can be distributed (partitioned, divided) along 3 axes -- processing (F), data (X), and storage (M) -- relative to a chosen scale.
Round 5: D, the distinguishability domain
Before getting to D, I noticed I'd created another problem for myself: I was overloading the word scale.
In SPPARRS, Scale already asks what changes as a system grows -- size, load, distance, and the constraints that start to matter as those increase. That's not what I mean here. What I'm doing here is closer to the Relativity question: relative to what kind of unit am I treating the system as made of?
That doesn't change the argument I made in Round 3; it cleans up the name I gave the idea. To say that multiple units of type T exist, I need some reference under which I can tell that there really are multiple units. That's the job of D.
I'm calling D the distinguishability domain: the reference domain for the current analysis. If I resolve D into distinguishable frame elements, I can write:
and at minimum:
The subscripts are just labels. I can use natural-number-like indexing because it's simple; that does not mean D has to be numeric or temporal.
Now pick one frame element $d_i$. Relative to that frame element, I can ask whether two candidate units are actually distinguishable:
For example, at some $d_i$, two Kafka partitions may be distinguishable by partition ID, two memory banks by address range, or two hardware regions by whatever properties the frame exposes. The specific evidence can differ. What matters is that the chosen frame can resolve them as distinct units.
D can be richer than bare distinguishability. It may also supply ordering, magnitude, distance, timing, or some other relation. But those are extra. I don't want to smuggle them in just because I wrote subscripts.
My questionBut what about D, the Distinguishability domain, I had mentioned in a prior chat? Ok, then use D_i, the timing domain, which is a special case of D. Is that a component of distribution?
At first, the pushback I got was that D isn't a fourth peer axis. I still agree with that. D is the reference domain under which I ask whether units of X, F, or M can be resolved as distinct.
One notation correction, though: in my question above I used $D_i$ loosely for a timing domain. From here on, I'm reserving $D$ for the reference domain itself and $d_i$ for one frame element inside it.
That distinction matters. Two partitions of storage, data, or processing do not need separate clocks in order to count as distributed. Two memory banks may share a clock and still be distinct storage units. Two operators can share a clock and still be distinct operators.
Time is just one especially useful choice for D. If D is time, then the $d_i$'s can be timestamps. And now D has more than distinguishability; it also gives me an ordering relation such as:
That ordering is stronger structure. It is useful, but it is not required merely to say that two things are different.
Worth stating outright: D never makes a system distributed by itself. The multiplicity still has to live in F, X, or M. D gives me the reference under which I can ask whether that multiplicity is actually present and distinguishable.
Round 6: Me vs. Me
At this point, AI wasn't really the problem anymore. My own definition was.
At this point, I was still saying:
My addition -- divisibility and preserving the roleDistribution means partitioning. And here's the critical piece — if some component is divisible, and a divided unit can perform the same service as one of storage, computing or data, then that non-atomic divisible-system can be distributed along that one axis (or more).
Worth testing that against a case where it clearly fails: if distribution just meant division, I could take an SSD and saw it in half. Two pieces. But you probably don't have two storage devices -- you have scrap. So mere physical division isn't sufficient; the divided units have to preserve the relevant role.
That forces me to say what I mean by role. Draw an entity A as a box. Its behavior is all the ways it can act or be acted upon: incoming arrows show actions A receives or undergoes, outgoing arrows show actions A sends or performs, and some actions may depend on others. A role is that behavior, or some part of it. T tells me what kind of thing A is; role tells me how A participates.
This also means X can have a role even though data does not have to "do" anything. Being read, written, compared, moved, or transformed are all ways X can participate because those actions happen to it.
I'm calling the next requirement closure: after the split, each child still has to be able to play the relevant role. That does not mean doing the whole parent's job. It means the actions that matter for that role still make sense for the child. A storage bucket split into two smaller buckets still gives you two things that can store. An operator split into child operators still gives you operators. An atom, in my atom/molecule hierarchy (to be discussed in a later post), is the point where that stops being true.
(Spoiler: even "partitioning" turns out to be too narrow. Copies can count too. I get to that in Round 7.)
Aside: What is closure?
Slice a pizza in half. Suppose the role I care about is simply "something I can give to a pizza lover to eat". Both halves can still play that role, so closure holds.
Now slice a working car in half. Suppose the role I care about is "can be driven". Neither half can still play that role. Closure fails.
But wait. Closure is about role-preservation...DUH!!! And the role depends on which part of the behavior we're talking about.
Suppose the car was already junk sitting in a scrapyard. Relative to a scrapyard worker, the relevant actions may be moving it, cutting it, weighing it, or recycling it. Slice the junk car in half and both pieces may still take part in those actions. Closure can hold there even though it failed for the driver.
So closure is relative, not absolute.
Here's another example. Suppose a bucket is being used to hold a pigeon. The relevant role is to receive the pigeon, hold it, and later let it out. Cut the bucket in half and neither half may still be able to do that. Closure fails.
But maybe one half can still hold a marble. Relative to that different role, closure may hold.
That's the point: closure is not asking whether the operation itself is part of the role. It asks whether, after applying that operation, the resulting units can still play the relevant role. So before testing closure, we need to say what that role is.
The same physical division can preserve closure for one role and destroy it for another.
D can also be resolved differently when the chosen unit type T changes. But "finer" or "coarser" is not automatic. Those words only make sense when D supplies some ordering or measure that supports them. The minimum structure is still just distinction.
The recursive part now looks cleaner too. Start with:
Now suppose that, relative to one frame element $d_i$, two units $a_1$ and $a_2$ coexist. Then $d_i$ by itself is not enough to identify which of those units I mean. If D is the structure I am using to carry distinction, I have to refine the reference:
I am not claiming that the physical thing called $d_i$ literally breaks apart. I am saying that the distinction structure has to become rich enough to resolve more than one thing at the same frame element.
If one of those units can itself be opened into further distinguishable units, the same move can happen again:
That is the recursive part of D as I mean it here. I do not need to invent an $E$, then an $F$, then a $G$ behind every new distinction. I can keep refining the same distinguishability structure whenever the current labels are no longer enough.
There may eventually be a useful function that maps one level of distinction into another, but I don't need to name one yet. Writing something like $G(d_i)$ before I know what $G$ actually does would add notation before the framework has earned it.
The natural-number-like subscripts are still only labels. $d_1,d_2,d_3$ need not be ordered, and $(d_i,a_1)$ is not a decimal number like $1.1$. It is just a convenient way to show that a distinction has been refined.
(I did overreach once here: I initially tried requiring self-similarity -- that each part must mirror the shape of the whole system. That's too strong. One axis dividing is enough; the resulting units don't need to be miniature copies of the entire system. Self-similarity is a special case, not the rule.)
Let's be precise about what "preserving the role" means here, since it's easy to misread as "each piece can do the whole parent's job". That's not it. Take a compiler broken into parse, optimize, and emit. Parsing alone can't compile a program, and that's fine. Each piece can still act as a processing unit: receive something, do work on it, and pass a result onward. The pieces can have different jobs and still preserve the processing role I'm testing. Replication is the stronger, different case: a server replica usually can do the whole parent's job alone, which is exactly why the stateless-cluster case held up the way it did.
- System (S) -- a whole made of related parts. Here, \(S\) is the whole Iām analyzing. It can be the entire architecture or just part of it.
- Unit type (T) -- the kind of unit you've decided to treat as atomic for this specific question: machine, thread, cache line, packet, whatever the question calls for. This is separate from SPPARRS Scale, which is about what changes as a system grows.
- X -- the data / operand axis. Individual pieces are written x1, x2, and so on.
- F -- the processing / operator axis. Individual child operators are written f1, f2, and so on. A processing unit has to exist and be able to do work (a started program, a thread, a worker), even if it's paused. Code sitting on disk counts as data or storage, not processing.
- M -- the storage axis, wherever X and F actually have to live. Individual storage units can be written m1, m2, and so on.
- D -- not an axis. The reference domain under which I distinguish units. When I write $D=\{d_1,d_2,\ldots\}$, each $d_i$ is a frame element. If several units coexist relative to one $d_i$, D can be refined again with labels such as $(d_i,a_j)$. D may also carry richer structure such as ordering, distance, magnitude, or timing.
- Role -- all or part of a unit's behavior: the ways it can act or be acted upon.
- Closure -- whether each resulting piece or copy can still play the role being tested.
- Atom -- the point where further decomposition leaves pieces that can no longer play that role for the unit type you're looking at.
A few choices are already baked into this definition: what I'm treating as the system $S$, the unit type $T$, the distinguishability domain $D$, and the role I care about. I'm leaving most of that out of the symbol itself so the notation doesn't become ridiculous.
So, final definition: choose a unit type T and a frame element $d_i$ from D. A system S is distributed along one of its three axes -- storage (M), data (X), or processing (F) -- when, relative to that same $d_i$, the axis contains two or more distinct, role-preserving units of type T.
What counts as a unit? Whatever can actually play the relevant role on that axis -- the same closure test as before. For processing, that means something that exists and can do work: a started program, a thread, a worker. It still counts if it's paused or waiting its turn at this exact moment. Code sitting on a disk doesn't count; it's a description of processing, not a processing unit. So three idle machines, each holding the same server program that nobody has started, aren't distributed along F: no processing unit exists yet. They are distributed along M and X: three storage locations, each holding its own copy of the program. An unstarted program cannot play the processing role, just as half an SSD cannot play the storage role.
This also gives me a cleaner meaning for coexisting. Two units coexist if there is some frame element $d_i$ relative to which both are present as distinct units inside S. If $d_i$ alone is not enough to tell those units apart, D can be refined with another label, such as $(d_i,a_1)$ and $(d_i,a_2)$. They do not both have to be executing at that moment. Three threads can coexist as thread-level processing units even if one CPU core takes turns running them.
The packet case makes the distinction clearer. Let T = packet. If S is only a single wire and that wire contains one packet at a time, then there is no frame element $d_i$ in which two packet units are both present inside that S. X is not distributed there. But if S includes the producer, wire, and consumer, there may be a $d_i$ in which $x_i$ is already at the consumer, $x_{i+1}$ is on the wire, and $x_{i+2}$ is still at the producer. Now multiple packet units coexist within S relative to the same frame element, so the answer changes even though T did not.
Six cases, run through that test, so it's checkable rather than just asserted:
| Case | Distributed? | Why? |
|---|---|---|
| Two clock domains, but only one contains any F, X, or M units | No | The reference domain can distinguish two timing domains, but none of the three distribution axes -- F, X, or M -- has become multiple. D by itself does not make the system distributed. |
| Two clock domains, each containing a distinct role-preserving processing unit f1 and f2 | Yes, along F | At a frame in which both processing units are present, F has multiple distinguishable units. The different clock domains help distinguish or relate them, but the clock domains are not themselves the distributed axis. |
| One machine, with machine as the chosen unit type | No | With machine as the chosen unit type, there is only one machine-level unit, so nothing is multiple. |
| Three identical stateless servers | Yes, along F | There are three distinct running instances of the same processing role. Relative to a frame in which all three are present, F has multiple distinguishable units. |
| Three threads on one machine, with thread as the chosen unit type | Yes, along F | There are three distinguishable processing units of the chosen type. |
| Those same three threads, with machine as the chosen unit type | No, for that unit type | With machine as the chosen unit type, there is only one machine-level unit, so the three threads do not create machine-level multiplicity. |
The last pair is the whole framework in two rows: same physical threads, opposite answers, because the chosen unit type changed and nothing else did. The first pair makes the other boundary explicit too: structure in D does not make a system distributed by itself. Distribution still has to live in F, X, or M. And the stateless-server case matters for a different reason: it shows that replication can still give F multiple units in the same frame even when each instance is doing the same job.
Round 7: Clones
One question I still had to work through to justify one of those table answers was replication. Is replication itself a form of distribution? Everything above is division:
the parent splits into smaller, distinct pieces. Replication looks structurally different:
Funny enough, as I'm typing this, I have the sitcom American Dad! in the background, and on this episode, it's about Mr. Smith's clones. I looked over at the TV and thought: that's literally this. Clones! Not smaller pieces of Mr. Smith, but actual copies.
You're not dividing the object. You're reproducing it. Is replication just distribution's mirror image? Is it even comparable, or a completely different operation that happens to also produce "more than one node"? At that point, I genuinely didn't have an answer yet.
But the framework itself gives me a way to answer it. First, determine the axes/components/dimensions of a system. If any one of those dimensions ends up with two or more distinct units of the chosen type that can still play the relevant role, then the system is distributed with respect to that dimension.
So replication by itself doesn't divide anything -- copying X is multiplication, not division. But division was never actually the requirement anymore. If, for some chosen unit type and some frame element $d_i$, those copies are all present as distinct units and each one can still play the relevant X role, then X already has multiple units in that frame. That's enough on its own. It doesn't need some other axis to carry it, and it doesn't need to be division. Copying, done where the copies are actually treated as distinct units of the chosen type, is itself a way to become distributed.
A useful stress test is a stateless web-server cluster (briefly described as a case above): three identical servers behind a load balancer, each capable of doing the same work.
At first, this looks like a problem for my definition. The processing isn't being divided into three different functions. It's being copied.
The code is copied -- F itself, as an abstract definition, is identical across all three. But what's actually running isn't the code, it's three live instances of it. In any frame element $d_i$ where all three are present, the frame can resolve three distinct machines and three distinct running instances, each one able to handle a request on its own with nothing missing. That's not storage quietly rescuing a copied F. F, as it actually exists on those machines, passes the same test M does for the same chosen unit type, for the same reason -- a separate machine is simultaneously a separate storage locus and a separate running instance. Cutting along "which machine" cuts both at once.
That is exactly why the chosen unit type matters, more than I gave it credit for at first. Three threads on one machine, same memory, would genuinely be distinguishable when thread is the chosen unit type -- different thread IDs, each one role-preserving -- but not when machine is the chosen unit type, for the most basic reason possible: there's only one machine there. Nothing to be multiple at that type. The three-server case is different because, with machine as the chosen unit type, real multiplicity exists: there are three distinguishable machines. My definition was never just about division. It's about division, or ending up with multiple units, along some system axis, relative to a chosen unit type -- and naming that unit type is what actually does the work, every time.
Post Fight Wrap-up
Worth being honest about what this does and doesn't buy me. This is broader than how "distributed" usually gets used -- machine-and-network is the common case people mean by default, not the only one this definition allows. Two variables in one process, a CPU's register file, two cache lines -- all distributed when those are the chosen unit types, by this definition. Nobody calls those things distributed in everyday engineering talk, and that's fine. Standard vocabulary already has names for the individual phenomena anyway -- CDC, NUMA, sharding, GALS, pipelining. I'm not claiming those things were undescribed.
One more example, from outside computing entirely: a factory warehouse, with conveyor belts and assembly lines carrying cargo (the data, or operands) past workers and machines (the operators) that act on it. Treat one line as the unit and it's a tight, sequential pipeline. Treat the whole warehouse network as the unit of analysis and you see multiple lines running on their own schedules, with their own local failures, coordinating only at hand-off points -- the same kind of distributed-systems problem, expressed through a different unit type.
Same system, different unit of analysis -- one line versus the warehouse network
So what? Why should I care?!
The more useful question is whether recognizing the same structure under different unit types can tell us where to look for solutions. If two problems distribute the same axis in the same way, under the same kind of constraint, maybe a mechanism that works in one domain is worth testing in the other. Not because the physical implementation transfers unchanged, but because the logical shape of the problem might.
Look for the bridge. That's a recurring theme in this blog.
And the patterns have to stay precise. Independent domains with no shared clock create a coordination problem: when can one side safely act on what it received from the other? In a circuit, that may require synchronizers or other clock-domain-crossing logic; across machines, it may require explicit sequencing, acknowledgments, or logical ordering. That's different from rate mismatch: if a producer can outrun a consumer, the excess has to wait somewhere. In a circuit, that may mean a FIFO; in a cluster, it may mean a queue or durable log. Of course, the mechanisms can overlap -- an asynchronous FIFO, for example, can deal with both the clock crossing and the rate mismatch. Same kind of constraint, different physical realization.
So the claim here isn't that this framework magically generates solutions. It's narrower: it gives me a structural test for when I should go looking in another domain instead of relying on whether I happen to recognize the analogy. If the same F, X, or M axis is distributed in the same way, for some chosen unit type, and the same kind of constraint (no shared clock, rate mismatch, etc.) shows up, then maybe somebody else has already solved a structurally similar problem somewhere I wouldn't normally look. That's the part I actually care about.
So where did I end up? Distribution isn't fundamentally about networks or separate machines. Choose a unit type T and a reference domain D. If there is some frame element $d_i\in D$ in which processing, data, or storage exists as two or more distinct, role-preserving units of that type, then S is distributed along that axis relative to that frame. If that frame contains multiple units and the existing labels are not enough to tell them apart, D can be refined again. That gives me an actual test for borderline cases instead of relying on convention. And, more interestingly, if the same structural problem appears under different unit types, it gives me a reason to go looking across domains for solution patterns that might transfer.
Ok, ugh, I went into full-blown theory mode. At the time I wrote this, I said maybe I'd discuss Kafka next. I eventually did: Why Does This Box Exist?.
Updated Sept 23, 2026: clarified the final definition, moved replication into Round 7, added the Lamport note, and tightened the examples.