← ::root

Why Does This Box Exist?

Derive the architecture. Don't memorize it.
Highlight any text, then choose Explain, Derive, Example, Extremum, or Attack.

In my last post, I admitted that I originally wanted to write about my system-design framework, then tried to make it practical by deriving Kafka, and somehow ended up arguing with AI about what the word "distributed" even means.

So now I'm coming back for the post I meant to write in the first place.

Open almost any system-design guide and sooner or later you'll see something like this:

$$ \text{Client} \rightarrow \text{Load Balancer} \rightarrow \text{API} \rightarrow \text{Kafka} \rightarrow \text{Workers} \rightarrow \text{Database} $$

Looks legit.

But why is Kafka there? Why the load balancer? Why multiple workers? Why replicas? Why a database at all?

If the answer is some version of "because that's what scalable architectures look like", I don't think the box has earned its place.

No box enters the architecture until the system earns it.

I don't want to memorize architectures. I want to derive them.

The rough rule I've been converging on is:

$$ \boxed{ \text{Physics} + \text{Intention} + \text{Failure} \Longrightarrow \text{required mechanisms} \Longrightarrow \text{tools} } $$

But before I can derive a system, I need to say what the smallest thing in that system even is.

Start With the Box

My first mental model looked like the block diagrams I had seen in software and digital logic:

$$ x_{\text{in}} \longrightarrow \boxed{f} \longrightarrow x_{\text{out}} $$

Input enters a bounded thing. Something happens. Output leaves.

Simple.

But while trying to generalize that into what I call an atom, I kept drawing this instead:

$$ \boxed{ \begin{array}{ccccc} x_{\text{in}} & \longrightarrow & f & \longrightarrow & x_{\text{out}} \end{array} } $$

Notice what I just did.

I put the "input" and "output" inside the box.

That looks wrong if you're thinking in the usual signal-flow sense. Inputs and outputs normally cross the component boundary.

So why did I keep wanting them inside?

The Atom Never Gets the Outside Directly

Suppose the box represents observer $A$.

In this model, whatever actually exists beyond $A$'s boundary cannot be directly possessed by $A$ as the thing-in-itself. What $A$ can operate on is whatever becomes available through observation and representation.

Aside: This starts sounding a little Matrix-ish/Black Mirror-ish if I push it too far. My fingers, the keyboard I'm typing on, even my idea of "Self" and the boundary between Self and Other are all things "I" only know through some representation available to "me". That's a rabbit hole 🕳 for another post.

In my earlier post on data, I separated this path more finely: a system gives rise to a signal; an observer measures, samples, or otherwise interacts with that signal; data records the outcome of that observation; and information appears when that data is interpreted through a reference frame. I'm not throwing those distinctions away here. I'm compressing them because this post only needs the boundary-level abstraction.

Notation cleanup across posts In that earlier post I used $S$ for signal, $F$ for the observer's reference frame, $D$ for datum/data, and $O_F$ for the observation map. The framework evolved, and later posts gave $F$ and $D$ different jobs: $F$ became the processing/operator axis and $D$ the distinguishability domain. The concepts carry forward; the letters don't. Here $\pi_A$ is a deliberately compressed observer-specific map, not a claim that signal, data, and information are the same thing.

So I am going to keep the familiar labels $x_{\text{in}}$ and $x_{\text{out}}$ for now, but change what I mean by their placement.

$x_{\text{in}}$ is not the outside thing itself, nor necessarily the physical signal through which $A$ observes it. It is the representation that becomes available inside $A$ as input.

From the modeler's point of view, I can write:

$$ W_{\text{in}} \xrightarrow{\pi_A} x_{\text{in}} $$

$W_{\text{in}}$ is only a placeholder for whatever outside state ultimately gives rise to the observation. $\pi_A$ compresses the path through signal, observation or sampling, and representation that I separated in the earlier post. Depending on the abstraction level, $x_{\text{in}}$ may correspond to recorded data, interpreted information, or some other internal representation. The important point here is that $A$ operates on what becomes available inside $A$, not on the outside thing-in-itself.

Likewise, $x_{\text{out}}$ is the representation produced by $A$ as output. It is still inside the atom's represented world until some relation carries an effect across a boundary.

The labels in and out describe where a representation sits relative to the atom's boundary. They do not mean the unknowable outside itself sits inside the box.

That explains the diagram I kept wanting to draw:

$$ \boxed{ \begin{array}{ccccc} x_{\text{in}} & \longrightarrow & f_A & \longrightarrow & x_{\text{out}}\end{array} } $$

The weirdness was useful. It forced me to separate what is outside from what the atom represents as input or output.

Retained State Versus Input Representation

Now add retained state.

Call the atom's current retained state $s$ and the state produced by the transformation $s'$. State is still data. The distinction is relative to the atom:

$$ s=\text{retained internal state} $$
$$ x_{\text{in}}=\text{representation available inside }A\text{ as input} $$

Then the atom can be written:

$$ \boxed{ (s,x_{\text{in}}) \xrightarrow{f_A} (s',x_{\text{out}}) } $$

or visually:

$$ \boxed{ \begin{array}{ccccc} && s && \\ && \downarrow && \\[3pt] x_{\text{in}} & \longrightarrow & f_A & \longrightarrow & x_{\text{out}} \\[3pt] && \downarrow && \\ && s' && \end{array} } $$

The arrow gives the transformation a direction, but I have not yet introduced a sequence of observations such as $i,i+1,i+2$. That extra structure will come later, if the problem earns it.

An atom is not necessarily a function, process, thread, machine, service, register, or logic gate. It is whatever unit I'm treating as indivisible for the current question.

An atom is a unit I treat as indivisible for the role and unit type I currently care about.

But What Creates the Boundary?

This is where I ran into a deeper problem while writing this post.

I've been casually drawing a box around $A$, then saying there is $A$ and there is "outside $A$".

But the very phrase outside already assumes a distinction:

$$ A \mid \neg A $$

Self versus not-Self. Inside versus outside.

So the boundary can't be logically prior to distinction. The boundary is what appears once a distinction has been made.

This connects directly to $D$, the distinguishability domain I introduced in the distributed-systems post.

At minimum, $D$ lets me say:

$$ A\neq_D\neg A $$

So a better order is:

$$ \boxed{ \text{distinction} \rightarrow A\mid\neg A \rightarrow \text{boundary} } $$

rather than:

$$ \text{boundary} \rightarrow \text{distinction} $$

So What Is $D$?

I did define $D$ operationally in the previous post. What I am deliberately not pretending I've completely solved is the deeper question of what $D$ ultimately is, or what makes distinction itself possible.

For the framework, I called $D$ the distinguishability domain: the reference domain under which things can be told apart.

I wrote it as:

$$ D=\{d_1,d_2,\ldots\} $$

where the $d_i$'s are frame elements. At minimum, relative to some $d_i$, I need to be able to say that two things are distinct:

$$ a\neq_{d_i}b $$

Nothing in that statement says which one came first. Nothing says how far apart they are. Nothing says one is greater than the other.

Those are additional structures.

A richer $D$ may optionally supply an ordering relation:

$$ d_i\prec_D d_j $$

or distance, timing, magnitude, or some other relation. But none of those comes for free merely because two things are distinguishable.

Do not let the notation smuggle in structure that the system has not earned.

The stranger question begins when one frame element still contains several things I need to distinguish. In the previous post, I handled that by refining the same distinguishability domain:

$$ d_i\rightarrow\{(d_i,a_1),(d_i,a_2),\ldots\} $$

and, if necessary, again:

$$ (d_i,a_j)\rightarrow\{(d_i,a_j,b_1),(d_i,a_j,b_2),\ldots\} $$

So I don't need an endless alphabet of new structures behind $D$. The same idea of distinction can recur as the reference is refined.

But writing the atom model makes a deeper problem visible. If I draw $D$ "outside" atom $A$, and I take the boundary seriously, then $A$ cannot directly possess that outside $D$ as the thing itself. Anything $A$ says about it is already some representation available inside $A$.

$D$ is defined well enough for the engineering framework: it is the reference domain under which distinctions are made. What remains open is the foundational question of what ultimately grounds that domain, and whether an observer can completely observe the very distinction structure under which it observes.
Danger: rabbit hole At this point I started asking whether distinction is more primitive than the atom itself, whether a represented boundary is really a projection, and whether representations of distinction can themselves become objects of distinction. That's where this starts drifting toward self-reference, foundations of mathematics, and probably a Gödel-shaped headache. I'm stopping here on purpose.

When Order Is Earned

Now I can finally introduce the notation I was tempted to use earlier:

$$ x_i, \qquad x_{i+1} $$

Why not use it from the beginning?

Because the subscripts do more than distinguish two representations. They order them.

$$ i \lt i+1 $$

So $x_i$ and $x_{i+1}$ are useful once I care about successive observations, repeated invocation, or a chain of transitions.

For a repeatedly operating atom:

$$ \boxed{ (s_i,x_i) \xrightarrow{f_A} (s_{i+1},x_{i+1}) } $$

Now the ordering is intentional. I have earned it by introducing succession.

There are still two different ideas here:

$$ \underbrace{x_i\neq_D x_{i+1}}_{\text{distinction}} \qquad \underbrace{i \lt i+1}_{\text{succession}} $$

The first only needs $D$'s minimum job, assuming the two representations actually differ. The second comes from the ordered transition index.

If $D$ itself has also been enriched with an order relation, I may additionally be able to say:

$$ x_i\prec_D x_{i+1} $$

And if $D$ has a metric:

$$ d_D(x_i,x_{i+1})=\Delta $$

Again, those are not free.

This also explains why the indexed notation becomes natural when atoms are chained:

$$ x_0\xrightarrow{f_0}x_1\xrightarrow{f_1}x_2\xrightarrow{f_2}\cdots $$

At a relation between two atoms, A's output representation may give rise to B's input representation:

$$ x_{\text{out}}^{(A)} \xrightarrow{\text{relation}} x_{\text{in}}^{(B)} $$

while an indexed description emphasizes its place in a sequence.

$x_{\text{in}},x_{\text{out}}$ describe where a representation sits relative to the atom's boundary. $x_i,x_{i+1}$ describe where representations sit in an ordered sequence. I use the stronger notation only when the stronger structure matters.

Atoms Form Molecules

Now back to engineering.

Let atom $A$ interact with atom $B$ through intermediary $R_1$:

$$ \boxed{A} \longrightarrow \boxed{R_1} \longrightarrow \boxed{B} $$

$R_1$ could be a network, queue, shared memory region, wire, protocol adapter, human, database, conveyor belt, or something I haven't named yet. At this level I can treat $R_1$ as an atom too, even if opening it later reveals another molecule.

Now suppose $B$ can send something back. That reverse path may not be the same physical path. Call it $R_2$:

$$ \begin{array}{ccccc} \boxed{A} & \longrightarrow & \boxed{R_1} & \longrightarrow & \boxed{B} \\[7pt] \boxed{A} & \longleftarrow & \boxed{R_2} & \longleftarrow & \boxed{B} \end{array} $$

That exposes things a one-way diagram hides: acknowledgements, responses, errors, retries, flow control, backpressure, feedback, and negotiation.

It also exposes separate contracts:

$$ A\xrightarrow{C_{AR_1}}R_1\xrightarrow{C_{R_1B}}B $$

The contract between $A$ and $R_1$ need not be the contract between $R_1$ and $B$. In practice, a contract might be an API, a function signature, a wire protocol, a message schema, or some other agreement about what may cross the relation and how it is interpreted.

Molecule = Atoms + Relations.

And the Outside Observer?

Now put observer $O$ above the molecule:

$$ \begin{array}{c} \boxed{O} \\[5pt] \downarrow \\[7pt] \left[\boxed{A}\rightarrow\boxed{R_1}\rightarrow\boxed{B}\right] \\[-1pt] \left[\boxed{A}\leftarrow\boxed{R_2}\leftarrow\boxed{B}\right] \end{array} $$

Being drawn outside the molecule does not magically make $O$ omniscient.

$O$ also only gets whatever projection of the molecule becomes available to $O$:

$$ \mathcal{M}_i\xrightarrow{\pi_O}x_i^{(O)} $$

So the recursion starts again.

$$ \boxed{ \text{Atom} \leftrightarrow \text{Molecule} } $$

This is why I don't think system design has one privileged node size. Threads, processes, machines, clusters, regions, factories, and even organizations can all become atoms or molecules depending on what role I'm analyzing.

The modeler is an observer too. When I reason about Physics or Intention, I'm reasoning from what I can establish about them, not claiming perfect access to reality.

Now Ask What Everybody Wants

Once I know the actors and relations, I need to know what the system is supposed to do.

That's Intention.

Until I state the intention, I cannot even define failure.

Failure is always failure relative to an intention.

Good Intentions, Bad Intentions

There's another thing I care about: not every atom necessarily wants the same outcome.

I use good and bad here in a narrow systems sense. A good intention is aligned with the intended behavior of the system. A bad intention conflicts with it.

$$ I_A\parallel I_{\text{system}} \qquad\text{versus}\qquad I_Z\not\parallel I_{\text{system}} $$

This is where I keep thinking about Cus D'Amato and fighting: imagine an adversary, even yourself, fighting your system with "bad intentions". That was my inspiration for the divide between Good and Bad Intentions. The other thing in front of you is not merely another physical object. It may have intention.

Add adversarial atom $Z$:

$$ \begin{array}{ccccc} \boxed{A} & \longrightarrow & \boxed{R_1} & \longrightarrow & \boxed{B} \\[7pt] && \uparrow && \\[-1pt] && \boxed{Z} && \end{array} $$

Now $Z$ may try to pretend to be $A$ or $B$, insert itself into $R_1$, observe a relation it was never meant to observe, replay an old interaction, modify a representation, delay or suppress an interaction, infer hidden state through a side channel, or push data across a privacy or governance boundary.

Spoofing attacks identity and distinguishability. Replay intentionally creates duplication. A man-in-the-middle attack inserts an unintended observer or mediator into a relation.

A side channel is especially interesting because the intended architecture might say:

$$ A\not\rightarrow Z $$

while Physics quietly permits:

$$ A\xrightarrow{\text{timing/cache/power/etc.}}Z $$

Accidents Are Different

Not every failure comes from bad intention.

If a disk dies, an operator mistypes a command, a packet disappears, a race occurs, a bit flips, or a process runs out of memory, there may be no adversary at all.

$$ \boxed{ \text{Failure Cause} \rightarrow \begin{cases} \text{non-adversarial} \\ \text{adversarial} \end{cases} } $$

That matters because reliability and security are related but not identical problems.

PIF: Physics, Intention, Failure

$$ \boxed{ \text{Physics} + \text{Intention} + \text{Failure} \Longrightarrow \text{required mechanisms} } $$

Physics is what reality gets a vote on: finite bandwidth, finite storage, finite compute, propagation delay, clock differences, rate mismatch, contention, power loss, noise, distance, hardware failure, and unintended channels.

Intention says what an actor is trying to make true, or what behavior the system is required to preserve. Those intentions may come from architects, engineers, users, business stakeholders, regulators, or other actors, and may include technical, economic, organizational, legal, or policy requirements. That includes both aligned and adversarial intentions, which is why I distinguish between good and "bad" intentions.

Failure asks how reality can violate those intentions.

The 4 Ds

$$ \boxed{ \text{Drop} \qquad \text{Duplicate} \qquad \text{Deform} \qquad \text{Delay} } $$

The 4 Ds classify the observable deviation. They don't tell me the cause.

A duplicate may come from an innocent retry or a replay attack. A deformation may come from corruption, schema mismatch, a bug, tampering, or spoofing. A delay may come from load, queueing, a slow dependency, or denial of service.

Note: I originally had 5 Ds, where the fifth D was Disagree, as in a consensus problem. I dropped it because disagreement isn't primitive in the same way the other four are. It can be derived from them. If one consumer misses an update, receives it late, applies it twice, or receives a deformed version of it, its state can disagree with the others. The disagreement is the inter-atom symptom; the underlying deviation is still one or more of the 4 Ds.

Mechanisms Are Earned

If something can Drop and my intention requires eventual delivery, I may earn acknowledgements and retries.

But retries can create Duplicates, so I may then earn identity, deduplication, or idempotency.

If information can Deform, I may earn validation, schemas, checksums, reconciliation, signatures, or authentication.

If Delay matters, I may earn buffering, timeouts, backpressure, admission control, ordering, or replication.

$$ \boxed{ \text{Intention} \rightarrow \text{Physical Constraint} \rightarrow \text{Failure} \rightarrow \text{Mechanism} \rightarrow \text{Technology} } $$

Technology comes last.

A mechanism can be necessary without deserving its own box. Whether I draw that mechanism as a separate box depends on what I need to reason about.

Let's Earn Kafka

Suppose atom $A$ produces work that atom $B$ consumes.

$$ \lambda_A=\text{production rate} \qquad \mu_B=\text{consumption rate} $$

If:

$$ \lambda_A\leq\mu_B $$

maybe direct handoff is enough.

But if:

$$ \lambda_A>\mu_B $$

Physics has created a rate mismatch.

Something now has to give. I can slow the producer through backpressure, reject or drop some work, increase consumption capacity, or allow the excess to wait.

Suppose my intention says two things: accepted work must not disappear, and the producer must be allowed to continue producing during a temporary mismatch.

Then the excess has to exist somewhere while $B$ catches up.

The system has now earned a buffer.

Not Kafka.

A buffer.

Now say pending work must survive failure of the component holding it. The system earns durable storage.

Suppose $B$ must restart and reread old work. It earns retention.

Suppose multiple consumers need independent positions. It earns offsets or cursors.

Suppose one sequential stream cannot satisfy the required throughput, and the workload contains distinguishable groups whose ordering does not need to be global. Now I can parallelize across multiple ordered streams. The system has earned partitioning.

Suppose events for the same customer must stay ordered:

$$ \operatorname{partition}(x)=h(\operatorname{customerID}(x)) $$

Notice what quietly entered here: $D$. customerID gives me a way to distinguish one customer's events from another's. The hash then maps that distinguishable key to a partition, and the partition becomes the domain within which I require ordering. Without some such distinction, "keep events for the same customer ordered" is not even operationally meaningful.

Suppose the log must survive machine failure. Now we earn replication and some rule for deciding which copy is authoritative enough for the intention.

Huh. This is starting to look Kafka-shaped.

Now Kafka gets a box.

The Extremum Principle

Deriving a plausible architecture isn't enough. Now I try to break it.

The Extremum Principle: push a parameter toward a boundary and ask what breaks.
$$ 0,\qquad 1,\qquad N,\qquad \min,\qquad \max,\qquad \infty $$

What happens at zero traffic?

What happens with one node?

What happens as $N\rightarrow\infty$?

What happens as $\lambda\rightarrow\mu$, and then $\lambda>\mu$?

What if one customer generates nearly all the traffic?

What if every client retries at once?

What if network delay approaches the timeout?

What if the data no longer fits on one machine?

What if the "rare" failure lasts for hours?

$$ \boxed{ \text{Extremum} \rightarrow \text{exposed assumption} \rightarrow \text{constraint or failure} \rightarrow \text{new mechanism} } $$

Extremum Forces the Atom Open

Suppose I have treated a database as one atom:

$$\boxed{DB}$$

Then I push storage until:

$$ \text{required storage} > \text{capacity of one }DB $$

The abstraction stops satisfying the intention. So I open the atom:

$$ \boxed{DB} \longrightarrow \{ \boxed{DB_1}, \boxed{DB_2}, \ldots, \boxed{DB_n} \} $$

What looked like one atom is now a molecule, and immediately I inherit new relations and new failure modes.

$$ \boxed{ \text{Atom} \rightarrow \text{stress} \rightarrow \text{decompose} \rightarrow \text{Molecule} \rightarrow \text{new relations} \rightarrow \text{new failures} } $$

Then from some higher level I can close the box around the whole thing and treat the molecule as one atom again.

My Architecture Lie Detector

"We'll just retry."

What if the destination is down for six hours? What if everybody retries at once?

"We'll put everything on one partition so ordering is easy."

What happens when throughput keeps growing?

"We'll keep it all in memory."

What happens when the dataset no longer fits?

"We'll add more nodes."

What happens when coordination costs more than the useful work?

"Timestamps tell us the order."

According to whose clock?

"That endpoint is internal."

Internal relative to which boundary, and what physically prevents an unintended observer from reaching it?

Extremum does not automatically hand me the solution. It tells me where my current story stops making sense.

Where SPPARRS Fits

$$ \boxed{ \text{Security} + \text{Power} + \text{Performance} + \text{Area} + \text{Reliability} + \text{Relativity} + \text{Scale} } $$

I don't want SPPARRS to generate an architecture by checklist.

Physics-Intention-Failure derives why mechanisms need to exist. Extremum attacks assumptions. SPPARRS audits the resulting system for major constraint classes I may have ignored.

The Framework, Compressed

$$ \boxed{ \text{Distinguish the Atom} \rightarrow \text{Build the Molecule} \rightarrow \text{State the Intention} \rightarrow \text{Respect Physics} \rightarrow \text{Enumerate Failure} } $$
$$ \boxed{ \text{4 Ds} \rightarrow \text{Extremum} \rightarrow \text{Required Mechanisms} \rightarrow \text{Technology} } $$
$$ \boxed{ \text{SPPARRS} \rightarrow \text{attack the design again} } $$

Look for the Bridge

A producer outrunning a consumer is not fundamentally a Kafka problem.

It's a rate-mismatch problem.

A fast clock domain feeding a slower one can encounter a structurally similar constraint. A factory line can encounter it. A network service can encounter it.

Different physics. Different unit type. Similar relation.

If I can identify the relation that failed before naming the product, I can go looking for mechanisms in fields I would not normally search.

Look for the bridge.

So Why Does This Box Exist?

The next time I draw an architecture, I want to be able to point at every box and explain why it exists.

Not:

"Because we need scalability."

"Because Netflix uses it."

"Because this is what the system-design book showed."

I want:

This is the atom and this is the relation I care about.

This is the intention.

Physics permits this failure.

The failure violates that intention.

Therefore I need this mechanism.

This box implements that mechanism.

That's an architecture I can defend.

And if I can't explain why the box exists?

Delete the box.

Explore this post

Highlight any sentence or formula to Explain, Derive, Attack, Example, or apply the Extremum Principle. Or interrogate the whole post below. This builds a context-preserving prompt for the LLM you already use.


© 2026 Clarence Bowen. Derived, not assumed.