Inference & PerformanceModel Architecture

Accelerating LLMs on a Mac with an iPhone Looks Crazy, but It Isn't That Crazy

An iPhone tethered next to a MacBook with a 10 Gb/s USB-C cable to accelerate local LLMs looks at first glance like a flashy geek toy. With a phone’s modest compute power combined with the communication latency over a physical cable, it could easily feel like using a sledgehammer to crack a nut.

That was my initial thought when I first looked at this open-source project Backburner. After examining the code, however, I realized many people had misunderstood it. The author is not trying to run the LLM on the phone. Computation remains on the Mac throughout; inference is never executed entirely on the phone.

In this system, the MacBook Pro acts as the host providing external services, running a regular local server. For instance, on a Mac with 24 GB of memory, ingesting long contexts puts memory under strain, so the cabled iPhone serves as an external memory pool and coprocessor attached to the other end of the wire. It takes on two responsibilities: first, sharing a portion of the network layers during the prefill phase (when reading long contexts); second, taking over the overflowing historical KV cache when the Mac cannot fit ultra-long contexts. You may already be familiar with these two concepts: prefill ingests the entire input in one shot, while decode outputs tokens one by one afterward. The phone plays completely different roles in these two phases, as will become clear below.

Even if you do not have an iPhone or a computer with lower memory at hand, this remains an intriguing scenario. It depicts a predicament many of us may encounter: physical resources are constrained, yet computational workloads must proceed. Long contexts saturate the host’s unified memory, the external device provides asymmetric compute, and the two are joined only by a single wire. Under such constraints, where do we go next? Looking beyond specific hardware models, this philosophy of navigating trade-offs within physical limits applies equally to multi-GPU parallelism, distributed inference, and edge-device collaboration: when available resources are scarce but tasks must be finished, where is the breakthrough?

First, Pick Workloads That Can Be Mathematically Split Without Loss

Human instinct naturally seeks to optimize the cable first—benchmarking whether communication bandwidth is adequate, squeezing protocol latency, and aggressively compressing transmitted data. But this easily breeds the misconception that as long as data travels fast enough, compute across both sides will seamlessly link up.

Anyone with background in distributed systems knows the biggest pitfall here. Whether two machines can collaborate on a single task is dictated not by how fast the link runs, but by whether the computation delegated to them can be solved independently in mathematics and merged without loss in the end. Link bandwidth is merely a secondary consideration.

If an operator, once partitioned, still requires extensive cross-device data dependencies, both ends must constantly synchronize intermediate states. Under such conditions, even with minimal cable latency, round-trip waiting times can easily wipe out the benefits of offloading compute across multiple devices. Such operators are ill-suited for this partitioning approach.

In this system, the computation chosen for offloading is attention. This is because attention possesses a mathematical property: it naturally lends itself to locally independent computation.

When the model’s full attention layers encounter long contexts, they must take the dot product of the current query vector with all past keys in history, and then compute a weighted sum over the values using attention weights. As context grows longer, the historical KV cache expands, eventually exceeding the memory capacity of a single machine. At this point, the system has the Mac offload the earlier portion of the KV cache en bloc to the phone for hosting, while feed-forward layers and model weights remain locally on the Mac.

During attention computation, the Mac transmits the current query vector over the data cable to the phone. The phone then performs local attention computation solely on its stored historical KV pages, obtaining a local result that includes a local output vector, a local maximum, and a sum of exponentials. Subsequently, the phone returns the normalized local output vector along with the log-normalizer computed from the maximum and exponential sum back to the Mac. Using the LogSumExp formula, the Mac merges the phone’s local result with its own local result.

On the left is an unsplittable operator, where global dependencies cause frequent synchronization between both ends and benefits are consumed by cable latency; on the right is a splittable operator, partitioned by position to evaluate independently and then merged precisely

Consequently, this partitioning yields the exact same outcome as before. Mathematically speaking, merging local attentions is completely equivalent; this split introduces no lossy quantization or approximate truncation. Under greedy decode testing in default mode, every generated token matches the single-machine, token-by-token output bit for bit. To make two devices cooperate, identifying computations that can be mathematically decoupled first is far more practical than searching everywhere for faster cables.

The Seams of State Organization Determine the Partition Points

Once the attention operator is confirmed to be splittable, the next natural intuition is to divide compute according to each device’s capabilities—allocating more layers to the stronger host and fewer layers to the weaker secondary device; as long as timings match on both ends, throughput remains stable.

Yet when actually attempting to partition the model, one often discovers that this approach simply does not work. This is because different layers in many models store state in fundamentally different ways.

Take Qwen3.8-27B used in this project as an example. Most of its layers do not preserve every historical token; instead, they maintain a fixed-size summary. As new tokens arrive, the summary is updated in a rolling fashion, folding prior information into itself. This saves memory, but the summary is monolithic—there is no way to carve out the first half of history for another machine, because it is no longer divisible. The summary also rolls forward, so if you ever need to revisit the past, you must recompute from scratch.

A minority of other layers use a different method, storing an individual row for every historical token and accumulating a table that grows with the context. Each row in this table is mutually independent, meaning the earliest rows can be relocated en bloc. These two types of layers alternate throughout the entire network: across every sequence of 4 layers, the first 3 are the summary layers described above, and only the 4th is a table-storing layer—repeating in a 4-layer cycle all the way to the end of the model.

This directly dictates what the phone can do. The first layer type features only a monolithic summary; to leverage it on the phone, one would have to migrate the entire layer, as it cannot be sliced by position. The second layer type features a table, allowing the oldest rows to be transferred to phone memory and retained, with the position sliced at will. Thus, the system’s full-layer partition points must align with the 4-layer cycle boundaries, following the data’s inherent seams. Where one can cut is constrained by the arrangement of these two layer types, though the specific boundary chosen remains up to the engineer.

Once partitioned, in split prefill under the default configuration of up to 64k context, the Mac executes the first 40 layers and hands the intermediate activations off to the phone to compute the remaining 24 layers; the final ubatch of each batch remains entirely on the Mac. At this point, one might wonder: if the Mac finishes 40 layers before handing off to the phone, isn’t that still sequential execution—so where does the speedup come from? The secret lies in pipelining. Long inputs are already split into small chunks fed in batches. Once the Mac finishes the first 40 layers of a chunk, instead of waiting for the phone, it transmits the intermediate results and immediately turns around to compute the first 40 layers of the next chunk. While the phone is still chewing through the tail of the first chunk, the Mac has already entered the second. Both devices stay busy concurrently within the same time window; sequential dependencies are partially masked by the pipeline, substantially reducing idle wait time.

Compiling Immutable Data into Hardware Compute

With state partitioning in place, the system was immediately bottlenecked by another common edge inference hurdle: the decode phase. In full attention layers, for every token generated, the chip must read the preceding KV cache; the longer it gets, the costlier it becomes, until memory bandwidth becomes the bottleneck.

The conventional approach to this bottleneck is optimizing cache layout or making minor adjustments to memory transfers. But looking closely at the autoregressive inference process reveals an easily overlooked reality: historical keys and values generated dozens of steps earlier are read repeatedly throughout decode, but never modified. Since this data is already fixed, continuing to read it from memory as dynamic data is highly inefficient.

Based on this insight, the development team adopted an alternative scheme. Inside the phone sits the Apple Neural Engine (ANE), capable of running models that perform matrix multiplications with fixed weights. The system takes the historical KV pages held by the secondary device and compiles them page by page directly into ANE model files; these pages are compiled sequentially in the background during the prefill phase, as detailed in the Neural Engine implementation documentation. Within the compiled model, the keys and values previously residing in memory are baked directly into static weights.

When the host transmits a new query vector during decode, the secondary device feeds the query vector as input into this ANE model to compute the local attention result for the compiled pages. While the query vector updates at every step, the historical keys and values in the compiled pages remain fixed as model weights, rather than eliminating memory access entirely.

An end-to-end pipeline from freezing historical pages, compiling into hardware weights, inputting queries, and hardware matrix multiplication to identical results; contrasted with the conventional path of repeatedly reading the memory bus and being constrained by bandwidth

This mechanism of offloading historical cache to the phone and having the phone share attention over old context is explained in greater detail in the project’s design notes. In benchmarks on the A18’s old-page attention operator, compiling static data into hardware weights proved 2-3x faster than the phone’s GPU. The page weights use fp16 rather than insufficiently accurate int8 page weights; across a 51k real-session test, all 33 tokens matched the exact path bit for bit. When data does not change, treating it as hardware weights accelerates old-page attention computation.

Even after establishing the data representation and operator partitioning, the two devices still need to communicate over a single link. The host device computes quickly while the secondary device is relatively slow, and any unnecessary latency across this link will drag down the throughput of the entire pipeline.

Faced with such situations, many distributed systems add layer upon layer to the communication stack—piling on scheduling protocols, multiplexing, and retry mechanisms, using increasingly intricate plumbing to cope with narrow interconnects. This system takes the opposite path: if the wire is narrow, subtract from the architecture so that the vast majority of data never needs to cross the wire at all.

It applied three subtractions across the two ends of the cable.

First, mirroring the KV cache and recurrent state for its own network layers directly on the phone, so that only newly added rows are synchronized across the wire.

Second, transmitting only residuals between layers. Instead of passing intra-layer intermediate tensors, it transmits inter-layer residual vectors.

Third, keeping destination computation local in split prefill mode. The final ubatch of each input batch is enforced to run locally on the Mac. Consequently, the vocabulary-sized logits tensor remains on the host, never crossing the wire.

Timeline of host preceding layers and phone succeeding layers: most work remains in local loops on each side, with only minimal data such as new increments and inter-layer residuals traversing the cable, while compute blocks on both ends overlap in time to hide transmission

After drastically cutting the volume of cross-wire data, temporal pipelining was introduced: beyond 64k, each 512-token ubatch is interleaved into two halves—the Mac handles the full network layer computation for one half of the tokens, while the phone simultaneously computes attention against old keys for the other half. The two interleave in time. In 140k context testing, prefill throughput increased from 58 to 68 tok/s, an improvement of roughly 17%; about a 30% boost came from split prefill tests at 32k and 48k. On narrow links, spending time building elaborate plumbing yields diminishing returns; relying on local loops to eliminate the vast majority of cross-wire interactions proves far more effective.

Discarding Negative Optimizations and Seeing the True Boundary of Gains

Evaluating system design requires looking not just at what features were implemented, but also at what was discovered and abandoned during experimentation. On paper, it is easy to envision seemingly plausible collaboration models: since the Mac has its own ANE, why not leverage it to participate in decode as well? The phone has both a GPU and an ANE, so why not use both? Yet on real hardware, such ideas often prove flawed. The project documents several approaches that were disabled or shelved after empirical measurement:

Early on, they attempted to use the Mac’s own ANE to assist with decode. But benchmarks revealed that generation speed actually slowed by 26%, as the ANE competed with the GPU for bus bandwidth during frequent memory reads and writes, dragging down the primary compute path. This idea was scrapped.

On the phone, they also experimented with running prefill concurrently across the GPU and ANE. When both were heavily loaded, contention ensued: ANE throughput dropped sharply, and the combined throughput was only 1.1-1.25x that of either engine running alone. Consequently, this scheme was shelved.

They also tried dynamically learning the partition ratio between the ANE and GPU at runtime. However, the ANE downclocks when idle, causing probe request latencies to fluctuate wildly and disrupting the online scheduling logic. In the end, they reverted to a fixed static constant.

Beyond weeding out negative optimizations, the system is remarkably candid about its limitations. In default mode, for short-to-medium contexts under 64k, it relies solely on the Mac for decode, meaning the secondary device does not participate in decode and thus offers no acceleration. For lightweight ingest requests of no more than roughly 512 tokens, split prefill is bypassed entirely in favor of standalone host execution to avoid uneconomical communication overhead. The system also imposes operational constraints: on mobile operating systems, where background GPU access is prohibited for apps, the phone must stay in the foreground with its screen on if it participates in compute. Furthermore, the entire system currently supports only single-request serial execution.

What this entire design truly solves is going from non-runnable to running reliably, rather than doubling routine inference speeds. An 8 GB host cannot fit the 27B model used in the project, and a 24 GB host under its stock configuration cannot fit long-context cache exceeding its memory capacity; tethering an external device via a single cable pushes past that barrier. Keeping unsplittable workloads out, pipelining along the seams of data structures, turning frozen data into forward compute, and subtracting ruthlessly over a narrow link—this decision path forged under extreme constraints remains valuable whenever resource bottlenecks appear, even if you take away the cable and the specific devices.

Evaluation Dimension Scenario & Configuration Baseline Measured Performance or Measurement Baseline
Long-context prefill throughput improvement 16k / 32k / 48k context Single-machine 109 / 101 / 87 tok/s increased to collaborative 157 / 130 / 113 tok/s
Cold-start large-context first token response 27k token session first response Original 245 seconds, standalone optimized 228 seconds, reduced to 168 seconds with collaborative acceleration
Dedicated hardware operator evaluation latency Per 16k KV page (4 heads) Latency of 1.85 ms (equivalent to 0.113 ms per 1k KV), reaching 2-3x that of general-purpose cores
Ultra-long context per-step latency 140k extreme depth token-by-token generation Pure general-purpose compute at 279 ms/token, reduced to 208 ms with 2 baked pages per layer, and 176 ms with 3 pages
Physical cable round-trip latency baseline USB-C 16-byte request/reply round-trip via phone app Median 96 µs, 90th percentile 183 µs
Zero-to-one under extreme constraints 8 GB MacBook Neo + iPhone Air, running 27B model with split decode Enables running an otherwise unfittable model, achieving collaborative decode at 3.9 tok/s and prefill at 64 tok/s

Now that large language model architectures are increasingly embracing hybrid attention, the balance between splittable and unsplittable state across the system is shifting as well. Future models will have more fixed-length recurrent states and fewer attention layers, but each attention layer will carry heavier computation. What new boundaries will this bring to the partitioning between host and secondary compute nodes?