Receiver-Conditioned Latent Communication gives 94% CacheBack
Maximillian Rossi
Prajwal Raghunath
Haoqing Xuan
Yusen Zhang
Eugene Wu
Columbia University
TLDR; Agents can share internal model state instead of generating text messages. Sending all of it, however, overwhelms the receiver.
The receiver asks for what it needs. CacheBack uses attention to that request to select which state gets sent. Simple, robust, and training-free.
At 16× compression, CacheBack passes on just 6.25% of sender positions.
Selected source positions can be sent as token IDs; latent steps must be sent as continuous vectors. The receiver prefills both to build its own state.
CacheBack on a coding task
36 seconds · 4K · Recorded coding task
Fixing a Django bug with seven coding agents
Seven Qwen3-8B workers inspect the code and send information to a coordinator, which writes the Django fix. The video compares selected latent state with generated text messages. CacheBack produces the patch 4.41× faster than with text communication on this case. Both runs make the same one-line code change and pass all 88 tests.
One recorded case: CacheBack 25.66 s; text 113.21 s. Separate recorded runs are aligned at their start. Startup and test grading excluded. Silent video; playback speed varies. Open in full window
Results
FanOutQA
Questions require combining evidence across Wikipedia pages. We split at least 120K tokens among three parallel senders; a receiver answers from their messages. Strict accuracy requires every reference-answer group. Higher is more accurate; left is faster. Each panel fixes the receiver model. Text labels give sender size; CacheBack uses same-size senders. 4× compression retains ¼ of sender positions; 16× retains ¹⁄₁₆.
Wikipedia pages ASender 1
Wikipedia pages BSender 2
Wikipedia pages CSender 3
messages
Combines 3 messagesReceiver↓ Answer
● ━ CacheBack · labels show compression factor▲ ┄ Text · labels show sender size× Question only
Qwen: 4× compression improves accuracy and completion time over 8B text; 1.7B text is slightly faster but less accurate.Nemotron: 8× compression improves on 12B text; 4B text is faster but less accurate.
LongBench v2 Easy
Long-document questions test whether evidence survives repeated handoffs. We split each 100K-246K-token document into four equal parts. Each sender reads its part plus the previous message; a final receiver answers from the fourth message. Compression factors retain a fraction of accumulated context (4× keeps ¼); fixed budgets (32K, 64K) cap message positions. Axes and text-size labels follow FanOutQA.
Document part 1Agent 1
Document part 2Agent 2
Document part 3Agent 3
Document part 4Agent 4
Final messageReceiver↓ Answer
● ━ CacheBack · labels show compression factor■ ━ CacheBack · labels show fixed position budget▲ ┄ Text · labels show sender size× Question only
Qwen: 4× compression gives the highest plotted accuracy; stronger compression loses accuracy.Nemotron: relative 8× compression beats 12B text on both axes; 4B text remains faster but less accurate.
Lines join the best accuracy-time tradeoffs within each channel; faded points are dominated settings. × means the receiver gets only the question. 50 concurrent tasks on 8 H100s; accuracy averaged over three receiver draws. Completion includes queueing. Full setup and results in the paper.
In multi-agent question answering, a large document collection can be split across agents, each with its own context. Answering a question requires combining information across those contexts.
With text communication, agents write messages summarising what they read. CacheBack lets them pass latent thoughts together with a subset of internal state selected for what the next agent needs. This example shows selected-state handoffs: three agents read separate documents in sequence, and a fourth agent produces the final answer.
The top row shows each agent’s input document: a booking confirmation, a location-change notice, then an ID policy. Agents 2 and 3 also receive the previous agent’s handoff; Agent 4 answers from the final handoff alone. The bottom row shows the messages generated by the text agents.
Question
Where should Maya collect her pass on Friday, and what ID should she bring?
Press Run to watch the handoffs.
Use CacheBack
Install rclc, bind your agents, and pass the receiver’s request.