AI Research••KR

Paper of the Week #4 — A Memory of Procedures, or a Memory of Examples?

Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.

Paper of the Week #4 — A Memory of Procedures, or a Memory of Examples?

Paper of the Week #4 — A Memory of Procedures, or a Memory of Examples?

An agent that writes down what it learned, in words, and reads it back later is the most appealing shape in the field right now: no fine-tuning, no weights, just a growing file of skills. This week's paper reports that shape taking execution success from 72.7% to 99.3%. I built the narrow version of it on a task where the answer is known, and the rules moved nothing, while handing the model five raw examples fixed 19 of the baseline's 27 misses.

Sep 17 – Sep 23, 2026 · Designer-RSI · SWE-bench leaderboard audit · TypeSafe's Jev evals

The Claim

"External memory" hides two different things. One stores procedures — natural-language rules about how to do the task. The other stores examples — past cases with their answers. Papers report the first and often ship both, because the procedures are distilled from examples that stay in the loop.

My claim: on a task with fixed labels and a known answer, the procedures are not where the gain lives. Write the rules down as carefully as you like; what changes the score is whether a relevant labelled example reaches the prompt. If that is right, a system credited to "evolving procedural memory" should lose little when you keep the memory and drop the examples, and lose most when you do the reverse.

The Receipt

Designer-RSI (Hongyang Du, Lan Yan, Christian Flores, Asim Kadav, submitted 18 September 2026) keeps a frozen frontier model and grows an external bank of natural-language skills from real usage. The bank "widens by acquiring procedures for recurring uncovered subtasks" and "deepens by revising existing procedures against their own successful and failed executions." Across 1,406 real user briefs and 1,869 automatically graded trajectories it grows from 76 to 139 procedures, with "no weight updates and no human labels," and reports execution success rising from 72.7% to 99.3%, win rates of 61.8% and 67.6% against a no-skill agent, and 58.5% (p = 0.025) on held-out briefs against 49.4% and 48.6% for widening or deepening alone.

The cell the paper does not fill is the one my claim needs: procedures against examples, same task, same model, same judge. Design quality has no answer key, so they grade with a model. I ran the comparison where there is one.

The task is BANKING77: 77 banking intents, human labels, 154 test messages (two per intent). The model is GPT-5.6 Terra at reasoning effort low, output constrained to one of the 77 labels. The baseline prompt is the label list and nothing else: 127 of 154 correct.

Then I built the narrow Designer-RSI. I ran the baseline over 600 training-set messages, collected the 72 it got wrong, grouped them into 52 confusion pairs, and asked the model to write one operating rule for each of the 40 most frequent pairs — "choose card_delivery_estimate when the customer asks about an expected delivery timeframe; choose card_arrival only when they report or ask whether the card has already arrived." No example sentences in the rules, only procedures, exactly the distinction the paper draws. That memory is 7,138 characters and rides in every prompt.

It fixed 4 items and broke 5. Against the same baseline, on the same 154 messages, an exact McNemar test gives p = 1.0.

Paired comparison against the same baseline run. Retrieval fixed 19 items and broke none; the rule memory fixed 4 and broke 5.

The contrast is the point. Swap the rules for five examples retrieved from the same training set by embedding similarity — no rules, no distillation, just the nearest labelled cases — and the same model fixes 19 items and breaks zero, p≈3.8×10⁻⁶. One example per label, statically, lands in between: 6 fixed, 2 broken, p=0.29. Shuffling the label list or grouping it by topic does nothing either way, which is the noise floor this sits on.

Where I am weak: this is one task, one model, one memory built one way. My rules were distilled in a single pass from 600 messages, while the paper iterates over five rounds and revises procedures against their own outcomes — I did not implement the revision loop, and a rule bank that never gets corrected is the weakest version of the idea. Classification is also the case most favourable to retrieval: the answer is a label that a nearby example carries outright. In a design workflow there is no nearest neighbour that hands you the answer, which is precisely why the paper's domain needs procedures.

The Contrast

Designer-RSI sets the claim up. Its strongest evidence for procedures specifically is the ablation: widening alone 49.4%, deepening alone 48.6%, both 58.5% (p = 0.025) against the no-skill agent on held-out briefs. That is an argument that the loop matters, not that natural-language rules beat examples — no examples-only arm appears in the comparison, and the grading is automatic throughout.

The SWE-bench leaderboard audit (arXiv 2609.17394) pushes from another angle. It reports within-model scaffold ranges reaching 29.8 points against an 8.8-point spread across the top thirty entries, and then does the thing most scaffold papers skip: exact paired McNemar tests, item by item. Its conclusion is procedural in the other sense — report model-scaffold provenance, stop reading small aggregate gaps as rank.

TypeSafe's Jev evaluation is the one I cannot use as evidence, and it is worth saying why. On its aggregate view across four workflows it reports 67.8% for Jev and 67.9% for GPT-5.6 Terra (workflow), but the reference answers are the averaged responses of GPT-6 Astra and Claude Fable 5.1 at high thinking effort. A number scored against two frontier models tells you about agreement, not correctness. Designer-RSI's automatic grading sits closer to this than to BANKING77's human labels.

Where I'd Be Wrong

Two conditions, both checkable. First: if I add the revision loop — re-scoring each rule against fresh training errors for three rounds, dropping the ones that never fire — and the rule memory then fixes more than it breaks at p < 0.05 on the same 154 messages, the single-pass distillation was the problem, not procedures. Second: if I run the same procedures-versus-examples comparison on a task where no retrieved case carries the answer, such as multi-step tool use, and the rules win there, then my claim is scoped to lookup-style tasks rather than true. I will run both and score them in issue #6.

Ship It

If you are about to build a skills file for an agent, measure the examples arm first. It took me an afternoon and a few dollars of API calls: run the baseline over your training set, keep the errors, and compare two prompts — one carrying distilled rules, one carrying the nearest labelled examples — on the same held-out items, scored with a paired test rather than two accuracy numbers. On my task the retrieval arm was the whole gain and the rules were free to drop. On yours it may be the opposite, and that is worth knowing before the skill bank becomes infrastructure you maintain.

The Ledger

Issue #3 promised three things for this issue: two were not kept, and one was conditional and went unchecked. DeepSeek-V2-Lite as the second non-renormalized control, which would test whether k₂ helps a model that also ships norm_topk_prob: false, did not run. Neither did forcing identical token sequences through every condition, the fix issue #3 named for the text divergence at batch 8 and 32. Scale-QLoRA's naive merge was promised only if the body and code surfaced; I did not check whether they had, so that one is open rather than failed. HoH at ten iterations stays on the ledger from issue #2, unchanged. The series also skipped two weeks, and those facts belong together: the GPU went to other measurements and the queue was not two weeks ahead, which is the exact trap this series warned itself about. This issue was written for the week of 23 September and goes out a week late. Reproduction streak: reset; this issue's Receipt starts it again at 1. The DeepSeek control and the token-forcing fix move to issue #5's ledger, DeepSeek first. Issue #3's other condition, a fused k₂ kernel reaching 1.5× at batch 1, stays open: no such kernel exists yet. Next issue: that control in the ledger, and a main pick with its own Receipt run before the issue is written.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts