← Back to notebook

Maybe the agent shouldn’t own the memory.

A lunch conversation, a graph-memory framework, a map of Basel and a strange detour through DeltaDB left me with a question: what if “agent memory” is really the beginning of a shared adaptive system layer?

The thesis in one view

This piece starts with a question, then sketches a model, then tests what it might imply.

Thesis diagram

Proposed model

Memory as system middleware

iThe agent receives a projection, not the source of truth.

system substrate → projection → actor → action → system update

Don’t let the metaphor lead your architecture.

We call it an agent. So we give it memory, skills, identity and tools.

But does the agent own those things because the system needs it to — or because the human metaphor made that architecture feel natural?

That question stayed with me after a lunch conversation about agent memory.

Today at EMPOWER in Zurich I met José Luis Latorre.

José was kind enough to endure a series of badly worded questions from me over lunch, and I am genuinely thankful for the time he took. I understood only part of what he was describing in the moment. The dangerous thing was that I understood enough for it to keep bothering me afterwards.

So I spent the evening looking through AgentMemory for .NET, AgentEval, and Neo4j’s NAMS / Agent Memory work around graph-native memory and skill distillation.

What I find so good about José’s work

The easy version of “agent memory” is retrieval: something happened before, we stored it, later we retrieve something vaguely relevant and put it back into the context window.

José’s AgentMemory for .NET structures experience into persistent short-, long- and reasoning memory, with provenance, temporal state, supersession and an explicit projection layer.

His AgentEval work provides an evidence-oriented evaluation and runtime-governance toolkit for agent systems.

Separately, Neo4j’s NAMS / Agent Memory work shows how graph-native memory can support the broader direction toward provenance-grounded skills and capability management.

Taken together, these architectures suggest something larger.

And then the really interesting thing happens: memory stops being only about remembering.

experience
    ↓
structured memory
    ↓
evaluation
    ↓
patterns
    ↓
capability
    ↓
changed behaviour
    ↺

A simple SKILL.md file is static prose. It tells a runtime how to behave, but it has no inherent relationship to the executions that produced it, no built-in notion of evidence, drift, failure or repair.

What I find especially interesting in the NAMS Agent Skills work is the shift in what a skill can be. A SKILL.md no longer has to be the deepest representation of capability. It can be a materialized runtime view of a capability whose grounding, review state, drift and repair live deeper in the system.

SKILL.md becomes a snapshot of a living capability.

Architecture evolution · 12 states Apparent agent ownership The agent appears to contain memory, skills and tools.
System / governed state
Canonical substrateone governed capability · evidence · provenance · policy · history
SKILL.md
wikirunbookchecklisttrainingAI skillworkflow
materialized from capability
deeper state→materialized checkout→existing tools
Runtime world
Runtime projection relevant capability
statehistoryskillstoolspermissionsinstructions
↓ scoped projection
Agent Actor A
memoryskillstools
Structured memory
shortlongreasoningprovenancetemporal statesupersession
experience→evaluation→capability
interpret · reason · decide · synthesize
Actor Bscoped context
Actor Cscoped context
Write-back / learning
↓ evidence, not mutation
Experienceoutcomes · failures · discoveries
Proposalcandidate change
Eval / governancepromote · reject · request evidence
↺ governed promotion to canonical substrate

One governed capability → many runtime / human representations

Explore diagram manually
Text version of the architecture
  1. Apparent agent ownership. A large agent appears to own memory, skills and tools.
  2. Structured memory. Memory separates into short-, long- and reasoning memory with provenance, temporal state and supersession.
  3. Capability emerges. Experience flows through evaluation into capability.
  4. Architectural inversion. Canonical substrate projects a runtime view into an actor that is only one layer.
  5. Materialized skill. SKILL.md is a capability projection, not agent property.
  6. Delta analogy. Deeper state becomes a materialized checkout used by existing tools.
  7. Context virtualization. Projection carries state, history, skills, tools, permissions and instructions.
  8. Experience return path. Actor outcomes return as evidence on a separate path.
  9. Proposal, not mutation. Evidence becomes a proposal evaluated by governance; the actor cannot write canonical state directly.
  10. Closed learning loop. Approved changes return to the canonical substrate.
  11. System learning. Actors A, B and C receive scoped projections from shared capability.
  12. Many projections. One governed capability becomes wiki, runbook, checklist, training, AI skill and workflow representations.

I think “materialized capability view” is the technically cleaner phrase. In Neo4j’s NAMS Agent Skills architecture, informed by the Agent Instruction Protocol (AIP) and its structured representation of skills, the deeper capability can carry grounding, provenance, procedure structure, review state, versions, drift and repair. SKILL.md is therefore not necessarily the canonical capability. It can be one portable runtime representation of its current approved state.

This is not something AgentMemory for .NET currently implements. It is one of the architectures that, alongside José’s work, triggered the broader question in this essay.

01 · Living capability
Living capabilityevidence · provenance · procedure · evaluation · versions · drift · repair
↓
Materialized viewsSKILL.md · executable graph · human review UI · API surface
Explain this simply

Most people imagine agent memory as something the agent owns — like a backpack it carries from one task to the next.

That feels natural because the runtime speaks as one voice. It says “I remember,” so we picture one actor collecting its own notes, habits and tools.

But the memory could work more like a library. The system holds the collection; each agent receives a temporary reading desk with only the material this user, task and moment permit.

That changes what we can control. Identity, permissions and context decide which view appears, while useful evidence can return to the shared collection without letting one agent rewrite it directly.

From the agent side, it still feels like memory. From the system side, it is a governed runtime view of shared state — the same architecture seen from opposite sides.

The DeltaDB detour

This is where another idea helped me understand what I was reaching for.

The team behind Delta does something conceptually strange and useful: shared worktree state can live deeper than the local filesystem representation that editors and compilers use.

Each participant still gets an ordinary synchronized checkout. VS Code edits files. Node reads them. Tests and builds run from them. Nothing about the familiar working interface has to disappear. But those files are a materialized working view of deeper shared state.

02 · DeltaDB analogy
DeltaDBshared worktree state · change history · collaboration state
↓
Materialized checkoutordinary project files on each machine
↓
Existing toolseditor · compiler · tests · application runtime

Delta does not abolish the filesystem. It demotes it from deepest source of truth to runtime representation.

That distinction lets existing tools keep the interface they need while the canonical model moves deeper into the system. And suddenly SKILL.md looked similar to me: the file still matters to the runtime, but it no longer has to be the place where capability fundamentally lives.

The repo already points in this direction

What makes this question more interesting is that José’s actual architecture is already less agent-centric than the phrase “agent memory” suggests.

Framework-specific surfaces sit around a reusable core. José’s implementation already separates stored memory, context assembly and runtime projection. Projection is an explicit architectural layer rather than one large prompt-building function. Owner and application scope are also represented as system concerns, although isolation remains configurable rather than automatic in every deployment.

Where this is in the repo

MemoryContextProjector runs inside MemoryContextAssembler after budgeting and reranking, in both the live and as-of paths, producing a surface-neutral MemoryContext.Projection. Three render surfaces consume that projection through ProjectionRenderer.

MemoryContextAssembler
        ↓
MemoryContextProjector
        ↓
MemoryContext.Projection
        ↓
ProjectionRenderer

The projection layer is deliberately opt-in: all six MemoryProjectionOptions flags default to false, and with projection disabled Projection is null while the existing render paths remain byte-identical, guarded by SHA256 fingerprint tests. That makes the feature feel less like a prompt rewrite and more like a carefully isolated architectural consolidation.

The framework boundary is explicit too: Microsoft Agent Framework, Semantic Kernel and MCP integrations depend inward on Core, not the reverse. Direct .NET usage can drive the same services without adopting an agent framework. Owner scope runs through MemoryScope; application identity and store routing are separate host-level concerns. MemoryIsolationMode defaults to SingleTenant for backward compatibility; WarnOnUnscoped and the fail-closed StrictMultiTenant mode are opt-in.

Architecture overview · Integration and scoping guide

This is why the projection idea felt less speculative to me. The implementation already distinguishes assembly, projection and rendering. My question is what happens if that same separation becomes the architectural primitive not only for memory, but for capability, policy and learned behaviour.

Verified against main on 28 August 2026.

This separation exists today for memory projection. Extending the same pattern to shared capability, write-back and governed system learning is my extrapolation.

So my question is not “why did you build it this way?” Much of the machinery is already there. His runtime integration is agent-lifecycle-centric, but the underlying architecture increasingly looks like a reusable substrate.

I am wondering what happens if the projection idea is pushed one step further: from projecting memory into a runtime to projecting a broader shared capability into different actors, interfaces and human views—and allowing execution to propose governed changes back.

Projection is only half the story

The Delta analogy is not only about reading a materialized view. A developer edits the checkout, and changes can flow back into the deeper shared state.

That return path matters for adaptive AI architecture. An actor can operate on a runtime projection and produce evidence without being allowed to mutate canonical capability directly.

A. Ephemeral

An actor changes or learns something locally. Nothing persists. The next materialization wipes it.

projection
    ↓
actor
    ↓
scratchpad
    ↓
gone

B. Proposal

This is the most interesting tier. The actor discovers a tool failure, a better procedure, a contradiction, a missing condition or a local exception. It does not commit directly.

experience
    ↓
proposal
    ↓
evaluation
    ↓
promote / reject
    ↓
canonical capability

This is where AgentEval became important to my thought experiment. AgentEval provides both development-time evaluation and, increasingly, runtime enforcement through its Gatekeeper architecture. But neither of those is the adaptive promotion loop I am describing here.

The speculative step is different: execution produces evidence and a proposed capability change; evaluation tests that proposal; and only then can governance decide whether it should become shared capability.

The loop above is a proposed architecture, not current AgentMemory, AgentEval or NAMS runtime behaviour.

C. Governed release

In regulated environments, evaluation may not be enough. Candidate capability can require human or policy approval before it becomes shared state.

Proposed system architecture — not a description of the current AgentMemory, AgentEval or NAMS learning loop.

candidate capability
    ↓
eval
    ↓
human / policy approval
    ↓
approved capability
    ↓
new runtime projections
03 · Projection and governed write-back
Canonical substrateapproved capability · evidence · provenance · policy · history
↓
Runtime projectionrelevant state · skills · tools · permissions · instructions
↓
Actorinterpret · reason · decide · synthesize
↓
Experienceoutcomes · failures · discoveries · exceptions
↓
Proposalproposal, not direct mutation
↓
Eval / governancepromote · reject · request evidence
↺ Canonical substrate

Why write-back matters

Capture at the point of discovery

The actor is present when a failure happens. Without a write-back path, a human has to notice the discovery, understand it, and transcribe it into some other system later. Much of the signal disappears between those steps.

Negative knowledge

Systems produce enormous amounts of evidence about what failed. Runbooks rarely contain it. Organisations repeatedly relearn that a tool breaks under a particular condition or that a procedure fails for a particular case.

Provenance

A capability change should link back to the executions that motivated it. Otherwise we cannot answer: Why does this instruction exist? Does the evidence still hold?

Repair latency

One policy changes. Five copied instructions become stale. Proposal, evaluation and promotion can update one canonical capability; projections refresh from there.

Local variants without forking

A deviation can be stored together with the condition under which it applies. Every local exception does not need to become a separate copied skill.

What if context itself is virtualized?

Instead of saying the agent has memory, skills and tools, we could say: the system materializes a runtime world for this execution.

04 · Context virtualization
Canonical substrateworld · experience · capabilities · provenance · evaluation · policy · permissions · history
↓
Runtime projectionrelevant state · history · skills · tools · permissions · instructions
↓
Actorinterpret · reason · decide · synthesize
↓
Experience / proposalevidence and candidate changes
↓
Eval / governancedecides what becomes shared learning
↺ Canonical substrate

The agent experiences that runtime world as its context.

From inside the execution it may look like “my memory, my tools, my skills, my context.” Architecturally, those may be projections of deeper system state. What flows back is not an unchecked rewrite of that state, but evidence and proposals for governed promotion.

From actor-local learning to system learning

José’s memory and evaluation work made me imagine an actor-level loop like this:

actor
    ↓
experience
    ↓
memory
    ↓
eval
    ↓
improved capability
    ↺

AgentEval spans development-time evaluation and emerging runtime governance. The move from those controls to an adaptive capability-promotion loop is the architectural step I am speculating about.

Lift that loop one level up and it scales differently:

05 · From actor learning to system learning
Actor-local loop
Actor
↓
Experience
↓
Memory
↓
Eval
↓
Improved capability
↺
Shared system loop
All actor experiences
↓
Shared evaluation
↓
Shared capability substrate
↓
Scoped projections
↓
Actor A / B / C
↺

If Actor A discovers that a tool is unreliable under a specific condition, Actor B should not have to rediscover that independently. If one workflow repeatedly fails because a policy changed, every assistant should not carry its own stale version until somebody repairs it.

Maybe the scaling move is not more agents. Maybe it is moving learning from actor-local memory into shared system learning.

Three things that make this hard

01

Permission-scoped learning

Actor A learns from information Actor B is not allowed to see. Can the derived lesson be shared? Provenance and classification may need to travel with the learned capability.

02

Contradictory lessons

Actor A learns X works. Actor B learns X fails. There is no compiler for semantic disagreement.

03

Missing shared success metric

System learning assumes we know what “worked” means. Faster? Safer? Cheaper? More compliant? Better for the user? Organisations often do not have one shared metric.

System learning needs a definition of good. Organisations are often much less explicit about that than software architectures assume.

Isn’t this just RAG again?

RAG is primarily a read-time architecture.

RAG · read-time
documents
↓
retrieve
↓
LLM
↓
answer
Adaptive substrate · closed loop
capability
↓
projection
↓
execution
↓
evidence
↓
proposal
↓
eval
↓
capability
↺

RAG can still be part of the substrate. This is not “better RAG.” The differentiator is the closed learning loop.

RAG externalises knowledge. The interesting change begins when execution can produce governed changes back into the system.

The easy version of the move: Basel

I already accept this architecture for computation in my Basel spatial graph: the LLM should not estimate routes, travel time or transit relationships. The deterministic spatial and temporal system owns those answers. The model extracts intent and constraints, then explains a computed result.

Taken together, these architectures made me wonder whether the same move applies to learned behaviour: not only move truth out of the model, but move capability and learning out of the actor.

Stop maintaining six truths

Organisations routinely take one real process and copy it into a wiki, a runbook, a checklist, training, AI instructions and workflow configuration. Each representation becomes a separately maintained truth. Each drifts.

ONE PROCESS
    ↓
wiki · runbook · checklist · training · AI instructions · workflow configuration
    ↓
six maintained truths
    ↓
drift

What if those are not six maintained truths, but six projections of one governed capability?

06 · One capability, many projections
Living capability substrateprocedure · evidence · policy · variants · evaluation · provenance
↓
Materialized viewswiki view · human checklist · AI runtime skill · reviewer view · workflow / executable representation

Stop maintaining six truths.

Where could this actually matter?

The model becomes interesting where one changing capability has to serve many consumers, roles or runtimes. Not every chatbot needs this. A simple document Q&A system probably does not.

Incident responseOne evolving operational capability, many runtime views.

Runbooks, wiki pages, postmortems and actual operator behaviour drift apart quickly. A shared capability substrate could accumulate successful and failed incident traces, tool behaviour, escalation patterns and postmortem findings.

shared incident capability
→ on-call engineer view
→ incident commander view
→ AI runtime view
→ audit / training view

The value is not “better chat”. It is one evolving operational capability that can be materialized differently for several roles.

Regulated workflowsPolicy, provenance, permissions and role-specific projections.

Banking, insurance, pharma and other regulated domains combine changing policy, permissions, evidence and auditability. Different actors need different representations of the same underlying process.

policy + procedure + provenance + eval
→ analyst runtime
→ reviewer runtime
→ compliance view
→ audit evidence

This is where keeping the source of truth deeper than a collection of prose instructions becomes especially attractive.

Organisational process knowledgeOne real process, many human and machine representations.

Think “How do we actually hire someone here?” The real process contains formal steps, systems, exceptions, local practice and recurring failure points.

one living process
→ hiring manager checklist
→ HR workflow
→ employee onboarding view
→ AI guidance
→ governance view

Instead of maintaining five slightly different truths, maintain one richer capability and materialize the views.

Software engineeringCapability as infrastructure rather than duplicated docs.

A deployment capability rarely lives only in DEPLOYMENT.md. It is distributed across repository structure, CI, tests, failed deploys, approval rules and environment constraints.

deployment capability substrate
→ SKILL.md
→ developer docs
→ CI validation
→ Codex / Copilot runtime
→ visual workflow

This is the Delta analogy at its strongest: existing tools keep the interface they need while the deeper capability lives somewhere more structured.

Customer supportShared learning from resolutions, failures and product change.

Support knowledge changes constantly as bugs appear, workarounds become obsolete and successful resolution patterns accumulate.

all support experience
→ evaluation
→ shared resolution capability
→ customer assistant
→ frontline support
→ specialist / QA

Learning once at system level is more attractive than teaching each assistant the same lesson separately.

AI adoption at scaleShared policy and capability underneath many copilots and assistants.

As organisations accumulate copilots, assistants, prompts and workflows, the real scaling problem may become maintaining what all of them are allowed to know and do.

approved tools + policy + data classification
+ workflows + incidents + eval
→ HR assistant
→ research assistant
→ M365 assistant
→ employee guidance
→ governance view

This may be where the architecture becomes organisational rather than merely technical: shared capability underneath many AI-facing interfaces.

The ownership question

From inside the runtime:

I remember.
I have these skills.
I can use these tools.

From outside:

The system projected relevant state,
capabilities and tools
into this execution context.

Maybe agent memory and adaptive system middleware are the same architecture viewed from opposite sides of the runtime boundary.

The question I want to ask José is now the short one:

Open question

At what point does agent memory become application architecture?

I don’t want to build intelligent little people with software attached. I want to build good software systems in which probabilistic reasoning is one capability among others.

Starting points

Start

Bring one real AI adoption question.

We can turn it into something clearer: a decision, a first experiment, a governance question, or a practical next step.