# Maybe the agent shouldn’t own the memory.

_A lunch conversation, a graph-memory framework, a map of Basel and a strange detour through DeltaDB left me with a question: what if “agent memory” is really the beginning of a shared adaptive system layer?_

Don’t let the metaphor lead your architecture.

We call it an agent. So we give it memory, skills, identity and tools.

But does the agent own those things because the system needs it to — or because the human metaphor made that architecture feel natural?

That question stayed with me after a lunch conversation about agent memory.

Today at EMPOWER in Zurich I met [José Luis Latorre](https://www.linkedin.com/in/joslat/).

José was kind enough to endure a series of badly worded questions from me over lunch, and I am genuinely thankful for the time he took. I understood only part of what he was describing in the moment. The dangerous thing was that I understood enough for it to keep bothering me afterwards.

So I spent the evening looking through [AgentMemory for .NET](https://github.com/joslat/agent-memory-dotnet), [AgentEval](https://github.com/AgentEvalHQ/AgentEval), and [Neo4j’s NAMS / Agent Memory work](https://github.com/neo4j-labs/agent-memory) around graph-native memory and skill distillation.

> **A note on perspective**
>
> I am not a deep .NET or Neo4j engineer. José is operating technically several levels below what I build day to day. I prototype, but I come at these questions from change, organisational development, interfaces, systems thinking and a strong preference for making probabilistic AI only one part of a software system. So this is not a technical review of his work. It is an architectural question that his work triggered for me.

> **What problem this solves**
>
> When one changing process gets copied into runbooks, checklists, training, AI instructions and workflows, organisations end up maintaining several drifting versions of the same capability.

## What I find so good about José’s work

The easy version of “agent memory” is retrieval: something happened before, we stored it, later we retrieve something vaguely relevant and put it back into the context window.

José’s [AgentMemory for .NET](https://github.com/joslat/agent-memory-dotnet) structures experience into persistent short-, long- and reasoning memory, with provenance, temporal state, supersession and an explicit projection layer.

His [AgentEval](https://github.com/AgentEvalHQ/AgentEval) work provides an evidence-oriented evaluation and runtime-governance toolkit for agent systems.

Separately, Neo4j’s [NAMS / Agent Memory work](https://github.com/neo4j-labs/agent-memory) shows how graph-native memory can support the broader direction toward provenance-grounded skills and capability management.

Taken together, these architectures suggest something larger.

And then the really interesting thing happens: **memory stops being only about remembering.**

```text
experience
    ↓
structured memory
    ↓
evaluation
    ↓
patterns
    ↓
capability
    ↓
changed behaviour
    ↺
```

A simple `SKILL.md` file is static prose. It tells a runtime how to behave, but it has no inherent relationship to the executions that produced it, no built-in notion of evidence, drift, failure or repair.

What I find especially interesting in the NAMS Agent Skills work is the shift in what a skill can be. A `SKILL.md` no longer has to be the deepest representation of capability. It can be a materialized runtime view of a capability whose grounding, review state, drift and repair live deeper in the system.

> SKILL.md becomes a snapshot of a living capability.

**Apparent agent ownership.**
 A large agent appears to own memory, skills and tools.

**Structured memory.**
 Memory separates into short-, long- and reasoning memory with provenance, temporal state and supersession.

**Capability emerges.**
 Experience flows through evaluation into capability.

**Architectural inversion.**
 Canonical substrate projects a runtime view into an actor that is only one layer.

**Materialized skill.**
 SKILL.md is a capability projection, not agent property.

**Delta analogy.**
 Deeper state becomes a materialized checkout used by existing tools.

**Context virtualization.**
 Projection carries state, history, skills, tools, permissions and instructions.

**Experience return path.**
 Actor outcomes return as evidence on a separate path.

**Proposal, not mutation.**
 Evidence becomes a proposal evaluated by governance; the actor cannot write canonical state directly.

**Closed learning loop.**
 Approved changes return to the canonical substrate.

**System learning.**
 Actors A, B and C receive scoped projections from shared capability.

**Many projections.**
 One governed capability becomes wiki, runbook, checklist, training, AI skill and workflow representations.

I think “materialized capability view” is the technically cleaner phrase. In Neo4j’s NAMS Agent Skills architecture, informed by the [Agent Instruction Protocol (AIP)](https://github.com/zach-blumenfeld/aip) and its structured representation of skills, the deeper capability can carry grounding, provenance, procedure structure, review state, versions, drift and repair. `SKILL.md` is therefore not necessarily the canonical capability. It can be one portable runtime representation of its current approved state.

This is not something AgentMemory for .NET currently implements. It is one of the architectures that, alongside José’s work, triggered the broader question in this essay.

**01 · Living capability**

A governed capability containing evidence and provenance is materialized into runtime and review views.

```text
Living capability

evidence · provenance · procedure · evaluation · versions · drift · repair

↓

Materialized views

SKILL.md · executable graph · human review UI · API surface
```

### Explain this simply

Most people imagine agent memory as something the agent owns — like a backpack it carries from one task to the next.
That feels natural because the runtime speaks as one voice. It says “I remember,” so we picture one actor collecting its own notes, habits and tools.
But the memory could work more like a library. The system holds the collection; each agent receives a temporary reading desk with only the material this user, task and moment permit.
That changes what we can control. Identity, permissions and context decide which view appears, while useful evidence can return to the shared collection without letting one agent rewrite it directly.
From the agent side, it still feels like memory. From the system side, it is a governed runtime view of shared state — the same architecture seen from opposite sides.

## The DeltaDB detour

This is where another idea helped me understand what I was reaching for.

The team behind [Delta](https://delta.dev/) does something conceptually strange and useful: shared worktree state can live deeper than the local filesystem representation that editors and compilers use.

Each participant still gets an ordinary synchronized checkout. VS Code edits files. Node reads them. Tests and builds run from them. Nothing about the familiar working interface has to disappear. But those files are a materialized working view of deeper shared state.

**02 · DeltaDB analogy**

Shared DeltaDB state becomes an ordinary checkout consumed by existing development tools.

```text
DeltaDB

shared worktree state · change history · collaboration state

↓

Materialized checkout

ordinary project files on each machine

↓

Existing tools

editor · compiler · tests · application runtime
```

> Delta does not abolish the filesystem. It demotes it from deepest source of truth to runtime representation.

That distinction lets existing tools keep the interface they need while the canonical model moves deeper into the system. And suddenly `SKILL.md` looked similar to me: the file still matters to the runtime, but it no longer has to be the place where capability fundamentally lives.

## The repo already points in this direction

What makes this question more interesting is that José’s actual architecture is already less agent-centric than the phrase “agent memory” suggests.

Framework-specific surfaces sit around a reusable core. José’s implementation already separates stored memory, context assembly and runtime projection. Projection is an explicit architectural layer rather than one large prompt-building function. Owner and application scope are also represented as system concerns, although isolation remains configurable rather than automatic in every deployment.

### Technical note: Where this is in the repo

`MemoryContextProjector` runs inside `MemoryContextAssembler` after budgeting and reranking, in both the live and as-of paths, producing a surface-neutral `MemoryContext.Projection`. Three render surfaces consume that projection through `ProjectionRenderer`.

```text
MemoryContextAssembler
        ↓
MemoryContextProjector
        ↓
MemoryContext.Projection
        ↓
ProjectionRenderer
```
The projection layer is deliberately opt-in: all six `MemoryProjectionOptions` flags default to false, and with projection disabled `Projection` is null while the existing render paths remain byte-identical, guarded by SHA256 fingerprint tests. That makes the feature feel less like a prompt rewrite and more like a carefully isolated architectural consolidation.
The framework boundary is explicit too: Microsoft Agent Framework, Semantic Kernel and MCP integrations depend inward on Core, not the reverse. Direct .NET usage can drive the same services without adopting an agent framework. Owner scope runs through `MemoryScope`; application identity and store routing are separate host-level concerns. `MemoryIsolationMode` defaults to `SingleTenant` for backward compatibility; `WarnOnUnscoped` and the fail-closed `StrictMultiTenant` mode are opt-in.
[Architecture overview](https://github.com/joslat/agent-memory-dotnet/blob/main/docs/architecture.md) · [Integration and scoping guide](https://github.com/joslat/agent-memory-dotnet/blob/main/docs/getting-started.md)
This is why the projection idea felt less speculative to me. The implementation already distinguishes assembly, projection and rendering. My question is what happens if that same separation becomes the architectural primitive not only for memory, but for capability, policy and learned behaviour.

Verified against main on 28 August 2026.

This separation exists today for memory projection. Extending the same pattern to shared capability, write-back and governed system learning is my extrapolation.

So my question is not “why did you build it this way?” Much of the machinery is already there. His runtime integration is agent-lifecycle-centric, but the underlying architecture increasingly looks like a reusable substrate.

I am wondering what happens if the projection idea is pushed one step further: from projecting memory into a runtime to projecting a broader shared capability into different actors, interfaces and human views—and allowing execution to propose governed changes back.

## Projection is only half the story

The Delta analogy is not only about reading a materialized view. A developer edits the checkout, and changes can flow back into the deeper shared state.

That return path matters for adaptive AI architecture. An actor can operate on a runtime projection and produce evidence without being allowed to mutate canonical capability directly.

### A. Ephemeral

An actor changes or learns something locally. Nothing persists. The next materialization wipes it.

```text
projection
    ↓
actor
    ↓
scratchpad
    ↓
gone
```

### B. Proposal

This is the most interesting tier. The actor discovers a tool failure, a better procedure, a contradiction, a missing condition or a local exception. It does **not** commit directly.

```text
experience
    ↓
proposal
    ↓
evaluation
    ↓
promote / reject
    ↓
canonical capability
```

This is where AgentEval became important to my thought experiment. AgentEval provides both development-time evaluation and, increasingly, runtime enforcement through its Gatekeeper architecture. But neither of those is the adaptive promotion loop I am describing here.

The speculative step is different: execution produces evidence and a proposed capability change; evaluation tests that proposal; and only then can governance decide whether it should become shared capability.

The loop above is a proposed architecture, not current AgentMemory, AgentEval or NAMS runtime behaviour.

### C. Governed release

In regulated environments, evaluation may not be enough. Candidate capability can require human or policy approval before it becomes shared state.

Proposed system architecture — not a description of the current AgentMemory, AgentEval or NAMS learning loop.

```text
candidate capability
    ↓
eval
    ↓
human / policy approval
    ↓
approved capability
    ↓
new runtime projections
```

**03 · Projection and governed write-back**

A canonical substrate projects context to an actor; experience returns as proposals that evaluation and governance may promote.

```text
Canonical substrate

approved capability · evidence · provenance · policy · history

↓

Runtime projection

relevant state · skills · tools · permissions · instructions

↓

Actor

interpret · reason · decide · synthesize

↓

Experience

outcomes · failures · discoveries · exceptions

↓

Proposal

proposal, not direct mutation

↓

Eval / governance

promote · reject · request evidence

↺ Canonical substrate
```

## Why write-back matters

### Capture at the point of discovery

The actor is present when a failure happens. Without a write-back path, a human has to notice the discovery, understand it, and transcribe it into some other system later. Much of the signal disappears between those steps.

### Negative knowledge

Systems produce enormous amounts of evidence about what failed. Runbooks rarely contain it. Organisations repeatedly relearn that a tool breaks under a particular condition or that a procedure fails for a particular case.

### Provenance

A capability change should link back to the executions that motivated it. Otherwise we cannot answer: **Why does this instruction exist? Does the evidence still hold?**

### Repair latency

One policy changes. Five copied instructions become stale. Proposal, evaluation and promotion can update one canonical capability; projections refresh from there.

### Local variants without forking

A deviation can be stored together with the condition under which it applies. Every local exception does not need to become a separate copied skill.

## What if context itself is virtualized?

Instead of saying the agent has memory, skills and tools, we could say: **the system materializes a runtime world for this execution.**

**04 · Context virtualization**

A system materializes a runtime world for an actor and evaluates evidence before incorporating shared learning.

```text
Canonical substrate

world · experience · capabilities · provenance · evaluation · policy · permissions · history

↓

Runtime projection

relevant state · history · skills · tools · permissions · instructions

↓

Actor

interpret · reason · decide · synthesize

↓

Experience / proposal

evidence and candidate changes

↓

Eval / governance

decides what becomes shared learning

↺ Canonical substrate
```

> The agent experiences that runtime world as its context.

From inside the execution it may look like “my memory, my tools, my skills, my context.” Architecturally, those may be projections of deeper system state. What flows back is not an unchecked rewrite of that state, but evidence and proposals for governed promotion.

## From actor-local learning to system learning

José’s memory and evaluation work made me imagine an actor-level loop like this:

```text
actor
    ↓
experience
    ↓
memory
    ↓
eval
    ↓
improved capability
    ↺
```

AgentEval spans development-time evaluation and emerging runtime governance. The move from those controls to an adaptive capability-promotion loop is the architectural step I am speculating about.

Lift that loop one level up and it scales differently:

**05 · From actor learning to system learning**

Actor-local memory and evaluation expand into shared evaluation, capability, and scoped projections for many actors.

```text
Actor-local loop

Actor
↓
Experience
↓
Memory
↓
Eval
↓
Improved capability
↺

Shared system loop

All actor experiences
↓
Shared evaluation
↓
Shared capability substrate
↓
Scoped projections
↓
Actor A / B / C
↺
```

If Actor A discovers that a tool is unreliable under a specific condition, Actor B should not have to rediscover that independently. If one workflow repeatedly fails because a policy changed, every assistant should not carry its own stale version until somebody repairs it.

> Maybe the scaling move is not more agents. Maybe it is moving learning from actor-local memory into shared system learning.

## Three things that make this hard

01

### Permission-scoped learning
Actor A learns from information Actor B is not allowed to see. Can the derived lesson be shared? Provenance and classification may need to travel with the learned capability.

02

### Contradictory lessons
Actor A learns X works. Actor B learns X fails. There is no compiler for semantic disagreement.

03

### Missing shared success metric
System learning assumes we know what “worked” means. Faster? Safer? Cheaper? More compliant? Better for the user? Organisations often do not have one shared metric.

> System learning needs a definition of good. Organisations are often much less explicit about that than software architectures assume.

## Isn’t this just RAG again?

RAG is primarily a read-time architecture.

RAG · read-time

```text
documents
↓
retrieve
↓
LLM
↓
answer
```

Adaptive substrate · closed loop

```text
capability
↓
projection
↓
execution
↓
evidence
↓
proposal
↓
eval
↓
capability
↺
```

RAG can still be part of the substrate. This is not “better RAG.” The differentiator is the closed learning loop.

> RAG externalises knowledge. The interesting change begins when execution can produce governed changes back into the system.

## The easy version of the move: Basel

I already accept this architecture for computation in my Basel spatial graph: the LLM should not estimate routes, travel time or transit relationships. The deterministic spatial and temporal system owns those answers. The model extracts intent and constraints, then explains a computed result.

Taken together, these architectures made me wonder whether the same move applies to learned behaviour: not only move truth out of the model, but move capability and learning out of the actor.

## Stop maintaining six truths

Organisations routinely take one real process and copy it into a wiki, a runbook, a checklist, training, AI instructions and workflow configuration. Each representation becomes a separately maintained truth. Each drifts.

```text
ONE PROCESS
    ↓
wiki · runbook · checklist · training · AI instructions · workflow configuration
    ↓
six maintained truths
    ↓
drift
```

What if those are not six maintained truths, but six projections of one governed capability?

**06 · One capability, many projections**

One governed capability is materialized as wiki, checklist, AI skill, review, and workflow views.

```text
Living capability substrate

procedure · evidence · policy · variants · evaluation · provenance

↓

Materialized views

wiki view · human checklist · AI runtime skill · reviewer view · workflow / executable representation
```

> Stop maintaining six truths.

## Where could this actually matter?

The model becomes interesting where one changing capability has to serve many consumers, roles or runtimes. Not every chatbot needs this. A simple document Q&A system probably does not.

### Incident response

One evolving operational capability, many runtime views.
Runbooks, wiki pages, postmortems and actual operator behaviour drift apart quickly. A shared capability substrate could accumulate successful and failed incident traces, tool behaviour, escalation patterns and postmortem findings.

```text
shared incident capability
→ on-call engineer view
→ incident commander view
→ AI runtime view
→ audit / training view
```
The value is not “better chat”. It is one evolving operational capability that can be materialized differently for several roles.

### Regulated workflows

Policy, provenance, permissions and role-specific projections.
Banking, insurance, pharma and other regulated domains combine changing policy, permissions, evidence and auditability. Different actors need different representations of the same underlying process.

```text
policy + procedure + provenance + eval
→ analyst runtime
→ reviewer runtime
→ compliance view
→ audit evidence
```
This is where keeping the source of truth deeper than a collection of prose instructions becomes especially attractive.

### Organisational process knowledge

One real process, many human and machine representations.
Think “How do we actually hire someone here?” The real process contains formal steps, systems, exceptions, local practice and recurring failure points.

```text
one living process
→ hiring manager checklist
→ HR workflow
→ employee onboarding view
→ AI guidance
→ governance view
```
Instead of maintaining five slightly different truths, maintain one richer capability and materialize the views.

### Software engineering

Capability as infrastructure rather than duplicated docs.
A deployment capability rarely lives only in DEPLOYMENT.md. It is distributed across repository structure, CI, tests, failed deploys, approval rules and environment constraints.

```text
deployment capability substrate
→ SKILL.md
→ developer docs
→ CI validation
→ Codex / Copilot runtime
→ visual workflow
```
This is the Delta analogy at its strongest: existing tools keep the interface they need while the deeper capability lives somewhere more structured.

### Customer support

Shared learning from resolutions, failures and product change.
Support knowledge changes constantly as bugs appear, workarounds become obsolete and successful resolution patterns accumulate.

```text
all support experience
→ evaluation
→ shared resolution capability
→ customer assistant
→ frontline support
→ specialist / QA
```
Learning once at system level is more attractive than teaching each assistant the same lesson separately.

### AI adoption at scale

Shared policy and capability underneath many copilots and assistants.
As organisations accumulate copilots, assistants, prompts and workflows, the real scaling problem may become maintaining what all of them are allowed to know and do.

```text
approved tools + policy + data classification
+ workflows + incidents + eval
→ HR assistant
→ research assistant
→ M365 assistant
→ employee guidance
→ governance view
```
This may be where the architecture becomes organisational rather than merely technical: shared capability underneath many AI-facing interfaces.

## The ownership question

From inside the runtime:

```text
I remember.
I have these skills.
I can use these tools.
```

From outside:

```text
The system projected relevant state,
capabilities and tools
into this execution context.
```

> Maybe agent memory and adaptive system middleware are the same architecture viewed from opposite sides of the runtime boundary.

The question I want to ask José is now the short one:

> **Open question**
>
> At what point does agent memory become application architecture?

> I don’t want to build intelligent little people with software attached. I want to build good software systems in which probabilistic reasoning is one capability among others.

### Starting points

- [José Luis Latorre — LinkedIn](https://www.linkedin.com/in/joslat/)
- [AgentMemory for .NET — GitHub](https://github.com/joslat/agent-memory-dotnet)
- [AgentEval — GitHub](https://github.com/AgentEvalHQ/AgentEval)
- [Neo4j Agent Memory / NAMS — GitHub](https://github.com/neo4j-labs/agent-memory)
- [Agent Instruction Protocol — GitHub](https://github.com/zach-blumenfeld/aip)
- [Delta](https://delta.dev/)
