<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/">
  <channel>
    <title>Bridge Work</title>
    <link>https://bridge-work.ai</link>
    <atom:link href="https://bridge-work.ai/rss.xml" rel="self" type="application/rss+xml" />
    <description>AI adoption, governance translation and architecture for real organisational work.</description>
    <language>de-CH</language>
    <lastBuildDate>Wed, 02 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Maybe the agent shouldn’t own the memory.</title>
      <link>https://bridge-work.ai/en/blog/maybe-the-agent-shouldnt-own-the-memory/</link>
      <guid isPermaLink="true">https://bridge-work.ai/en/blog/maybe-the-agent-shouldnt-own-the-memory/</guid>
      <description>What if memory, skills and learning belong to the system rather than the agent? A reflection on adaptive middleware, runtime projection and system learning.</description>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      
      <dc:language>en</dc:language>
      <category>AI Architecture</category><category>Memory</category><category>Graphs</category><category>System Learning</category>
      <content:encoded><![CDATA[Don’t let the metaphor lead your architecture.

We call it an agent. So we give it memory, skills, identity and tools.

But does the agent own those things because the system needs it to — or because the human metaphor made that architecture feel natural?

That question stayed with me after a lunch conversation about agent memory.

Today at EMPOWER in Zurich I met José Luis Latorre.

José was kind enough to endure a series of badly worded questions from me over lunch, and I am genuinely thankful for the time he took. I understood only part of what he was describing in the moment. The dangerous thing was that I understood enough for it to keep bothering me afterwards.

So I spent the evening looking through AgentMemory for .NET, AgentEval, and Neo4j’s NAMS / Agent Memory work around graph-native memory and skill distillation.

A note on perspective
I am not a deep .NET or Neo4j engineer. José is operating technically several levels below what I build day to day. I prototype, but I come at these questions from change, organisational development, interfaces, systems thinking and a strong preference for making probabilistic AI only one part of a software system. So this is not a technical review of his work. It is an architectural question that his work triggered for me.

What problem this solves
When one changing process gets copied into runbooks, checklists, training, AI instructions and workflows, organisations end up maintaining several drifting versions of the same capability.

What I find so good about José’s work

The easy version of “agent memory” is retrieval: something happened before, we stored it, later we retrieve something vaguely relevant and put it back into the context window.

José’s AgentMemory for .NET structures experience into persistent short-, long- and reasoning memory, with provenance, temporal state, supersession and an explicit projection layer.

His AgentEval work provides an evidence-oriented evaluation and runtime-governance toolkit for agent systems.

Separately, Neo4j’s NAMS / Agent Memory work shows how graph-native memory can support the broader direction toward provenance-grounded skills and capability management.

Taken together, these architectures suggest something larger.

And then the really interesting thing happens: memory stops being only about remembering.

experience
    ↓
structured memory
    ↓
evaluation
    ↓
patterns
    ↓
capability
    ↓
changed behaviour
    ↺

A simple SKILL.md file is static prose. It tells a runtime how to behave, but it has no inherent relationship to the executions that produced it, no built-in notion of evidence, drift, failure or repair.

What I find especially interesting in the NAMS Agent Skills work is the shift in what a skill can be. A SKILL.md no longer has to be the deepest representation of capability. It can be a materialized runtime view of a capability whose grounding, review state, drift and repair live deeper in the system.

SKILL.md becomes a snapshot of a living capability.

Apparent agent ownership. A large agent appears to own memory, skills and tools.

Structured memory. Memory separates into short-, long- and reasoning memory with provenance, temporal state and supersession.

Capability emerges. Experience flows through evaluation into capability.

Architectural inversion. Canonical substrate projects a runtime view into an actor that is only one layer.

Materialized skill. SKILL.md is a capability projection, not agent property.

Delta analogy. Deeper state becomes a materialized checkout used by existing tools.

Context virtualization. Projection carries state, history, skills, tools, permissions and instructions.

Experience return path. Actor outcomes return as evidence on a separate path.

Proposal, not mutation. Evidence becomes a proposal evaluated by governance; the actor cannot write canonical state directly.

Closed learning loop. Approved changes return to the canonical substrate.

System learning. Actors A, B and C receive scoped projections from shared capability.

Many projections. One governed capability becomes wiki, runbook, checklist, training, AI skill and workflow representations.

I think “materialized capability view” is the technically cleaner phrase. In Neo4j’s NAMS Agent Skills architecture, informed by the Agent Instruction Protocol (AIP) and its structured representation of skills, the deeper capability can carry grounding, provenance, procedure structure, review state, versions, drift and repair. SKILL.md is therefore not necessarily the canonical capability. It can be one portable runtime representation of its current approved state.

This is not something AgentMemory for .NET currently implements. It is one of the architectures that, alongside José’s work, triggered the broader question in this essay.

01 · Living capability
Living capability

evidence · provenance · procedure · evaluation · versions · drift · repair

↓

Materialized views

SKILL.md · executable graph · human review UI · API surface

The DeltaDB detour

This is where another idea helped me understand what I was reaching for.

The team behind Delta does something conceptually strange and useful: shared worktree state can live deeper than the local filesystem representation that editors and compilers use.

Each participant still gets an ordinary synchronized checkout. VS Code edits files. Node reads them. Tests and builds run from them. Nothing about the familiar working interface has to disappear. But those files are a materialized working view of deeper shared state.

02 · DeltaDB analogy
DeltaDB

shared worktree state · change history · collaboration state

↓

Materialized checkout

ordinary project files on each machine

↓

Existing tools

editor · compiler · tests · application runtime

Delta does not abolish the filesystem. It demotes it from deepest source of truth to runtime representation.

That distinction lets existing tools keep the interface they need while the canonical model moves deeper into the system. And suddenly SKILL.md looked similar to me: the file still matters to the runtime, but it no longer has to be the place where capability fundamentally lives.

The repo already points in this direction

What makes this question more interesting is that José’s actual architecture is already less agent-centric than the phrase “agent memory” suggests.

Framework-specific surfaces sit around a reusable core. José’s implementation already separates stored memory, context assembly and runtime projection. Projection is an explicit architectural layer rather than one large prompt-building function. Owner and application scope are also represented as system concerns, although isolation remains configurable rather than automatic in every deployment.

Where this is in the repo

MemoryContextProjector runs inside MemoryContextAssembler after budgeting and reranking, in both the live and as-of paths, producing a surface-neutral MemoryContext.Projection. Three render surfaces consume that projection through ProjectionRenderer.

MemoryContextAssembler
        ↓
MemoryContextProjector
        ↓
MemoryContext.Projection
        ↓
ProjectionRenderer

The projection layer is deliberately opt-in: all six MemoryProjectionOptions flags default to false, and with projection disabled Projection is null while the existing render paths remain byte-identical, guarded by SHA256 fingerprint tests. That makes the feature feel less like a prompt rewrite and more like a carefully isolated architectural consolidation.

The framework boundary is explicit too: Microsoft Agent Framework, Semantic Kernel and MCP integrations depend inward on Core, not the reverse. Direct .NET usage can drive the same services without adopting an agent framework. Owner scope runs through MemoryScope; application identity and store routing are separate host-level concerns. MemoryIsolationMode defaults to SingleTenant for backward compatibility; WarnOnUnscoped and the fail-closed StrictMultiTenant mode are opt-in.

Architecture overview · Integration and scoping guide

This is why the projection idea felt less speculative to me. The implementation already distinguishes assembly, projection and rendering. My question is what happens if that same separation becomes the architectural primitive not only for memory, but for capability, policy and learned behaviour.

Verified against main on 28 August 2026.

This separation exists today for memory projection. Extending the same pattern to shared capability, write-back and governed system learning is my extrapolation.

So my question is not “why did you build it this way?” Much of the machinery is already there. His runtime integration is agent-lifecycle-centric, but the underlying architecture increasingly looks like a reusable substrate.

I am wondering what happens if the projection idea is pushed one step further: from projecting memory into a runtime to projecting a broader shared capability into different actors, interfaces and human views—and allowing execution to propose governed changes back.

Projection is only half the story

The Delta analogy is not only about reading a materialized view. A developer edits the checkout, and changes can flow back into the deeper shared state.

That return path matters for adaptive AI architecture. An actor can operate on a runtime projection and produce evidence without being allowed to mutate canonical capability directly.

A. Ephemeral

An actor changes or learns something locally. Nothing persists. The next materialization wipes it.

projection
    ↓
actor
    ↓
scratchpad
    ↓
gone

B. Proposal

This is the most interesting tier. The actor discovers a tool failure, a better procedure, a contradiction, a missing condition or a local exception. It does not commit directly.

experience
    ↓
proposal
    ↓
evaluation
    ↓
promote / reject
    ↓
canonical capability

This is where AgentEval became important to my thought experiment. AgentEval provides both development-time evaluation and, increasingly, runtime enforcement through its Gatekeeper architecture. But neither of those is the adaptive promotion loop I am describing here.

The speculative step is different: execution produces evidence and a proposed capability change; evaluation tests that proposal; and only then can governance decide whether it should become shared capability.

The loop above is a proposed architecture, not current AgentMemory, AgentEval or NAMS runtime behaviour.

C. Governed release

In regulated environments, evaluation may not be enough. Candidate capability can require human or policy approval before it becomes shared state.

Proposed system architecture — not a description of the current AgentMemory, AgentEval or NAMS learning loop.

candidate capability
    ↓
eval
    ↓
human / policy approval
    ↓
approved capability
    ↓
new runtime projections

03 · Projection and governed write-back
Canonical substrate

approved capability · evidence · provenance · policy · history

↓

Runtime projection

relevant state · skills · tools · permissions · instructions

↓

Actor

interpret · reason · decide · synthesize

↓

Experience

outcomes · failures · discoveries · exceptions

↓

Proposal

proposal, not direct mutation

↓

Eval / governance

promote · reject · request evidence

↺ Canonical substrate

Why write-back matters

Capture at the point of discovery

The actor is present when a failure happens. Without a write-back path, a human has to notice the discovery, understand it, and transcribe it into some other system later. Much of the signal disappears between those steps.

Negative knowledge

Systems produce enormous amounts of evidence about what failed. Runbooks rarely contain it. Organisations repeatedly relearn that a tool breaks under a particular condition or that a procedure fails for a particular case.

Provenance

A capability change should link back to the executions that motivated it. Otherwise we cannot answer: Why does this instruction exist? Does the evidence still hold?

Repair latency

One policy changes. Five copied instructions become stale. Proposal, evaluation and promotion can update one canonical capability; projections refresh from there.

Local variants without forking

A deviation can be stored together with the condition under which it applies. Every local exception does not need to become a separate copied skill.

What if context itself is virtualized?

Instead of saying the agent has memory, skills and tools, we could say: the system materializes a runtime world for this execution.

04 · Context virtualization
Canonical substrate

world · experience · capabilities · provenance · evaluation · policy · permissions · history

↓

Runtime projection

relevant state · history · skills · tools · permissions · instructions

↓

Actor

interpret · reason · decide · synthesize

↓

Experience / proposal

evidence and candidate changes

↓

Eval / governance

decides what becomes shared learning

↺ Canonical substrate

The agent experiences that runtime world as its context.

From inside the execution it may look like “my memory, my tools, my skills, my context.” Architecturally, those may be projections of deeper system state. What flows back is not an unchecked rewrite of that state, but evidence and proposals for governed promotion.

From actor-local learning to system learning

José’s memory and evaluation work made me imagine an actor-level loop like this:

actor
    ↓
experience
    ↓
memory
    ↓
eval
    ↓
improved capability
    ↺

AgentEval spans development-time evaluation and emerging runtime governance. The move from those controls to an adaptive capability-promotion loop is the architectural step I am speculating about.

Lift that loop one level up and it scales differently:

05 · From actor learning to system learning
Actor-local loop

Actor
↓
Experience
↓
Memory
↓
Eval
↓
Improved capability
↺

Shared system loop

All actor experiences
↓
Shared evaluation
↓
Shared capability substrate
↓
Scoped projections
↓
Actor A / B / C
↺

If Actor A discovers that a tool is unreliable under a specific condition, Actor B should not have to rediscover that independently. If one workflow repeatedly fails because a policy changed, every assistant should not carry its own stale version until somebody repairs it.

Maybe the scaling move is not more agents. Maybe it is moving learning from actor-local memory into shared system learning.

Three things that make this hard

01
Permission-scoped learning

Actor A learns from information Actor B is not allowed to see. Can the derived lesson be shared? Provenance and classification may need to travel with the learned capability.

02
Contradictory lessons

Actor A learns X works. Actor B learns X fails. There is no compiler for semantic disagreement.

03
Missing shared success metric

System learning assumes we know what “worked” means. Faster? Safer? Cheaper? More compliant? Better for the user? Organisations often do not have one shared metric.

System learning needs a definition of good. Organisations are often much less explicit about that than software architectures assume.

Isn’t this just RAG again?

RAG is primarily a read-time architecture.

RAG · read-time
documents
↓
retrieve
↓
LLM
↓
answer

Adaptive substrate · closed loop
capability
↓
projection
↓
execution
↓
evidence
↓
proposal
↓
eval
↓
capability
↺

RAG can still be part of the substrate. This is not “better RAG.” The differentiator is the closed learning loop.

RAG externalises knowledge. The interesting change begins when execution can produce governed changes back into the system.

The easy version of the move: Basel

I already accept this architecture for computation in my Basel spatial graph: the LLM should not estimate routes, travel time or transit relationships. The deterministic spatial and temporal system owns those answers. The model extracts intent and constraints, then explains a computed result.

Taken together, these architectures made me wonder whether the same move applies to learned behaviour: not only move truth out of the model, but move capability and learning out of the actor.

Stop maintaining six truths

Organisations routinely take one real process and copy it into a wiki, a runbook, a checklist, training, AI instructions and workflow configuration. Each representation becomes a separately maintained truth. Each drifts.

ONE PROCESS
    ↓
wiki · runbook · checklist · training · AI instructions · workflow configuration
    ↓
six maintained truths
    ↓
drift

What if those are not six maintained truths, but six projections of one governed capability?

06 · One capability, many projections
Living capability substrate

procedure · evidence · policy · variants · evaluation · provenance

↓

Materialized views

wiki view · human checklist · AI runtime skill · reviewer view · workflow / executable representation

Stop maintaining six truths.

Where could this actually matter?

The model becomes interesting where one changing capability has to serve many consumers, roles or runtimes. Not every chatbot needs this. A simple document Q&A system probably does not.

Incident response One evolving operational capability, many runtime views.

Runbooks, wiki pages, postmortems and actual operator behaviour drift apart quickly. A shared capability substrate could accumulate successful and failed incident traces, tool behaviour, escalation patterns and postmortem findings.

shared incident capability
→ on-call engineer view
→ incident commander view
→ AI runtime view
→ audit / training view

The value is not “better chat”. It is one evolving operational capability that can be materialized differently for several roles.

Regulated workflows Policy, provenance, permissions and role-specific projections.

Banking, insurance, pharma and other regulated domains combine changing policy, permissions, evidence and auditability. Different actors need different representations of the same underlying process.

policy + procedure + provenance + eval
→ analyst runtime
→ reviewer runtime
→ compliance view
→ audit evidence

This is where keeping the source of truth deeper than a collection of prose instructions becomes especially attractive.

Organisational process knowledge One real process, many human and machine representations.

Think “How do we actually hire someone here?” The real process contains formal steps, systems, exceptions, local practice and recurring failure points.

one living process
→ hiring manager checklist
→ HR workflow
→ employee onboarding view
→ AI guidance
→ governance view

Instead of maintaining five slightly different truths, maintain one richer capability and materialize the views.

Software engineering Capability as infrastructure rather than duplicated docs.

A deployment capability rarely lives only in DEPLOYMENT.md. It is distributed across repository structure, CI, tests, failed deploys, approval rules and environment constraints.

deployment capability substrate
→ SKILL.md
→ developer docs
→ CI validation
→ Codex / Copilot runtime
→ visual workflow

This is the Delta analogy at its strongest: existing tools keep the interface they need while the deeper capability lives somewhere more structured.

Customer support Shared learning from resolutions, failures and product change.

Support knowledge changes constantly as bugs appear, workarounds become obsolete and successful resolution patterns accumulate.

all support experience
→ evaluation
→ shared resolution capability
→ customer assistant
→ frontline support
→ specialist / QA

Learning once at system level is more attractive than teaching each assistant the same lesson separately.

AI adoption at scale Shared policy and capability underneath many copilots and assistants.

As organisations accumulate copilots, assistants, prompts and workflows, the real scaling problem may become maintaining what all of them are allowed to know and do.

approved tools + policy + data classification
workflows + incidents + eval
→ HR assistant
→ research assistant
→ M365 assistant
→ employee guidance
→ governance view

This may be where the architecture becomes organisational rather than merely technical: shared capability underneath many AI-facing interfaces.

The ownership question

From inside the runtime:

I remember.
I have these skills.
I can use these tools.

From outside:

The system projected relevant state,
capabilities and tools
into this execution context.

Maybe agent memory and adaptive system middleware are the same architecture viewed from opposite sides of the runtime boundary.

The question I want to ask José is now the short one:

Open question
At what point does agent memory become application architecture?

I don’t want to build intelligent little people with software attached. I want to build good software systems in which probabilistic reasoning is one capability among others.

Starting points

José Luis Latorre — LinkedIn
AgentMemory for .NET — GitHub
AgentEval — GitHub
Neo4j Agent Memory / NAMS — GitHub
Agent Instruction Protocol — GitHub
Delta]]></content:encoded>
    </item>
<item>
      <title>Vielleicht sollte der Agent das Gedächtnis nicht besitzen.</title>
      <link>https://bridge-work.ai/blog/vielleicht-sollte-der-agent-das-gedaechtnis-nicht-besitzen/</link>
      <guid isPermaLink="true">https://bridge-work.ai/blog/vielleicht-sollte-der-agent-das-gedaechtnis-nicht-besitzen/</guid>
      <description>Was, wenn Gedächtnis, Skills und Lernen zum System gehören statt zum Agenten? Eine Reflexion über adaptive Middleware, Laufzeitprojektion und Systemlernen.</description>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      
      <dc:language>de</dc:language>
      <category>KI-Architektur</category><category>Gedächtnis</category><category>Graphen</category><category>Systemlernen</category>
      <content:encoded><![CDATA[Lass nicht die Metapher deine Architektur bestimmen.

Wir nennen es einen Agenten. Also geben wir ihm Gedächtnis, Fähigkeiten, Identität und Werkzeuge.

Aber gehören diese Dinge wirklich dem Agenten, weil das System es so braucht — oder weil die menschliche Metapher diese Architektur selbstverständlich erscheinen lässt?

Diese Frage ging mir nach einem Gespräch über Agent Memory beim Mittagessen nicht mehr aus dem Kopf.

Heute habe ich am EMPOWER in Zürich José Luis Latorre getroffen.

José war so freundlich, beim Mittagessen eine Reihe schlecht formulierter Fragen von mir auszuhalten, und ich bin ihm für seine Zeit wirklich dankbar. Im Moment selbst verstand ich nur einen Teil dessen, was er beschrieb. Das Gefährliche war: Ich verstand genug, damit es mich danach nicht mehr losliess.

Also verbrachte ich den Abend mit AgentMemory for .NET, AgentEval und Neo4js NAMS Agent Skills rund um graphenbasiertes Gedächtnis und Skill-Destillation.

Eine Anmerkung zur Perspektive
Ich bin kein tiefgehender .NET- oder Neo4j-Engineer. José arbeitet technisch mehrere Ebenen unter dem, was ich im Alltag baue. Ich entwickle Prototypen, komme an diese Fragen aber aus Veränderung, Organisationsentwicklung, Schnittstellen, Systemdenken und einer starken Präferenz dafür, probabilistische KI nur zu einem Teil eines Softwaresystems zu machen. Dies ist deshalb kein technisches Review seiner Arbeit. Es ist eine architektonische Frage, die seine Arbeit bei mir ausgelöst hat.

Welches Problem das löst
Wenn ein veränderlicher Prozess in Runbooks, Checklisten, Schulungen, KI-Anweisungen und Workflows kopiert wird, pflegen Organisationen am Ende mehrere auseinanderdriftende Versionen derselben Fähigkeit.

Was ich an Josés Arbeit so gut finde

Die einfache Version von „Agentengedächtnis“ ist Retrieval: Etwas ist früher passiert, wir speichern es, später holen wir etwas ungefähr Relevantes zurück und legen es wieder ins Kontextfenster.

Josés AgentMemory strukturiert Erfahrung als persistentes Kurzzeit-, Langzeit- und Schlussfolgerungsgedächtnis – mit Provenienz, zeitlichem Zustand, Ablösung und einer expliziten Projektionsschicht.

Seine Arbeit an AgentEval ergänzt dieses System um eine evidenzorientierte Evaluationsschicht für die Entwicklung.

Separat zeigt Neo4js NAMS-Agent-Skills-Arbeit, wie angesammeltes Gedächtnis in provenienzbasierte Skills mit menschlichem Review, Drifterkennung und Reparatur verdichtet werden kann.

Zusammen deuten diese Architekturen auf etwas Grösseres hin.

Und dann geschieht das eigentlich Interessante: Gedächtnis dient nicht mehr nur dem Erinnern.

Erfahrung
    ↓
strukturiertes Gedächtnis
    ↓
Evaluation
    ↓
Muster
    ↓
Fähigkeit
    ↓
verändertes Verhalten
    ↺

Eine einfache SKILL.md-Datei ist statischer Text. Sie sagt einer Laufzeit, wie sie sich verhalten soll, hat aber keine inhärente Beziehung zu den Ausführungen, aus denen sie entstanden ist, und kein eingebautes Konzept von Evidenz, Drift, Fehlern oder Reparatur.

Besonders interessant finde ich an NAMS Agent Skills die Verschiebung dessen, was ein Skill sein kann. Eine SKILL.md muss nicht mehr die tiefste Repräsentation einer Fähigkeit sein. Sie kann eine materialisierte Laufzeitansicht einer Fähigkeit sein, deren Grundlage, Review-Status, Drift und Reparatur tiefer im System liegen.

SKILL.md wird zur Momentaufnahme einer lebendigen Fähigkeit.

Scheinbarer Besitz des Agenten. Ein grosser Agent scheint Gedächtnis, Skills und Werkzeuge zu besitzen.

Strukturiertes Gedächtnis. Gedächtnis trennt sich in kurz, lang und schlussfolgernd, mit Provenienz, zeitlichem Zustand und Ablösung.

Fähigkeit entsteht. Erfahrung fliesst durch Evaluation in Fähigkeit.

Architektonische Umkehrung. Das kanonische Substrat projiziert eine Laufzeitansicht in einen Akteur, der nur eine Ebene ist.

Materialisierter Skill. SKILL.md ist eine Projektion der Fähigkeit, kein Eigentum des Agenten.

Delta-Analogie. Tieferer Zustand wird zum materialisierten Checkout für bestehende Werkzeuge.

Kontextvirtualisierung. Die Projektion trägt Zustand, Historie, Skills, Werkzeuge, Berechtigungen und Anweisungen.

Rückweg der Erfahrung. Ergebnisse des Akteurs kehren als Evidenz über einen getrennten Pfad zurück.

Vorschlag statt Mutation. Evidenz wird zum Vorschlag für Governance; der Akteur schreibt nicht direkt in den kanonischen Zustand.

Geschlossener Lernkreislauf. Genehmigte Änderungen kehren ins kanonische Substrat zurück.

Systemlernen. Akteure A, B und C erhalten begrenzte Projektionen der gemeinsamen Fähigkeit.

Viele Projektionen. Eine gesteuerte Fähigkeit wird zu Wiki, Runbook, Checkliste, Training, KI-Skill und Workflow.

„Materialisierte Fähigkeitsansicht“ ist wohl der technisch sauberere Ausdruck. In dieser NAMS/AIP-Architektur enthält die tiefere Fähigkeit Evidenz, Provenienz, Verfahren, Evaluation, Versionierung, Drift und Reparatur. SKILL.md ist nicht zwingend die kanonische Fähigkeit. Sie ist eine materialisierte Laufzeitansicht ihres aktuell genehmigten Zustands.

AgentMemory for .NET implementiert dies heute nicht. Es ist eine der Architekturen, die zusammen mit Josés Arbeit die grössere Frage dieses Essays ausgelöst hat.

01 · Lebendige Fähigkeit
Lebendige Fähigkeit

Evidenz · Provenienz · Verfahren · Evaluation · Versionen · Drift · Reparatur

↓

Materialisierte Ansichten

SKILL.md · ausführbarer Graph · Review-Oberfläche · API

Der Umweg über DeltaDB

Hier half mir eine andere Idee zu verstehen, wonach ich suchte.

Das Team hinter Delta macht etwas konzeptionell Seltsames und Nützliches: Gemeinsamer Worktree-Zustand kann tiefer liegen als die lokale Dateisystemdarstellung, die Editoren und Compiler verwenden.

Alle Beteiligten erhalten weiterhin einen gewöhnlichen synchronisierten Checkout. VS Code bearbeitet Dateien. Node liest sie. Tests und Builds laufen damit. Nichts an der vertrauten Arbeitsschnittstelle muss verschwinden. Doch diese Dateien sind eine materialisierte Arbeitsansicht eines tieferen gemeinsamen Zustands.

02 · DeltaDB-Analogie
DeltaDB

gemeinsamer Worktree-Zustand · Änderungshistorie · Kollaborationszustand

↓

Materialisierter Checkout

gewöhnliche Projektdateien auf jedem Rechner

↓

Bestehende Werkzeuge

Editor · Compiler · Tests · Anwendungslaufzeit

Delta schafft das Dateisystem nicht ab. Es stuft es von der tiefsten Wahrheitsquelle zur Laufzeitrepräsentation zurück.

Diese Unterscheidung lässt bestehenden Werkzeugen ihre gewohnte Schnittstelle, während das kanonische Modell tiefer ins System wandert. Plötzlich sah SKILL.md für mich ähnlich aus: Die Datei bleibt für die Laufzeit wichtig, muss aber nicht der Ort sein, an dem die Fähigkeit grundsätzlich lebt.

Das Repository weist bereits in diese Richtung

Die Frage wird interessanter, weil Josés tatsächliche Architektur schon weniger agentenzentriert ist, als der Ausdruck „Agentengedächtnis“ vermuten lässt.

Framework-spezifische Oberflächen liegen um einen wiederverwendbaren Kern. Josés Implementierung trennt gespeichertes Gedächtnis, Kontextzusammenstellung und Laufzeitprojektion. Projektion ist eine explizite Architekturschicht statt einer einzigen grossen Prompt-Funktion. Owner- und Anwendungsscope sind ebenfalls als Systembelange repräsentiert, auch wenn Isolation nicht in jedem Deployment automatisch aktiviert ist.

Wo dies im Repository zu finden ist

MemoryContextProjector läuft innerhalb von MemoryContextAssembler nach Budgetierung und Reranking – sowohl im Live- als auch im As-of-Pfad – und erzeugt eine oberflächenneutrale MemoryContext.Projection. Drei Render-Oberflächen konsumieren diese Projektion über ProjectionRenderer.

MemoryContextAssembler
        ↓
MemoryContextProjector
        ↓
MemoryContext.Projection
        ↓
ProjectionRenderer

Die Projektionsschicht ist bewusst opt-in: Alle sechs Flags in MemoryProjectionOptions sind standardmässig false. Bei deaktivierter Projektion ist Projection null, während die bestehenden Render-Pfade byte-identisch bleiben, abgesichert durch SHA256-Fingerprint-Tests.

Auch die Framework-Grenze ist explizit: Microsoft Agent Framework, Semantic Kernel und MCP-Integrationen hängen nach innen vom Core ab, nicht umgekehrt. Direkte .NET-Nutzung kann dieselben Services ohne Agentenframework verwenden. Der Owner-Scope läuft durch MemoryScope; Anwendungsidentität und Store-Routing sind separate Host-Belange. MemoryIsolationMode verwendet aus Gründen der Rückwärtskompatibilität standardmässig SingleTenant; WarnOnUnscoped und der geschlossen fehlschlagende Modus StrictMultiTenant sind opt-in.

Architekturübersicht · Integrations- und Scoping-Leitfaden

Darum fühlte sich die Projektionsidee für mich weniger spekulativ an. Die Implementierung unterscheidet bereits Zusammenstellung, Projektion und Rendering. Meine Frage ist, was geschieht, wenn dieselbe Trennung nicht nur für Gedächtnis, sondern auch für Fähigkeit, Richtlinien und gelerntes Verhalten zum Architekturprinzip wird.

Mit main am 28. August 2026 abgeglichen.

Diese Trennung existiert heute für die Gedächtnisprojektion. Dasselbe Muster auf gemeinsame Fähigkeiten, Rückschreiben und gesteuertes Systemlernen auszudehnen, ist meine Extrapolation.

Meine Frage lautet also nicht: „Warum hast du es so gebaut?“ Ein grosser Teil der Mechanik ist bereits da. Die Laufzeitintegration ist auf den Agentenlebenszyklus ausgerichtet, doch die zugrunde liegende Architektur sieht zunehmend wie ein wiederverwendbares Substrat aus.

Was passiert, wenn wir die Projektionsidee einen Schritt weiterdenken: vom Projizieren von Gedächtnis in eine Laufzeit zum Projizieren einer breiteren gemeinsamen Fähigkeit in unterschiedliche Akteure, Schnittstellen und menschliche Ansichten – während Ausführungen gesteuerte Änderungen zurück vorschlagen können?

Projektion ist nur die halbe Geschichte

Bei der Delta-Analogie geht es nicht nur darum, eine materialisierte Ansicht zu lesen. Eine Entwicklerin bearbeitet den Checkout, und Änderungen können in den tieferen gemeinsamen Zustand zurückfliessen.

Dieser Rückweg ist für adaptive KI-Architektur wichtig. Ein Akteur kann mit einer Laufzeitprojektion arbeiten und Evidenz erzeugen, ohne die kanonische Fähigkeit direkt verändern zu dürfen.

A. Flüchtig

Ein Akteur verändert oder lernt lokal etwas. Nichts bleibt bestehen. Die nächste Materialisierung löscht es.

Projektion
    ↓
Akteur
    ↓
Notizblock
    ↓
verschwunden

B. Vorschlag

Das ist die interessanteste Stufe. Der Akteur entdeckt einen Werkzeugfehler, ein besseres Verfahren, einen Widerspruch, eine fehlende Bedingung oder eine lokale Ausnahme. Er schreibt nicht direkt fest.

Erfahrung
    ↓
Vorschlag
    ↓
Evaluation
    ↓
übernehmen / ablehnen
    ↓
kanonische Fähigkeit

Hier wurde AgentEval für mein Gedankenexperiment wichtig. José verwendet es heute als Evaluationsschicht für die Entwicklung, nicht als Promotion-Gate in der Produktion. Doch dieselbe evidenzorientierte Denkweise legt eine mögliche Architektur nahe: Eine Ausführung erzeugt einen Vorschlag, die Evaluation testet ihn, und erst danach kann eine Änderung zur gemeinsamen Fähigkeit werden. Der obige Kreislauf ist eine vorgeschlagene Architektur, nicht das heutige Laufzeitverhalten von AgentEval.

C. Gesteuerte Freigabe

In regulierten Umgebungen reicht Evaluation womöglich nicht aus. Eine Kandidatenfähigkeit kann eine menschliche oder richtlinienbasierte Genehmigung brauchen, bevor sie zum gemeinsamen Zustand wird.

Vorgeschlagene Systemarchitektur – kein heutiges Verhalten von AgentMemory oder AgentEval.

Kandidatenfähigkeit
    ↓
Evaluation
    ↓
menschliche / richtlinienbasierte Genehmigung
    ↓
genehmigte Fähigkeit
    ↓
neue Laufzeitprojektionen

03 · Projektion und gesteuertes Rückschreiben
Kanonisches Substrat

genehmigte Fähigkeit · Evidenz · Provenienz · Richtlinien · Historie

↓

Laufzeitprojektion

relevanter Zustand · Skills · Werkzeuge · Berechtigungen · Anweisungen

↓

Akteur

interpretieren · schlussfolgern · entscheiden · synthetisieren

↓

Erfahrung

Ergebnisse · Fehler · Entdeckungen · Ausnahmen

↓

Vorschlag

Vorschlag statt direkter Mutation

↓

Evaluation / Governance

übernehmen · ablehnen · Evidenz verlangen

↺ Kanonisches Substrat

Warum Rückschreiben wichtig ist

Erfassen am Ort der Entdeckung

Der Akteur ist anwesend, wenn ein Fehler passiert. Ohne Rückweg muss ein Mensch die Entdeckung bemerken, verstehen und später in ein anderes System übertragen. Zwischen diesen Schritten geht viel Signal verloren.

Negatives Wissen

Systeme erzeugen enorme Mengen an Evidenz darüber, was nicht funktioniert hat. Runbooks enthalten sie selten. Organisationen lernen immer wieder neu, dass ein Werkzeug unter einer bestimmten Bedingung versagt.

Provenienz

Eine Fähigkeitsänderung sollte auf die Ausführungen zurückverweisen, die sie motiviert haben. Sonst können wir nicht beantworten: Warum existiert diese Anweisung? Gilt die Evidenz noch?

Reparaturgeschwindigkeit

Eine Richtlinie ändert sich. Fünf kopierte Anweisungen veralten. Vorschlag, Evaluation und Promotion können eine kanonische Fähigkeit aktualisieren; die Projektionen werden von dort erneuert.

Lokale Varianten ohne Fork

Eine Abweichung kann zusammen mit der Bedingung gespeichert werden, unter der sie gilt. Nicht jede lokale Ausnahme muss zu einem separat kopierten Skill werden.

Was, wenn der Kontext selbst virtualisiert wird?

Statt zu sagen, der Agent habe Gedächtnis, Skills und Werkzeuge, könnten wir sagen: Das System materialisiert für diese Ausführung eine Laufzeitwelt.

04 · Kontextvirtualisierung
Kanonisches Substrat

Welt · Erfahrung · Fähigkeiten · Provenienz · Evaluation · Richtlinien · Berechtigungen · Historie

↓

Laufzeitprojektion

relevanter Zustand · Historie · Skills · Werkzeuge · Berechtigungen · Anweisungen

↓

Akteur

interpretieren · schlussfolgern · entscheiden · synthetisieren

↓

Erfahrung / Vorschlag

Evidenz und mögliche Änderungen

↓

Evaluation / Governance

entscheidet, was zu gemeinsamem Lernen wird

↺ Kanonisches Substrat

Der Agent erlebt diese Laufzeitwelt als seinen Kontext.

Aus der Ausführung heraus mag es wie „mein Gedächtnis, meine Werkzeuge, meine Skills, mein Kontext“ aussehen. Architektonisch können dies Projektionen eines tieferen Systemzustands sein. Zurück fliesst keine ungeprüfte Umschreibung dieses Zustands, sondern Evidenz und Vorschläge für eine gesteuerte Übernahme.

Vom akteurslokalen Lernen zum Systemlernen

Josés Gedächtnis- und Evaluationsarbeit liess mich einen akteursbezogenen Kreislauf wie diesen vorstellen:

Akteur
    ↓
Erfahrung
    ↓
Gedächtnis
    ↓
Evaluation
    ↓
verbesserte Fähigkeit
    ↺

AgentEval wird heute vor allem als Evaluationskreislauf in der Entwicklung eingesetzt. Der Schritt von dort zum kontinuierlichen Systemlernen ist die architektonische Spekulation, die ich hier mache.

Eine Ebene höher skaliert der Kreislauf anders:

05 · Vom Lernen des Akteurs zum Systemlernen
Akteurslokaler Kreislauf

Akteur
↓
Erfahrung
↓
Gedächtnis
↓
Evaluation
↓
Verbesserte Fähigkeit
↺

Gemeinsamer Systemkreislauf

Erfahrungen aller Akteure
↓
Gemeinsame Evaluation
↓
Gemeinsames Fähigkeitssubstrat
↓
Begrenzte Projektionen
↓
Akteur A / B / C
↺

Wenn Akteur A entdeckt, dass ein Werkzeug unter einer bestimmten Bedingung unzuverlässig ist, sollte Akteur B das nicht unabhängig neu entdecken müssen. Wenn ein Workflow wiederholt scheitert, weil sich eine Richtlinie geändert hat, sollte nicht jede Assistenz ihre eigene veraltete Version tragen, bis jemand sie repariert.

Vielleicht besteht der Skalierungsschritt nicht in mehr Agenten. Vielleicht besteht er darin, Lernen aus dem lokalen Gedächtnis eines Akteurs in gemeinsames Systemlernen zu verschieben.

Drei Dinge, die das schwierig machen

01
Berechtigungsgebundenes Lernen

Akteur A lernt aus Informationen, die Akteur B nicht sehen darf. Darf die abgeleitete Lektion geteilt werden? Provenienz und Klassifizierung müssen womöglich mit der gelernten Fähigkeit reisen.

02
Widersprüchliche Lektionen

Akteur A lernt, dass X funktioniert. Akteur B lernt, dass X scheitert. Für semantischen Widerspruch gibt es keinen Compiler.

03
Fehlende gemeinsame Erfolgsmetrik

Systemlernen setzt voraus, dass wir wissen, was „funktioniert“ bedeutet. Schneller? Sicherer? Günstiger? Regelkonformer? Besser für Nutzende? Organisationen haben oft keine gemeinsame Metrik.

Systemlernen braucht eine Definition des Guten. Organisationen sind darin oft viel weniger explizit, als Softwarearchitekturen annehmen.

Ist das nicht einfach wieder RAG?

RAG ist primär eine Lesezeitarchitektur.

RAG · Lesezeit
Dokumente
↓
Abruf
↓
LLM
↓
Antwort

Adaptives Substrat · geschlossener Kreislauf
Fähigkeit
↓
Projektion
↓
Ausführung
↓
Evidenz
↓
Vorschlag
↓
Evaluation
↓
Fähigkeit
↺

RAG kann weiterhin Teil des Substrats sein. Das ist kein „besseres RAG“. Der Unterschied ist der geschlossene Lernkreislauf.

RAG externalisiert Wissen. Die interessante Veränderung beginnt, wenn Ausführung gesteuerte Änderungen ins System zurückbringen kann.

Die einfache Version des Schritts: Basel

Für Berechnungen in meinem Basler Raumgraphen akzeptiere ich diese Architektur bereits: Das LLM soll keine Routen, Reisezeiten oder ÖV-Beziehungen schätzen. Das deterministische räumliche und zeitliche System besitzt diese Antworten. Das Modell extrahiert Absicht und Einschränkungen und erklärt danach ein berechnetes Resultat.

Zusammen liessen mich diese Architekturen fragen, ob derselbe Schritt für gelerntes Verhalten gilt: nicht nur Wahrheit aus dem Modell, sondern Fähigkeit und Lernen aus dem Akteur zu verschieben.

Schluss mit sechs Wahrheiten

Organisationen kopieren einen realen Prozess routinemässig in Wiki, Runbook, Checkliste, Schulung, KI-Anweisungen und Workflow-Konfiguration. Jede Repräsentation wird zu einer separat gepflegten Wahrheit. Jede driftet.

EIN PROZESS
    ↓
Wiki · Runbook · Checkliste · Schulung · KI-Anweisungen · Workflow-Konfiguration
    ↓
sechs gepflegte Wahrheiten
    ↓
Drift

Was, wenn das nicht sechs gepflegte Wahrheiten sind, sondern sechs Projektionen einer gesteuerten Fähigkeit?

06 · Eine Fähigkeit, viele Projektionen
Lebendiges Fähigkeitssubstrat

Verfahren · Evidenz · Richtlinien · Varianten · Evaluation · Provenienz

↓

Materialisierte Ansichten

Wiki · menschliche Checkliste · KI-Laufzeitskill · Review · Workflow / ausführbare Repräsentation

Schluss mit sechs Wahrheiten.

Wo könnte das tatsächlich relevant sein?

Das Modell wird dort interessant, wo eine veränderliche Fähigkeit vielen Konsumenten, Rollen oder Laufzeiten dienen muss. Nicht jeder Chatbot braucht das. Ein einfaches Frage-Antwort-System für Dokumente wahrscheinlich nicht.

Incident Response Eine sich entwickelnde operative Fähigkeit, viele Laufzeitansichten.

Runbooks, Wikis, Postmortems und tatsächliches Verhalten driften schnell auseinander. Ein gemeinsames Substrat könnte erfolgreiche und gescheiterte Abläufe, Werkzeugverhalten, Eskalationen und Erkenntnisse sammeln.

gemeinsame Incident-Fähigkeit
→ Bereitschaftsdienst
→ Einsatzleitung
→ KI-Laufzeit
→ Audit / Schulung

Der Wert ist nicht „besserer Chat“, sondern eine operative Fähigkeit für mehrere Rollen.

Regulierte Abläufe Richtlinien, Provenienz, Berechtigungen und rollenspezifische Projektionen.

Banken, Versicherungen und Pharma verbinden veränderliche Richtlinien, Berechtigungen, Evidenz und Auditierbarkeit. Unterschiedliche Akteure brauchen unterschiedliche Darstellungen desselben Prozesses.

Richtlinie + Verfahren + Provenienz + Evaluation
→ Analyse
→ Review
→ Compliance
→ Audit-Evidenz

Organisatorisches Prozesswissen Ein realer Prozess, viele menschliche und maschinelle Darstellungen.

Wie stellen wir hier tatsächlich jemanden ein? Der reale Prozess enthält formale Schritte, Systeme, Ausnahmen, lokale Praxis und wiederkehrende Fehler.

ein lebendiger Prozess
→ Checkliste
→ HR-Workflow
→ Onboarding
→ KI-Unterstützung
→ Governance

Softwareentwicklung Fähigkeit als Infrastruktur statt duplizierter Dokumentation.

Eine Deployment-Fähigkeit lebt selten nur in DEPLOYMENT.md. Sie verteilt sich über Repository, CI, Tests, fehlgeschlagene Deployments, Freigaben und Umgebungsbedingungen.

Deployment-Substrat
→ SKILL.md
→ Dokumentation
→ CI-Validierung
→ Codex / Copilot
→ visueller Workflow

Kundensupport Gemeinsames Lernen aus Lösungen, Fehlern und Produktänderungen.

Supportwissen verändert sich ständig. Lernen auf Systemebene ist attraktiver, als jeder Assistenz dieselbe Lektion einzeln beizubringen.

alle Supporterfahrungen
→ Evaluation
→ gemeinsame Lösungsfähigkeit
→ Kundenassistenz
→ Frontline-Support
→ Spezialisten / QA

KI-Adoption im grossen Massstab Gemeinsame Richtlinien und Fähigkeiten unter vielen Copilots.

Wenn Organisationen immer mehr Copilots, Assistenzen, Prompts und Workflows ansammeln, wird die eigentliche Skalierungsfrage womöglich, was sie alle wissen und tun dürfen.

genehmigte Werkzeuge + Richtlinien + Datenklassifizierung
Workflows + Vorfälle + Evaluation
→ HR-Assistenz
→ Forschungsassistenz
→ M365-Assistenz
→ Mitarbeitendenhilfe
→ Governance

Die Eigentumsfrage

Aus der Laufzeit heraus:

Ich erinnere mich.
Ich habe diese Skills.
Ich kann diese Werkzeuge verwenden.

Von aussen:

Das System hat relevanten Zustand,
Fähigkeiten und Werkzeuge
in diesen Ausführungskontext projiziert.

Vielleicht sind Agentengedächtnis und adaptive System-Middleware dieselbe Architektur, von entgegengesetzten Seiten der Laufzeitgrenze betrachtet.

Die Frage, die ich José stellen möchte, ist nun die kurze:

Offene Frage
Ab wann wird Agentengedächtnis zur Anwendungsarchitektur?

Ich möchte keine intelligenten kleinen Menschen bauen, an denen Software hängt. Ich möchte gute Softwaresysteme bauen, in denen probabilistisches Schlussfolgern eine Fähigkeit unter anderen ist.

Ausgangspunkte

José Luis Latorre — LinkedIn
AgentMemory for .NET — GitHub
AgentEval — GitHub
Neo4j NAMS / Agent Skills
Delta]]></content:encoded>
    </item>
<item>
      <title>A note from the other side of the same frontier</title>
      <link>https://bridge-work.ai/en/blog/a-note-from-the-other-side/</link>
      <guid isPermaLink="true">https://bridge-work.ai/en/blog/a-note-from-the-other-side/</guid>
      <description>Generative UI changes the economics of task-specific interfaces. Structure makes model power controllable, while governance belongs at the attachment between context, model and interface.</description>
      <pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate>
      <dcterms:modified>2026-09-02T00:00:00.000Z</dcterms:modified>
      <dc:language>en</dc:language>
      <category>Generative UI · Governance</category><category>Generative UI</category><category>MCP Apps</category><category>Context Governance</category>
      <content:encoded><![CDATA[TL;DR — Generative UI changes the economics of task-specific interfaces. Structure makes model power controllable. Context should stay outside any one model, while governance belongs at the attachment between context, model and interface.

Archive update · September 2026
Ruben Casas’ talk remains a useful source, and MCP Apps have since become an official MCP extension with a stable 2026-01-26 specification. Host support still varies, so “MCP Apps provide the delivery layer” should not be read as universal deployment support.

This started as a private letter to Ruben Casas after his talk Beyond Components — Designing Generative UI for MCP Apps.

The talk gave me language for something I had been approaching from a different direction. Ruben came from protocols and rendering. I came from adoption, context and governance.

Three thoughts stayed with me.

Argument map · From generative interface to governed attachment
01 · Static components

The host selects a pre-built component.
Data and properties change; the available surface remains fixed.
host application → pre-built component

02 · Declarative UI

The model produces a structured descriptor.
The host validates and renders a bounded component, properties and layout.
model → descriptor → host-rendered UI

03 · Generative UI

The task receives a task-shaped surface.
Context informs a temporary interface whose interaction can shape the next turn.
task + context → surface → interaction

04 · Separate the layers

Interface, model, context and governance have different jobs.
Replaceability and portability are goals that depend on keeping these concerns explicit.
interface · model · context · governance

05 · Govern the attachment

Policy controls what connects.
Consent, permissions and audit mediate how context reaches a model, tool and task interface.
portable context → governed attachment → model + UI

06 · Collaborative canvas

Conversation becomes a working surface.
The user changes the temporary interface, and that interaction becomes input to the next model turn.
conversation → canvas → interaction ↺ next turn

01 — Specificity economics, not screen evolution

The temptation with generative UI is to imagine that the revolution must be a new kind of screen: spatial canvases, floating windows, interfaces we have not yet named.

That may happen. But I think the more immediate change is economic.

Photoshop, Premiere, VS Code, DAWs and Figma already taught us a durable design pattern: dense, task-shaped surfaces. The interface adapts to the work.

What generative UI changes is the cost of producing that specificity. Instead of shipping years of fixed panels and menus, a system can assemble a useful surface per task, user and moment.

The revolution may not be that the screen becomes alien. It may be that specificity becomes cheap.

02 — Structure turns power into usable control

Ruben called declarative UI a useful balance between flexibility and consistency. That point has aged well.

The broader pattern is visible across structured generation: schemas, constrained components, declarative descriptors, validation, tool contracts and action guards.

The goal is not to make the model weak. It is to give a powerful model a precisely shaped place to act.

model capability + structured contract → bounded action → inspectable result

A model can generate arbitrary HTML, but a production system often benefits when it can instead produce a constrained representation that the host renders and validates.

03 — Centralize context, not interface. Govern the attachment.

The rendering conversation often starts with “where does the UI run?” I think there is a prior question: what should be central, and what should remain at the edges?

Interface Distributed, task-specific, per tool and moment.

Model Reasoning and compute utility; ideally replaceable.

Context Accumulated knowledge and working state outside any single model.

Governance Rules for what context may attach to which model and interface.

I still like this architecture, but one sentence from the 2026 draft should be softened. “The model becomes a swappable utility” is a design goal, not a guaranteed property. Provider-specific tools, context semantics and model behaviour can make substitution expensive.

Context portability across tools and models also remains incomplete and fragmented. That is the architectural direction, not a claim that the required product layer is already solved.

MCP Apps make the attachment concrete

MCP Apps provide a standardized pattern for servers to declare interactive UI resources that supporting hosts can render in sandboxed iframes, with bidirectional communication through the host. The official specification and SDK repository marks version 2026-01-26 as stable.

That gives us a real attachment point between tool, UI and conversational context. Support still varies by host.

But the extension does not eliminate governance. It makes governance more specific:

Which context reaches the tool?
Which fragments may be displayed or sent onward?
What actions can the embedded UI invoke?
What does the host log or audit?
What happens when the host does not support the extension?

Conversation as canvas as conversation

The user-side moment that first made this intuitive for me was much simpler. I asked an LLM to draw a BPMN diagram of making a margherita pizza.

The interesting part was not that it could draw. The diagram became a temporary task surface inside the conversation: something I could point at, change and reason through.

conversation → task-shaped surface → human interaction → next model turn → changed surface

We are now much closer to that pattern being a standard product primitive than when I first had the thought.

The frontier from the adoption side

The protocol and rendering layers are becoming real. The unresolved institutional question is still the attachment.

Inside a regulated organisation, the demo rarely fails because the generated card is unattractive. It fails when nobody can answer who routed which context, under what authority, to which model and tool, with what audit trail.

Distribute the interface.
Keep context portable.
Govern the attachment.

Sources

Ruben Casas — Beyond Components: Designing Generative UI for MCP Apps
Model Context Protocol — MCP Apps overview
Official MCP Apps specification and SDK repository]]></content:encoded>
    </item>
<item>
      <title>Eine Notiz von der anderen Seite derselben Grenze</title>
      <link>https://bridge-work.ai/blog/notiz-von-der-anderen-seite/</link>
      <guid isPermaLink="true">https://bridge-work.ai/blog/notiz-von-der-anderen-seite/</guid>
      <description>Generative UI verändert die Ökonomie aufgabenspezifischer Schnittstellen. Struktur macht Modellleistung kontrollierbar; Governance gehört an die Verbindung zwischen Kontext, Modell und Schnittstelle.</description>
      <pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate>
      <dcterms:modified>2026-09-02T00:00:00.000Z</dcterms:modified>
      <dc:language>de</dc:language>
      <category>Generative UI · Governance</category><category>Generative UI</category><category>MCP Apps</category><category>Kontext-Governance</category>
      <content:encoded><![CDATA[Kurzfassung — Generative UI verändert die Ökonomie aufgabenspezifischer Schnittstellen. Struktur macht Modellleistung kontrollierbar. Kontext sollte ausserhalb eines einzelnen Modells bleiben; Governance gehört an die Verbindung zwischen Kontext, Modell und Schnittstelle.

Archiv-Update · September 2026
Ruben Casas’ Talk bleibt eine hilfreiche Quelle. MCP Apps sind inzwischen eine offizielle MCP-Erweiterung mit einer stabilen Spezifikation vom 26. Januar 2026. Die Unterstützung unterscheidet sich weiterhin je nach Host; „MCP Apps liefern die Auslieferungsschicht“ bedeutet deshalb keine universelle Verfügbarkeit.

Dieser Text begann als privater Brief an Ruben Casas nach seinem Talk Beyond Components — Designing Generative UI for MCP Apps.

Der Talk gab mir Sprache für etwas, dem ich mich aus einer anderen Richtung genähert hatte. Ruben kam von Protokollen und Rendering. Ich kam von Adoption, Kontext und Governance.

Drei Gedanken blieben hängen.

Argumentkarte · Von generativer UI zur gesteuerten Verbindung
01 · Statische Komponenten

Der Host wählt eine vorgefertigte Komponente.
Daten und Eigenschaften ändern sich; die verfügbare Oberfläche bleibt fest.
Host-Anwendung → vorgefertigte Komponente

02 · Deklarative UI

Das Modell erzeugt eine strukturierte Beschreibung.
Der Host validiert und rendert begrenzte Komponenten, Eigenschaften und Layouts.
Modell → Beschreibung → Host-gerenderte UI

03 · Generative UI

Die Aufgabe erhält eine aufgabenspezifische Oberfläche.
Kontext formt eine temporäre Schnittstelle, deren Interaktion den nächsten Schritt beeinflussen kann.
Aufgabe + Kontext → Oberfläche → Interaktion

04 · Schichten trennen

Schnittstelle, Modell, Kontext und Governance haben unterschiedliche Aufgaben.
Austauschbarkeit und Portabilität bleiben Ziele, die explizite Grenzen voraussetzen.
Schnittstelle · Modell · Kontext · Governance

05 · Die Verbindung steuern

Richtlinien kontrollieren, was verbunden wird.
Einwilligung, Berechtigungen und Audit vermitteln, wie Kontext Modell, Werkzeug und Oberfläche erreicht.
portabler Kontext → gesteuerte Verbindung → Modell + UI

06 · Kollaborativer Canvas

Konversation wird zur Arbeitsoberfläche.
Die Person verändert die temporäre Oberfläche; diese Interaktion wird Eingabe für den nächsten Modellschritt.
Konversation → Canvas → Interaktion ↺ nächster Schritt

01 — Ökonomie der Spezifität statt Bildschirmevolution

Bei generativer UI liegt die Versuchung nahe, die Revolution als völlig neuen Bildschirm zu denken: räumliche Canvases, schwebende Fenster, Schnittstellen ohne Namen.

Das mag kommen. Die unmittelbarere Veränderung ist für mich jedoch ökonomisch.

Photoshop, Premiere, VS Code, DAWs und Figma haben uns längst ein dauerhaftes Designmuster gezeigt: dichte, aufgabenspezifische Oberflächen. Die Schnittstelle passt sich der Arbeit an.

Generative UI verändert die Kosten dieser Spezifität. Statt über Jahre feste Panels und Menüs auszuliefern, kann ein System pro Aufgabe, Person und Moment eine passende Oberfläche zusammensetzen.

Die Revolution muss nicht darin liegen, dass der Bildschirm fremdartig wird. Vielleicht wird Spezifität einfach günstig.

02 — Struktur macht Leistung kontrollierbar

Ruben bezeichnete deklarative UI als nützliche Balance zwischen Flexibilität und Konsistenz. Dieser Punkt ist weiterhin stark.

Das breitere Muster zeigt sich überall in strukturierter Generierung: Schemata, begrenzte Komponenten, deklarative Beschreibungen, Validierung, Tool-Verträge und Aktionsgrenzen.

Das Ziel ist nicht, das Modell schwach zu machen. Es soll einen präzise geformten Ort erhalten, an dem es handeln kann.

Modellfähigkeit + strukturierter Vertrag → begrenzte Aktion → prüfbares Ergebnis

Ein Modell kann beliebiges HTML erzeugen. Ein Produktionssystem profitiert aber oft davon, wenn es stattdessen eine begrenzte Darstellung erzeugt, die der Host rendert und validiert.

03 — Kontext zentralisieren, nicht Schnittstellen. Die Verbindung steuern.

Das Rendering-Gespräch beginnt oft mit der Frage: „Wo läuft die UI?“ Davor liegt für mich eine andere: Was sollte zentral sein, und was sollte an den Rändern bleiben?

Schnittstelle Verteilt und aufgabenspezifisch, je Werkzeug und Moment.

Modell Reasoning und Rechenleistung; idealerweise austauschbar.

Kontext Gesammeltes Wissen und Arbeitszustand ausserhalb eines einzelnen Modells.

Governance Regeln dafür, welcher Kontext an welches Modell und welche Schnittstelle darf.

Ich mag diese Architektur weiterhin. Einen Satz aus dem Entwurf von 2026 würde ich aber abschwächen: „Das Modell wird zum austauschbaren Utility“ ist ein Designziel, keine garantierte Eigenschaft. Anbieterspezifische Werkzeuge, Kontextsemantik und Modellverhalten können einen Wechsel teuer machen.

Auch die Kontextportabilität über Werkzeuge und Modelle hinweg bleibt unvollständig und fragmentiert. Das ist die architektonische Richtung, keine Behauptung, dass die nötige Produktschicht bereits gelöst wäre.

MCP Apps machen die Verbindung konkret

MCP Apps bieten ein standardisiertes Muster, mit dem Server interaktive UI-Ressourcen deklarieren können. Unterstützende Hosts rendern sie in Sandbox-iFrames und vermitteln die bidirektionale Kommunikation. Das offizielle Spezifikations- und SDK-Repository kennzeichnet die Version 2026-01-26 als stabil.

Damit entsteht ein realer Verbindungspunkt zwischen Werkzeug, UI und Gesprächskontext. Die Unterstützung unterscheidet sich weiterhin je nach Host.

Die Erweiterung beseitigt Governance jedoch nicht. Sie macht sie konkreter:

Welcher Kontext erreicht das Werkzeug?
Welche Fragmente dürfen angezeigt oder weitergegeben werden?
Welche Aktionen darf die eingebettete UI auslösen?
Was protokolliert oder auditiert der Host?
Was geschieht, wenn ein Host die Erweiterung nicht unterstützt?

Konversation als Canvas als Konversation

Der Moment auf der Nutzerseite, der mir das zuerst verständlich machte, war viel einfacher. Ich bat ein LLM, ein BPMN-Diagramm für die Zubereitung einer Margherita-Pizza zu zeichnen.

Interessant war nicht, dass es zeichnen konnte. Das Diagramm wurde im Gespräch zu einer temporären Arbeitsoberfläche: etwas, auf das ich zeigen, das ich verändern und mit dem ich denken konnte.

Konversation → aufgabenspezifische Oberfläche → menschliche Interaktion → nächster Modellschritt → veränderte Oberfläche

Wir sind heute deutlich näher daran, dass dieses Muster zu einem Standardbaustein von Produkten wird.

Die Grenze von der Adoptionsseite aus

Protokoll- und Rendering-Schichten werden real. Die ungelöste institutionelle Frage bleibt die Verbindung.

In einer regulierten Organisation scheitert eine Demo selten daran, dass die erzeugte Karte nicht attraktiv genug ist. Sie scheitert, wenn niemand beantworten kann, wer welchen Kontext unter welcher Befugnis an welches Modell und Werkzeug geleitet hat – und mit welcher Auditspur.

Schnittstellen verteilen.
Kontext portabel halten.
Die Verbindung steuern.

Quellen

Ruben Casas — Beyond Components: Designing Generative UI for MCP Apps
Model Context Protocol — MCP Apps Überblick
Offizielles MCP Apps Spezifikations- und SDK-Repository]]></content:encoded>
    </item>
<item>
      <title>From Agents to AI Capabilities</title>
      <link>https://bridge-work.ai/en/blog/from-agents-to-ai-capabilities/</link>
      <guid isPermaLink="true">https://bridge-work.ai/en/blog/from-agents-to-ai-capabilities/</guid>
      <description>The agent metaphor is a useful adoption doorway, but real systems are better designed as coordinated capabilities with explicit controls and human judgment.</description>
      <pubDate>Sun, 07 Jun 2026 00:00:00 GMT</pubDate>
      <dcterms:modified>2026-09-02T00:00:00.000Z</dcterms:modified>
      <dc:language>en</dc:language>
      <category>AI Architecture · Adoption</category><category>AI Adoption</category><category>Capability Design</category><category>Agent Architecture</category>
      <content:encoded><![CDATA[TL;DR — The agent metaphor is a useful adoption doorway, but real systems are better designed as coordinated capabilities: reasoning, automation, data, interfaces, controls and human judgment.

The agent metaphor did something useful for AI adoption. It gave people a person-shaped doorway into an abstract technology.

An assistant, helper or digital colleague is easier to imagine than orchestration, tool calls, retrieval, policies and execution graphs.

People often need imagination before architecture.

Argument map · From agent metaphor to capability system
01 · The adoption doorway

“Agent” makes AI imaginable.
A person-shaped assistant gives people a familiar way into an abstract system.
agent → assistant · helper · colleague

02 · The Bond illusion

One clever agent appears to own everything.
Memory, tools, data and actions collapse into a compelling but misleading character.
one agent → memory + tools + data + actions

03 · Coordinated system

Different capabilities do different work.
Retrieval, reasoning, validation, action and human approval are coordinated by workflow and policy.
retrieve → reason → validate → act → approve

04 · Capabilities before personas

Design the system around the work.
Reasoning, automation, data, interface, control and human judgment become explicit design choices.
work → capabilities → boundaries

05 · Persona as projection

The visible assistant becomes optional.
The same capability system can appear as chat, a background workflow, inline UI or an approval surface.
capability system → task-appropriate interface

Why agents worked

The agent metaphor lowers abstraction. It gives people something familiar to talk to and can be an excellent adoption doorway.

But a doorway is not a floor plan. The metaphor can help a person enter the subject without deciding how the system should be built.

The Bond illusion

A lot of early AI enthusiasm had a James Bond imagination: one brilliant agent in the centre, fast and competent enough to handle everything.

That is useful as a story. It is often misleading as an architecture.

The more real the use case gets, the less it looks like one person.

Why Ocean’s Eleven is the better architecture metaphor

Real systems combine different strengths. One component retrieves. One reasons. One validates. One acts. Another applies deterministic rules. A human approves the consequential step.

data → retrieval → reasoning → validation → action → human / policy gate

The magic is not one genius. The magic is coordination.

This does not mean “multi-agent is always better.” Microsoft’s current architecture guidance recommends starting with a single-agent test for most use cases and moving to multiple agents only when real separation boundaries or demonstrated limitations justify the added coordination, state, cost and latency.

The chatbot trap

If every AI opportunity is framed as an agent, teams tend to overproduce visible assistants. Sometimes chat is exactly right. Often it is not.

The better intervention may be a hidden workflow, an inline decision aid, a task-specific form, a background classifier, a retrieval step or a human review surface.

A better starting question
Not “where can we put an agent?” but “what capability is missing from the work?”

Design capabilities before personas

Reasoning Where probabilistic interpretation actually adds value.

Automation Repeatable steps that should be deterministic.

Data What context must be retrieved and scoped.

Interface What the user needs to see or manipulate.

Control Policies, permissions, validation and audit.

Human judgment Where approval, interpretation or accountability remains necessary.

What this means for adoption

The agent metaphor can remain an excellent adoption doorway. It lowers abstraction and gives people something familiar to talk to.

But the framing should mature with the use case. A user-facing persona is a product decision. It should not silently determine where memory, tools, permissions or workflow state live.

This is the bridge to the later articles in the archive: once we stop treating the agent as the whole system, questions about task-shaped interfaces, context, projection and governance become easier to ask.

Practical design questions

What workflow are we trying to improve?
Which capability is actually missing?
Which part needs probabilistic language reasoning?
Which part should remain deterministic?
What data is needed, and under whose permissions?
Where does a human need to approve?
What should be monitored, evaluated or audited?
Does a person-like assistant genuinely help the user, or merely simplify our story?

Agents are a useful doorway.
Capabilities are the architecture.

Source

Microsoft Cloud Adoption Framework — choosing single-agent and multi-agent systems]]></content:encoded>
    </item>
<item>
      <title>Von Agenten zu KI-Fähigkeiten</title>
      <link>https://bridge-work.ai/blog/von-agenten-zu-ki-faehigkeiten/</link>
      <guid isPermaLink="true">https://bridge-work.ai/blog/von-agenten-zu-ki-faehigkeiten/</guid>
      <description>Die Agenten-Metapher ist eine nützliche Eingangstür. Reale Systeme sollten aber als koordinierte Fähigkeiten mit klaren Kontrollen und menschlichem Urteil gestaltet werden.</description>
      <pubDate>Sun, 07 Jun 2026 00:00:00 GMT</pubDate>
      <dcterms:modified>2026-09-02T00:00:00.000Z</dcterms:modified>
      <dc:language>de</dc:language>
      <category>KI-Architektur · Adoption</category><category>KI-Adoption</category><category>Fähigkeitsdesign</category><category>Agentenarchitektur</category>
      <content:encoded><![CDATA[Kurzfassung — Die Agenten-Metapher ist eine nützliche Eingangstür. Reale Systeme sollten aber als koordinierte Fähigkeiten gestaltet werden: Reasoning, Automatisierung, Daten, Schnittstellen, Kontrollen und menschliches Urteil.

Die Agenten-Metapher hat für die KI-Adoption etwas Wichtiges geleistet. Sie gab Menschen einen personenförmigen Zugang zu einer abstrakten Technologie.

Eine Assistenz, ein Helfer oder eine digitale Kollegin sind leichter vorstellbar als Orchestrierung, Tool-Aufrufe, Retrieval, Richtlinien und Ausführungsgraphen.

Menschen brauchen oft Vorstellungskraft vor Architektur.

Argumentkarte · Von der Agenten-Metapher zum Fähigkeitssystem
01 · Die Eingangstür

„Agent“ macht KI vorstellbar.
Eine personenförmige Assistenz gibt Menschen einen vertrauten Zugang zu einem abstrakten System.
Agent → Assistenz · Helfer · Kollegin

02 · Die Bond-Illusion

Ein cleverer Agent scheint alles zu besitzen.
Gedächtnis, Werkzeuge, Daten und Aktionen verschmelzen zu einer überzeugenden, aber irreführenden Figur.
ein Agent → Gedächtnis + Werkzeuge + Daten + Aktionen

03 · Koordiniertes System

Unterschiedliche Fähigkeiten erledigen unterschiedliche Arbeit.
Retrieval, Reasoning, Validierung, Aktion und menschliche Freigabe werden durch Workflow und Richtlinien koordiniert.
abrufen → schlussfolgern → validieren → handeln → freigeben

04 · Fähigkeiten vor Personas

Das System wird um die Arbeit herum gestaltet.
Reasoning, Automatisierung, Daten, Schnittstelle, Kontrolle und menschliches Urteil werden explizite Designentscheidungen.
Arbeit → Fähigkeiten → Grenzen

05 · Persona als Projektion

Die sichtbare Assistenz wird optional.
Dasselbe Fähigkeitssystem kann als Chat, Hintergrund-Workflow, Inline-UI oder Freigabeoberfläche erscheinen.
Fähigkeitssystem → aufgabengerechte Schnittstelle

Warum Agenten funktioniert haben

Die Agenten-Metapher senkt den Abstraktionsgrad. Sie gibt Menschen etwas Vertrautes, über das sie sprechen können, und kann eine hervorragende Eingangstür zur Adoption sein.

Aber eine Eingangstür ist kein Grundriss. Die Metapher kann den Einstieg erleichtern, ohne die Architektur des Systems festzulegen.

Die Bond-Illusion

Viel frühe KI-Begeisterung folgte einer James-Bond-Vorstellung: ein brillanter Agent im Zentrum, schnell und kompetent genug für alles.

Als Geschichte ist das nützlich. Als Architektur ist es oft irreführend.

Je realer der Anwendungsfall wird, desto weniger sieht er nach einer einzelnen Person aus.

Warum Ocean’s Eleven die bessere Architekturmetapher ist

Reale Systeme verbinden unterschiedliche Stärken. Eine Komponente ruft Daten ab. Eine schlussfolgert. Eine validiert. Eine handelt. Eine weitere wendet deterministische Regeln an. Ein Mensch gibt den folgenreichen Schritt frei.

Daten → Retrieval → Reasoning → Validierung → Aktion → Mensch / Richtlinien-Gate

Die Magie liegt nicht in einem Genie. Sie liegt in der Koordination.

Das bedeutet nicht, dass Multi-Agenten-Systeme grundsätzlich besser sind. Microsofts aktuelle Architekturleitlinie empfiehlt für die meisten Anwendungsfälle zuerst einen Test mit einem einzelnen Agenten. Mehrere Agenten sind dann sinnvoll, wenn echte Trennungsgrenzen oder nachgewiesene Einschränkungen den zusätzlichen Aufwand für Koordination, Zustand, Kosten und Latenz rechtfertigen.

Die Chatbot-Falle

Wenn jede KI-Chance als Agent gerahmt wird, produzieren Teams zu viele sichtbare Assistenzen. Manchmal ist Chat genau richtig. Oft ist er es nicht.

Die bessere Intervention kann ein unsichtbarer Workflow, eine Entscheidungshilfe im Prozess, ein aufgabenspezifisches Formular, ein Klassifikator im Hintergrund, ein Retrieval-Schritt oder eine menschliche Review-Oberfläche sein.

Eine bessere Ausgangsfrage
Nicht „Wo können wir einen Agenten einsetzen?“, sondern „Welche Fähigkeit fehlt in der Arbeit?“

Fähigkeiten vor Personas gestalten

Reasoning Wo probabilistische Interpretation tatsächlich Wert schafft.

Automatisierung Wiederholbare Schritte, die deterministisch sein sollten.

Daten Welcher Kontext abgerufen und begrenzt werden muss.

Schnittstelle Was Nutzende sehen oder bearbeiten müssen.

Kontrolle Richtlinien, Berechtigungen, Validierung und Audit.

Menschliches Urteil Wo Freigabe, Interpretation oder Verantwortung nötig bleiben.

Was das für die Adoption bedeutet

Die Agenten-Metapher kann weiterhin eine hervorragende Eingangstür sein. Sie senkt die Abstraktion und gibt Menschen etwas Vertrautes, über das sie sprechen können.

Doch mit dem Anwendungsfall muss auch die Rahmung reifen. Eine sichtbare Persona ist eine Produktentscheidung. Sie sollte nicht stillschweigend bestimmen, wo Gedächtnis, Werkzeuge, Berechtigungen oder Workflow-Zustand liegen.

Hier liegt die Brücke zu den späteren Texten im Archiv: Sobald wir den Agenten nicht mehr als das ganze System behandeln, lassen sich Fragen zu aufgabenspezifischen Schnittstellen, Kontext, Projektion und Governance leichter stellen.

Praktische Designfragen

Welchen Arbeitsablauf wollen wir verbessern?
Welche Fähigkeit fehlt tatsächlich?
Welcher Teil braucht probabilistisches Sprach-Reasoning?
Welcher Teil sollte deterministisch bleiben?
Welche Daten werden gebraucht und unter wessen Berechtigungen?
Wo muss ein Mensch freigeben?
Was sollte überwacht, evaluiert oder auditiert werden?
Hilft eine personenähnliche Assistenz den Nutzenden wirklich oder vereinfacht sie nur unsere Erzählung?

Agenten sind eine nützliche Eingangstür.
Fähigkeiten sind die Architektur.

Quelle

Microsoft Cloud Adoption Framework — Einzel- und Multi-Agenten-Systeme auswählen]]></content:encoded>
    </item>
<item>
      <title>Ausführung ist keine Schnittstelle</title>
      <link>https://bridge-work.ai/blog/ausfuehrung-ist-keine-schnittstelle/</link>
      <guid isPermaLink="true">https://bridge-work.ai/blog/ausfuehrung-ist-keine-schnittstelle/</guid>
      <description>KI kann immer mehr Softwarearbeit ausführen. Dadurch werden Lesbarkeit, Provenienz und Navigation wichtiger, nicht weniger wichtig.</description>
      <pubDate>Tue, 10 Mar 2026 00:00:00 GMT</pubDate>
      <dcterms:modified>2026-09-02T00:00:00.000Z</dcterms:modified>
      <dc:language>de</dc:language>
      <category>KI-Entwicklung · Schnittstellen</category><category>Developer Experience</category><category>Agentische Ausführung</category><category>Generative UI</category>
      <content:encoded><![CDATA[Kurzfassung — KI kann immer mehr Softwarearbeit ausführen. Das löst das menschliche Schnittstellenproblem nicht, sondern macht Lesbarkeit, Provenienz und Navigation wichtiger.

Archiv-Update · September 2026
Der Text, auf den dieser Artikel antwortet, ist real: Gwen Davis veröffentlichte am 10. März 2026 im GitHub Blog „The era of ‘AI as text’ is over. Execution is the new interface.“ Der folgende Text widerspricht weiterhin der Formulierung, nicht der zugrunde liegenden Entwicklung hin zu programmierbarer agentischer Ausführung.

In der Softwareentwicklung hat sich etwas Wesentliches verschoben. KI-Systeme können zunehmend planen, Werkzeuge aufrufen, Dateien verändern, Tests ausführen und Arbeit erledigen, statt nur Text zurückzugeben.

Aber die Formulierung execution is the new interface stört mich weiterhin.

Ausführung ist keine Schnittstelle. Ausführung ist das, was hinter einer Schnittstelle geschieht.

Wer beides verwechselt, verdeckt genau das Designproblem, das wichtiger wird, wenn Maschinen immer grössere Teile eines Systems für uns zusammensetzen: Wie können Menschen verstehen, navigieren und kontrollieren, was gebaut wurde?

Argumentkarte · Von Autorschaft zu Systemverständnis
01 · Manuelle Autorschaft

Die Entwicklerin baut das System.
Autorschaft schafft eine mentale Route durch Code und Entscheidungen.
Entwicklerin → Code → Ausführung

02 · Geteilte Autorschaft

KI beteiligt sich an der Konstruktion.
Die Verantwortung bleibt beim Menschen, während die direkte Vertrautheit mit der Umsetzung abnimmt.
Entwicklerin → KI-Konstruktion → System

03 · Ausführung ist nicht die Karte

Logs, Diffs, Terminal und Chat zeigen Aktivität.
Keines davon liefert allein ein zusammenhängendes mentales Modell des Systems.
Ausführungsevidenz → Systemverständnis

04 · Struktur als gemeinsame Oberfläche

Graphartige Artefakte machen Beziehungen sichtbar.
Knoten, Verbindungen und Verträge bieten Menschen und Maschinen eine Darstellung oberhalb von Rohtext.
Knoten + Verbindungen + Verträge → Karte

05 · Ein entstehendes Paradigma

Befehlsoberflächen und generative UI bestehen nebeneinander.
Die nächste Entwicklungsschnittstelle ist erkennbar, aber noch nicht gefestigt.
Terminal + Agenten-CLI + MCP Apps / GenUI

06 · Koordinierte Projektionen

Ein System, mehrere nützliche Ansichten.
Architektur, Ausführung, Provenienz und Narrativ verbinden sich zu einer mentalen Karte.
mehrere Systemansichten → Verständnis

Die eigentliche Verschiebung: Systeme nicht mehr nur schreiben, sondern navigieren

Über Jahrzehnte liess sich die dominante Programmierschleife einfach beschreiben:

Entwicklerin → Code → Ausführung

Niemand hatte buchstäblich jedes Detail im Kopf. Doch Autorschaft schuf ein kognitives Gerüst. Man erinnerte sich an die schwierige Funktion, die heikle Grenze und die Abkürzung, die man später noch korrigieren wollte.

KI-gestützte Entwicklung lockert diese Beziehung:

Entwicklerin → KI → erzeugtes / verändertes System → Ausführung

Die Verantwortung bleibt, während die Autorschaft partiell wird. Die Arbeit verschwindet nicht. Sie verändert ihre Form.

Der Schwerpunkt verschiebt sich vom Schreiben jedes einzelnen Systemteils hin zum Navigieren, Prüfen und Steuern eines Systems, dessen Konstruktion zunehmend mit Maschinen geteilt wird.

Das Sichtbarkeitsproblem

Wenn grosse Teile eines Systems fertig zusammengesetzt eintreffen, fehlt oft nicht der Code. Es fehlen Geschichte und Struktur, die den Code bewohnbar machen.

Struktur Was hängt wovon ab?

Absicht Warum existiert dieses Modul?

Provenienz Wer oder was hat es verändert – und warum?

Risiko Wo liegen die fragilen Grenzen?

Darum sind Dateibaum und Editor weiterhin nötig, aber nicht ausreichend. Sie zeigen die Implementierung. Sie zeigen nicht automatisch das System als mentales Modell.

Ich mag weiterhin den Vergleich mit einem langen Text. Wer ihn selbst schreibt, erinnert sich an den Weg durch das Argument. Wenn ein Modell den grössten Teil erzeugt, braucht es zusätzlich eine Karte, die zeigt, wie alles zusammenhängt.

Ausführung löst Infrastruktur, nicht Verständlichkeit

Agentische Entwicklung hat echte Fortschritte in der Infrastruktur gebracht. Systeme können Aufgaben planen, APIs aufrufen, Repositories verändern, Tests ausführen und sich von manchen Fehlern erholen.

Das ist wichtig. Aber es beantwortet die Schnittstellenfrage nicht.

Die Schnittstellenfrage
Wie kann ein Mensch das System auf jener Ebene prüfen, auf der er dafür Verantwortung übernehmen soll?

Logs sind nützlich. Diffs sind nützlich. Terminals sind nützlich. Chat ist nützlich. Keines davon ist für sich allein die Karte.

Das frühe Signal sind strukturierte Systeme

Werkzeuge wie n8n sind hier interessant, weil das Artefakt zugleich ausführbar und strukturell lesbar ist: Ein Workflow wird als Knoten, Parameter und Verbindungen dargestellt.

Die alte Behauptung würde ich heute vorsichtiger formulieren. Nicht: n8n „funktioniert gut mit KI, weil es Struktur offenlegt“. Das wäre zu kausal. Die nützlichere Beobachtung ist einfacher: Graphartig strukturierte Artefakte geben Menschen und Maschinen eine Darstellung oberhalb von Rohtext.

Strukturen verstehen wir über Karten.

Die Renaissance des Terminals – und was sich seit März verändert hat

Die Rückkehr befehlsorientierter Agentenwerkzeuge war folgerichtig. Terminals sind schnell, kombinierbar und gut instrumentierbar. Wenn KI sich unvorhersehbar verhält, bieten sie einen Ort zum Prüfen, Wiederholen und Eingreifen.

Einen Satz aus dem ursprünglichen Artikel würde ich heute ändern. Im März schrieb ich, das nächste Schnittstellenparadigma sei „noch nicht erfunden“. Im September ist das zu absolut.

MCP Apps sind heute eine offizielle Erweiterung für interaktive UI in MCP-Hosts. Gleichzeitig untersucht generative UI ausdrücklich aufgabenspezifische, dynamisch gerenderte Oberflächen. Das Paradigma ist nicht ausgeblieben. Es hat sich nur noch nicht gefestigt.

Was eine KI-native Entwicklungsschnittstelle zeigen könnte

Statt den Quellbaum als einzige Ansicht zu behandeln, könnten Entwicklungsumgebungen mehrere koordinierte Ansichten desselben Systems projizieren:

Architektur Module, Dienste, Verträge und Abhängigkeiten.

Ausführung Aufrufe, Werkzeuge, Traces und Zustandswechsel.

Provenienz Änderungen durch Menschen und Modelle, Begründungen, Diffs und Evidenz.

Narrativ Wie sich das System entwickelt hat und wo Entscheidungen entstanden.

Diese Ansichten ersetzen den Code nicht. Sie machen maschinell unterstützten Code navigierbar.

KI kann den Code erzeugen.
Doch Entwickelnde brauchen weiterhin die Karte.

Quellen

Gwen Davis, GitHub Blog — „Execution is the new interface“ (10. März 2026)
Model Context Protocol — MCP Apps Überblick]]></content:encoded>
    </item>
<item>
      <title>Execution is Not an Interface</title>
      <link>https://bridge-work.ai/en/blog/execution-is-not-an-interface/</link>
      <guid isPermaLink="true">https://bridge-work.ai/en/blog/execution-is-not-an-interface/</guid>
      <description>AI can increasingly execute software work. That makes system legibility, provenance and navigation more important, not less.</description>
      <pubDate>Tue, 10 Mar 2026 00:00:00 GMT</pubDate>
      <dcterms:modified>2026-09-02T00:00:00.000Z</dcterms:modified>
      <dc:language>en</dc:language>
      <category>AI Development · Interface</category><category>Developer Experience</category><category>Agentic Execution</category><category>Generative UI</category>
      <content:encoded><![CDATA[TL;DR — AI can increasingly execute software work. That does not remove the human interface problem; it makes system legibility, provenance and navigation more important.

Archive update · September 2026
The article this responds to is real: Gwen Davis published “The era of ‘AI as text’ is over. Execution is the new interface.” on the GitHub Blog on 10 March 2026. The argument below still disagrees with the phrase, not with the underlying move toward programmable agentic execution.

Something important has shifted in software development. AI systems can increasingly plan, call tools, change files, run tests and execute work instead of only returning text.

But the phrase execution is the new interface still bothers me.

Execution is not an interface. Execution is what happens behind one.

And confusing the two hides the design problem that becomes more important as machines assemble more of the system for us: how do humans understand, navigate and control what was built?

Argument map · From authorship to system legibility
01 · Manual authorship

The developer builds the system.
Authorship creates a mental route through the code and its decisions.
developer → code → execution

02 · Partial authorship

AI shares construction.
The developer stays responsible while direct familiarity with the implementation weakens.
developer → AI construction → system

03 · Execution is not the map

Logs, diffs, terminals and chat reveal activity.
None of them alone provides a coherent mental model of the system.
execution evidence → system legibility

04 · Structure as shared surface

Graph-shaped artifacts expose relationships.
Nodes, connections and contracts give humans and machines a representation above raw text.
nodes + connections + contracts → map

05 · An emerging paradigm

Command surfaces and generative UI now coexist.
The next development interface has begun to appear, but it has not settled.
terminal + agent CLI + MCP Apps / GenUI

06 · Coordinated projections

One system, several useful views.
Architecture, execution, provenance and narrative combine into a human mental map.
multiple system views → developer understanding

The real shift: from writing systems to navigating systems

For decades the dominant programming loop was easy to describe:

developer → code → execution

The developer did not literally hold every detail in their head, but authorship created cognitive scaffolding. You remembered the difficult function, the awkward boundary, the shortcut you promised yourself you would revisit.

AI-assisted development weakens that relationship:

developer → AI → generated / modified system → execution

The developer remains responsible, while authorship becomes partial. The job does not disappear. It changes shape.

The centre of gravity moves from writing every part of the system toward navigating, reviewing and steering a system whose construction is increasingly shared with machines.

The visibility problem

When large portions of a system arrive fully formed, the missing thing is often not code. It is the story and structure that make the code inhabitable.

Structure What depends on what?

Intent Why does this module exist?

Provenance Who or what changed it, and why?

Risk Where are the fragile boundaries?

This is why the familiar file tree and editor are necessary but insufficient. They show the implementation. They do not automatically show the system as a mental model.

The analogy I still like is writing a long paper. When you write it yourself, you remember the route through the argument. If a model produces most of it, you need an additional map to understand how the argument hangs together.

Execution solves infrastructure, not legibility

Agentic development has made genuine infrastructure progress. Systems can plan tasks, invoke APIs, manipulate repositories, run tests and recover from some failures.

That matters. But it does not answer the interface question.

The interface question
How can a human inspect the system at the level at which they are now expected to take responsibility for it?

Logs are useful. Diffs are useful. Terminals are useful. Chat is useful. None of them, alone, is the map.

The early signal is structured systems

Tools such as n8n are interesting here because the artifact is already both executable and structurally legible: a workflow is represented as nodes, parameters and connections.

I would now state the old claim more carefully. It is not that n8n “works with AI because it exposes structure.” That is too causal. The useful observation is simpler: graph-structured artifacts give both humans and machines a representation above raw text.

Structures are understood through maps.

The terminal renaissance — and what changed since March

The return of command-driven agent tools made sense. Terminals are fast, composable and easy to instrument. When AI behaves unpredictably, a terminal gives developers a place to inspect, retry and intervene.

But I would change one sentence from the original article. In March I wrote that the next interface paradigm “has not been invented yet.” By September that is too absolute.

MCP Apps are now an official extension for interactive UI inside MCP hosts, and generative-UI work is explicitly exploring task-shaped, dynamically rendered surfaces. The paradigm has not failed to appear. It has not settled.

What an AI-native development interface might expose

Instead of treating the source tree as the only view, development environments can project several coordinated views over the same system:

Architecture Modules, services, contracts and dependencies.

Execution Calls, tools, traces and state transitions.

Provenance Human and model changes, rationale, diffs and evidence.

Narrative How the system evolved and where decisions accumulated.

These do not replace code. They make machine-assisted code navigable.

AI may generate the code.
But developers still need the map.

Sources

Gwen Davis, GitHub Blog — “Execution is the new interface” (10 March 2026)
Model Context Protocol — MCP Apps overview]]></content:encoded>
    </item>
  </channel>
</rss>
