Those of us who spend any time thinking about AI agents tend to think about them as loops around a language model. Something wakes the agent, like a message. The agent sends that message, plus whatever context it has assembled, to a model. The model returns. Maybe it requests a tool, and the tool executes, and the result goes back, and so on. The loop continues until the model decides it is finished. Then the process terminates. Between that and the next wake, nothing is running.

That is fine for an isolated task. Agent frameworks like OpenClaw and its variants have introduced schedulers and timers to awaken the agent periodically to poll an inbox, check a calendar, or run a batch job. But these remain low-frequency wake-and-loop iterations, sleeping until the next timer tick. And they have to be. Waking a frontier model every second to perform reasoning would be financially ruinous - especially when, ninety-nine times out of a hundred, there is nothing to think about.

As agents take on continuous cognitive processing and embed into daily life, the wake cycle must become a high-frequency, high-bandwidth activity. An agent at a human’s side requires audio, visual, and OS input at a high refresh rate, even when that input is mostly background noise. Frontier language models are structurally ill-suited to participate in cognition at high refresh rates.

Which raises an architectural question: what does continuous cognition actually look like?

The Robotics Split

Kahneman’s labels for this divide are System 1 and System 2, from Thinking, Fast and Slow. System 1 is fast, automatic, and cheap enough to run continuously. System 2 is slow, deliberative, and expensive enough that it must remain occasional. Brains are not software, and the two-system framing is an abstraction, but the core lesson is about frequency: a high refresh rate and a frontier model do not belong on the same clock.

That is a fundamentally different claim from “use a smaller model in the same loop.” The loop we have in almost all AI agents is inherently a System 2 loop. Speeding it up does not yield System 1. System 1 must be an entirely different substrate running while the language model is asleep, deciding whether that loop should wake at all.

Robotics has understood this from the beginning. A bipedal robot cannot wait for a deliberative planner to finish a chain-of-thought before it balances. The field did not spend a decade pretending otherwise. Reactive control, deliberative planning, and an intermediate translation layer are standard robotics architecture, stretching from hybrid deliberative-reactive systems back to Rodney Brooks’ subsumption architecture. Modern foundation model robotics papers are simply new implementations of an old split:

  • Helix 02 is the cleanest recent statement. Each subsystem operates at its natural timescale: System 2 reasons over scenes, language, and high-level goals; System 1 translates those goals into full-body joint targets at 200 Hz; System 0 stabilizes balance and contact dynamics at 1 kHz. System 2 never chooses finger trajectories a thousand times a second, and System 0 does not need to know why the dishwasher is being unloaded. Helix 2.5 preserved that identical stack while generalizing zero-shot across unfamiliar homes. The hierarchy was the prerequisite for the experiment, not the experiment itself.
  • GR00T N1 splits along the same boundary: a vision-language planner operating over a diffusion action model.
  • Diffusion-VLA couples an autoregressive reasoner to a diffusion policy, driving the action policy at 82 Hz on a single workstation A6000.
  • Gemini Robotics places multi-step embodied planning in one model and motor translation in another, deploying an on-device action model for the tight feedback loops that cannot tolerate network round trips.

In none of these systems is the robot treated as “the model.” The model is one subsystem among many.

Personal software agents, like Muse or Dot, have resisted this split. They remain almost entirely single-clock System 2 thinkers. A robot that falls over makes a single-clock architecture untenable immediately; an agent checking an email inbox does not.

Three Clocks

Software agents got away with one clock because a chat interface is forgiving. A personal agent wired to a microphone, video feeds, and an active desktop sits much closer to the robotics constraint than the chat constraint. It simply operates with slightly softer real-time deadlines.

Tokens cost money. A cloud round-trip is an explicit data exposure. Someone asking aloud where their phone went will not wait for a 4-second reasoning trace. And an open microphone running for six hours produces almost nothing a frontier model ought to hear.

What must watch the stream is not a prompt loop, but a mesh:

SYSTEM 0   Reflexes, interlocks, window focus, VAD, scene changes (10–100 Hz)
SYSTEM 1   Perception, situation tracking, association, prediction, fast judgment (1–10 Hz)
SYSTEM 2   The deliberative loop: planning, deep retrieval, tool orchestration (On Demand)
EFFECTORS  Browser, terminal, APIs, audio synthesis, external agents

In practice, this resolves into three distinct clocks:

1. The Fast Clock (Reflex & Ingestion)

The fastest layer should barely qualify as cognition. Voice activity detection (VAD), rough scene-change deltas, location pings, operating system signals (VS Code coming to the foreground, a terminal buffer receiving new output), and safety interlocks. This layer is intentionally simple. Its purpose is to discard 99% of the stream at a refresh rate a language model should never encounter.

2. The Middle Clock (The Perceptual Mesh)

This is the System 1 layer that personal agency actually hinges on. Speech recognition, speaker embeddings, object trackers, OCR on active window regions, and small local models compressing raw perceptual deltas into state:

  • Who is in the room?
  • What are they doing?
  • Was that muttered sentence addressed to the agent or to the room?
  • Is this repeated build failure the same failure as two minutes ago?

This layer will frequently make noisy inferences. It must be cheap enough that being wrong carries zero marginal cost, because it runs continuously while the frontier model sleeps.

3. The Slow Clock (The Deliberative Loop)

This is the loop we already know how to build: a frontier model, an explicit objective, tools, and a broad context window. It executes only when the middle clock produces an actionable trigger, although not because a frame arrived, or a generic timer ticked. It does not own the raw sensors, nor does it maintain the event log. Waking it is an explicit architectural decision made by the middle layer.

Judgments Without Essays

For the middle clock to gate the slow clock, it must be able to evaluate closed questions without generating prose:

unusual?
relevant to an active goal?
addressed to the agent?
assistance welcome?
promote to memory?
wake the slow model?

Some of these evaluations are deterministic code: a known face ID, a thresholded delta, a regex over a terminal error, or a failed prediction. But the evaluations that require nuanced judgment are where dedicated System 1 models belong.

In September, TypeSafe released Jev and framed it explicitly as a System 1 model. This is arguably the first time a lab has treated Kahneman’s label as an architectural product category. Jev does not generate freeform tokens. You pass it a structured state vector alongside a set of typed questions, and it returns classification probabilities, choices, or categorical confidence scores in a single forward pass (quoting 70 to 500 milliseconds). Output generation costs effectively nothing because there is no autoregressive token decoding to meter. It cannot answer an open-ended question, and it should never be asked to. The questions required to manage an agent’s attention are almost entirely closed.

Consider two events in the same room:

Case 1: Routine Retrieval
An operator mutters that they cannot find their phone. Speech recognition transcribes the sentence; speaker identification matches the voice; an associative vector lookup recalls an object detection coordinate on the kitchen counter from twenty minutes prior. The judgment model scores the interruption threshold as low-cost and welcome. The agent speaks without ever touching a frontier model: “I think you left it on the kitchen counter.”

Case 2: Deliberative Intervention
The operator remarks that the consolidation path from situation to episodic memory is malfunctioning. Speech recognition flags the utterance, but this time the context triggers different scores: the sentence concerns the agent’s internal architecture, it follows a sequence of failed unit tests, and it cannot be resolved with a cached fact.

The slow subsystem wakes to a structured handoff brief:

reason: operator challenging memory consolidation architecture
situation: debugging pAI cognitive architecture
evidence: utterance followed three consecutive episodic-memory test failures
memories: current consolidation design, provenance graph schema
local hypothesis: transient situation state is being promoted too eagerly
confidence: 0.76

The frontier model does not ingest the sensor stream; it receives an executive handoff synthesized by the subsystems that did.

The Situation Model and Memory Promotion

The artifact that enables this handoff is the situation model.

A situation model is not long-term memory. It is a short-lived, transient representation of present reality, maintained whether or not the slow loop is awake:

{
  "people": {
    "dagny": {
      "present": true,
      "activity": "coding",
      "attention": "terminal"
    }
  },
  "environment": {
    "location": "office",
    "other_people_present": false
  },
  "task": {
    "inferred": "debugging pAI",
    "confidence": 0.84
  },
  "recent_events": ["build_failed", "build_failed", "build_failed"],
  "predictions": [
    { "event": "inspect_build_logs", "probability": 0.72 }
  ]
}

A user looking at a terminal at 09:42 does not belong in permanent episodic storage. Three consecutive identical build failures might. The eventual discovery that a specific dependency pin was broken definitely does, and it must carry provenance back to the event that demonstrated it.

observation ──▶ situation ──▶ attention ──▶ experience ──▶ memory

Storing everything is a trivial failure mode that poisons retrieval. Selective promotion is the real architectural challenge. In my own agent work, memory is modeled as a projection over an underlying event log, maintaining strict distinctions between what was directly observed, what was inferred, and what was inherited from prior belief. That model remains sound only if the vast majority of raw sensory input expires inside the situation model before it ever hardens into an event. Otherwise, your knowledge graph degenerates into a bloated, noisy transcript.

A small local multimodal model (such as a quantized 9B parameter model) is more than capable of keeping this situation JSON updated. But even that model should not parse raw video frames. Simpler components must identify what changed: YOLO, window-focus hooks, accessibility trees, or lightweight frame diffs. The operating system’s raw state is already structured; turning structured OS events into unstructured English prose just so a massive LLM can re-extract the structure is computational theater.

This insight is validated by recent research in Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?. When user activity is structured as a stream of (actor, verb, object, timestamp), a compact temporal graph model can determine whether to trigger an agent wake, and identify which entities are relevant in 14 milliseconds on a laptop, using roughly 220 MiB of RAM. That is 12 to 83 times faster than delegating the wake decision to an LLM. The language model should craft the response after the decision is made; it should not be the trigger.

Similar separations are emerging across the literature:

  • Talker-Reasoner separates a low-latency conversational agent from a slower background reasoner.
  • System-1.x balances fast heuristic inference against deliberate search, modulated by task complexity.
  • Cluster, Route, Escalate preserves 97% to 99% of frontier model accuracy while routing routine queries to cheap, fast classifiers.

Association, Prediction, and the Privacy Boundary

Running association and prediction on the middle clock fundamentally changes how an agent interacts with memory.

Most contemporary agent architectures treat recall as a deliberate, slow-loop step: the model identifies an information gap, issues a vector search, reads the retrieved chunks, and synthesizes. But human recall is overwhelmingly involuntary. An error string appears on screen, and the memory of last month’s resolution surfaces unbidden.

In a multi-clock system, associative retrieval runs continuously against the middle clock. Most recalled memories simply decay inside the situation model. A few alter intermediate predictions. Very few ever get serialized into a frontier context window. As argued in The Substrate Post, models are swappable commodities; history and state are not. History lives outside the model loop; the loop merely reads from it.

Prediction is how the middle clock suppresses unnecessary compute. Keys jingling in a pocket predicts departure. A failed compilation predicts log inspection. A colleague appearing in the doorway predicts that conversational boundaries have changed. A confirmed prediction requires zero downstream action. A violated prediction generates surprise, and surprise is one of the few legitimate signals that justifies spending tokens on deliberate System 2 thought.

Crucially, the clock boundary is also the privacy boundary.

system-0 system-1 and system-3 agent diagram

Raw audio streams, desktop captures, face embeddings, and the situation model remain entirely local. Durable memory promotion happens on-device. The slow model receives only a sanitized, purposeful brief, and the compilation of that brief is governed by who is physically present.

When an operator is alone in their office, the situation model permits a broad disclosure envelope. The moment another person walks into the room, a personal AI’s contextual access control restricts what can be surfaced: certain memories remain queryable for internal inference, but are flagged as ineligible for speech synthesis or UI rendering. Knowing a fact and being authorized to disclose it to a specific room are two entirely separate concerns.

Calling a model was already a data exposure decision when the payload was a single prompt. It becomes an unacceptable hazard when the payload encompasses continuous real-time audio and video. If the agent is the model, there is no place to enforce privacy before the call occurs. If the agent is the mesh, privacy filtering is just another subsystem running on the middle clock.

The Agent Is the Mesh

In our earlier reference architecture, the partition was horizontal: execution, reasoning, agency. Deterministic code at the bottom, models where structure breaks down, and an execution loop where goals outlive individual turns.

Within the agency layer itself, the architecture is not a monolithic ladder with a massive model sitting at the peak. It is a decoupled mesh:

  • System 0: Reflexes, interlocks, window focus, timers.
  • System 1: Continuous perception, transient situation tracking, associative recall, prediction, closed-loop judgment.
  • System 2: The deliberative frontier loop: multi-step planning, complex retrieval, and novel tool synthesis.
  • Effectors: The terminal, browser, system APIs, voice synthesizers, and external agents.

System 2 is an expensive capability to be bought sparingly. The boundaries between these layers are fluid: an inference that requires a frontier call today might run as a Jev pass or a local 3B model next year, and become a deterministic rule shortly after. But the foundational architecture is defined by the capacity to choose. An agent that lacks the machinery to decline deliberative thought is merely a chatbot strapped to an aggressive polling loop.

The industry placed the frontier model at the center because modern agents originated from chat interfaces, and our vocabulary followed our interfaces. Context was bound to the model. Tool use was bound to the model. Memory was bound to the model. Periodic schedulers did not change that structure; they simply set an alarm clock on the same monolithic loop.

Robotics never had the luxury of that illusion. A personal agent operating in the real world does not have it either. The model called by the system might be a local 9B on the desk, a specialized classification head, or an API call to Claude or GPT. Sometimes it is none of them.

The agent is not the language model. The agent is the mesh that decides what runs, what sleeps, and what deserves attention.